← README (repo overview) Glossary Full report →

Ordo Take-Home — a Vision Model on a Phone

One page to understand the whole repo: what was built, how it flows, and what every acronym means.
Qwen2-VL-2BSmolVLM-256M / 500Mllama.cpp b10549 · MetaliPhone 15 (A16) · Mac M2 Pro
28
eval photos (own)
270
ladder runs
46%
best on-device AR (Qwen @768)
+27pts
LoRA accuracy gain
200ms
multi-turn TTFT (cached)
11.9MB
git repo size

1 · Architecture

Two measurement surfaces share one manifest (eval/eval.json) and one scoring script (eval/scripts/metrics.py). The Mac runs the heavy ladder; the iPhone app is both the production target and a second measurement surface.

Mac — M2 Pro (measurement & training)
eval/28 photos → photos_tiny/ (≤100KB)
eval.json (questions + GTs)
→
bench/tool/bench_vlm.cppllama.cpp harness · encode / prefill /
decode / TTFT per stage + RSS
→
bench/mac/run_bench.pyladder driver: 15 configs × 28 photos
→
results/ladder_full.csv270 rows of per-stage timings
→
bench/mac/report.pyAR / EM / F1 scoring
→
docs/report.mdthe deliverable (14 sections)
model/Qwen2-VL-2B: f16·Q8·Q4·Q2·IQ2 + mmproj f16/Q8/Q4
SmolVLM-256M·500M Q8 + mmproj
→
lora/ (MLX)build_dataset.py → run_lora.py (300 iters)
eval_lora_mlx.py (base vs adapter)
→
lora/adapters/rank-8 LoRA weights (37MB)
iPhone 15 — A16 (the actual product)
QwenBench appbundle com.ordo.qwenbench
→
ContentView.swiftUI · camera · autorun dispatch
→
BenchmarkEngine.swifteval · stress · soak · photo flow
→
MtmdContext.swiftllama.cpp MTMD bindings · vision cache ·
prepareImage (pre-compute) · runOnce
→
llama.frameworkstatic Metal build (b10549)
→
GPU (Metal)encode + prefill + decode
model_data/6 GGUFs bundled:
SmolVLM-256M·500M, Qwen2-VL-2B Q4, mmprojs
→
Resources/eval/eval.json + 512/1024 variants + photos
→
Documents/results.jsonl · diag.txt (pulled via devicectl)
⇅
same weights, same llama.cpp codenumbers transfer Mac ↔ phone
⇅
Key insight baked into this design: the vision encoder is separately measurable (encode_ms) from prefill and decode. That one decision — timing the three stages apart — is what made the whole analysis possible (the encoder is 87% of TTFT on Mac, and its A16 failure mode was found by inspecting its embedding outputs).

2 · User flow (in the app)

2a · The photo flow — "look at what you're looking at"

1Tap CameraUIImagePickerController
→
2Auto-compress≤768px · quality sweep
target <100KB JPEG
→
3Pre-compute visionprepareImage() runs the
tower while you type
→
4Type a question → Askanswers use cached embeddings
→ turn-2 TTFT ≈ 200ms
→
5Answer + timingenc / pre / dec / TTFT shown
under the thumbnail
ask again about the same photo → enc=0ms, CACHE_HIT on all 3 image chunks

2b · The test flows (buttons, kept as-is so anyone can re-run)

①Load modelpicks model + mmproj from
model_data/
②Eval set28 photos · question each
→ results.jsonl
③Stress 20x20 sequential queries
→ TTFT/thermal/mem per query
④Soak 10mincontinuous inference
→ thermal phases
⑤Exportshares results file
CLIautorun·512·1024·500·qwen·qwen-fast·multiturnlaunch args drive headless runs —
each re-install forces a fresh launch

3 · Task flow — how the assignment maps to the repo

P1Model on a phonellama.cpp b10549 · Metal
static framework build
→ model/fetch.sh
→
P2Trustworthy eval set28 own photos · difficulty/condition
tags · scripts/capture_wizard.py
draft_manifest · second_annotator
→
P3Quantization ladder5 text × 3 vision quants = 270 runs
→ where it breaks · encoder vs decoder
→
P4Under stress20 queries · thermals · memory
→ results/device_stress_*
→
P5LoRA fine-tuneMLX · 300 iters on eval-style data
→ AR 38% → 65%
→
✓Reportdocs/report.md — 14 sections
incl. "what surprised me"
Repo layout cheatsheet — model/ weights · bench/tool/ harness · bench/mac/ ladder + scoring · bench/ios/QwenBench/ the app · eval/ manifest + photos + integrity scripts · lora/ fine-tuning · results/ all raw measurements · docs/report.md the deliverable.

4 · LoRA fine-tuning (the assignment's final question)

The last part of the assignment: take the quantized model, LoRA fine-tune it on examples in the style of the eval task, and check whether fine-tuning recovers the accuracy that quantization cost. Answer here: it does — AR 38% → 65% on the same 26 items (+27 points).

Mac — MLX training pipeline
eval/eval.json28 items · question + GT
(skip empty-GT items)
→
lora/build_dataset.pyparaphrases questions ×2
→ 40 train + 6 val rows
→
lora/run_lora.pymlx-vlm trainer: 300 iters,
rank 8, alpha 16, lr 2e-5
language LoRA only (vision
frozen — tower OOMs)
→
lora/adapters/rank-8 LoRA weights (37MB)
→
lora/eval_lora_mlx.pybase vs adapter on the same
26 items → AR 38% vs 65%
lora/mlx-model/Qwen2-VL-2B fp16 MLX conversion
(from HF safetensors)
→
train + evalruns entirely on the M2 Pro (Metal)
300 iters ≈ 10 min · eval ≈ 5 min

Setup process (reproduce)

1Convertmlx_vlm.convert — HF safetensors
→ lora/mlx-model/ (fp16, unquantized)
→
2Build datasetlora/build_dataset.py — eval.json
→ train.jsonl (40) + val.jsonl (6)
paraphrased questions, 768px photos
→
3Trainlora/run_lora.py — 300 iters, batch 1,
grad-checkpoint, image capped 672px
(bigger = Metal OOM on 16GB)
→
4Evallora/eval_lora_mlx.py both — base vs
adapter, same prompt template
(apply_chat_template + num_images)
→
5Reportresults/lora_mlx_base.jsonl +
_lora.jsonl → report §10
Gotchas learned the hard way: (a) mlx-vlm 0.6.15 CLI dropped --export and --val-dataset — use the Python trainer API; (b) model.eval() breaks generation in 0.6.15 (empty outputs) — don't call it; (c) apply_chat_template(…, num_images=1) is required or the image is silently ignored (text-only hallucinations); (d) unfreezing the vision tower OOMs a 16GB M2 — language-LoRA only; (e) the GGUF-adapter seam for on-device LoRA is still pending (the 0.6.15 export path is gone).

5 · Glossary

Every acronym with its meaning and a real example — now on its own page: docs/glossary.html (25 terms: VLM, GGUF, mmproj, quants, AR, TTFT, KV cache, LoRA, MLX, ANE/NPU, MTMD, CACHE_HIT, …).