Ordo Take-Home — Glossary

Every acronym used in the repo, with a meaning and a real example from this project.

Acronyms & terms

TermMeaningExample from this project
VLMVision-Language Model — reads images + text, answers in text.Qwen2-VL-2B, SmolVLM-256M
GGUFllama.cpp's single-file weight format (model + tokenizer, memory-mappable).SmolVLM-256M-Instruct-Q8_0.gguf — 167MB
mmprojMultimodal projector — the vision-encoder weights file, loaded alongside the text GGUF.mmproj-Qwen2-VL-2B-Instruct-Q8_0.gguf — 677MB
f16 / Q8_0 / Q4_K_M / Q2_K / IQ2_XXSWeight precisions: 16-bit float, 8-bit, 4-bit (K-quants), 2-bit; IQ = imatrix-calibrated 2-bit.Q4_K_M: 940MB, 81% AR · IQ2_XXS: 487MB, 53% — the cliff
imatrixImportance matrix — calibration data used to choose what a 2-bit quant should preserve.computed over wikitext-2 for IQ2_XXS
ARAnswer Recall — fraction of ground-truth tokens present in the prediction. The chosen accuracy metric.GT "JIGSAW PUZZLE" ⊆ "The label says Jigsaw Puzzle" → 1.0
EMExact Match — prediction must equal the GT verbatim. Too strict here (0%).reported as the "strict check"
TTFTTime To First Token — what the user feels before the answer starts.Qwen @768: 5.8s · SmolVLM @512: 611ms
encode / prefill / decodeThe three measurable stages: vision tower pass · prompt through decoder · one-token-at-a-time generation.Qwen: 4124ms / 1485ms / 650ms
tok/sTokens per second — throughput for prefill/decode.Q4_K_M decode ~78 tok/s on Mac
RSSResident Set Size — peak physical RAM used during inference.Qwen 1.13GB vs SmolVLM 355MB
KV cacheKey-Value cache — stored attention state; grows with tokens, reused across turns.n_ctx 8192 · turn-2 reuses KV
n_batch / n_ctxTokens per batch (throughput knob) / context window size.2048→4096 bought prefill −16%
MetalApple's GPU API — where all inference actually runs (not CPU, not NPU)."GPU via Metal, yes; NPU, no"
ANE / NPUApple Neural Engine / Neural Processing Unit — the phone's AI chip; llama.cpp doesn't use it.the CoreML/ANE path is the untapped 4x encode lever
MTMDllama.cpp's Multi-Modal (mtmd) inference API — tokenizes images + text into chunks.chunks=TITITIT = 3 image chunks per SmolVLM photo
LoRALow-Rank Adaptation — tiny trainable matrices (rank 8) added to frozen weights; fine-tunes for ~1% of full-training cost.300 iters on eval-style data → AR 38% → 65%
MLXApple's ML framework for Apple Silicon — used for LoRA training on the Mac.lora/mlx-model/ (fp16 conversion)
FAFlash Attention — memory-efficient attention kernel.on/off tried; no fix for the A16 tower bug
EOGEnd Of Generation token — the model's "I'm done". Instant EOG = silent failure.un-strdup'd role → confident instant EOG
CACHE_HITDebug marker — vision embeddings reused from the pre-computed cache instead of re-encoding.multi-turn: enc=0ms, TTFT 200ms
GT (DRAFT)Ground Truth — the expected answer per photo. DRAFT = model-written, awaiting human verification.the one remaining task before submission
HEICApple's photo format — the raw captures were converted to JPEG for eval.sips -s format jpeg
JSONL / CSVLine-delimited JSON / comma-separated values — the raw results formats.results/device_eval_smolvlm_18.jsonl
devicectlApple's CLI for installing/launching/pulling files on a physical device.device copy from … Documents/results.jsonl
Reading the results table — every row in docs/report.md §4 is: text-quant × vision-quant, model size, AR, then the three stages (encode/prefill/decode) + TTFT + peak RSS. The story the numbers tell: encode dominates TTFT, the encoder tolerates quantization, and decode is a token-count problem.