| Term | Meaning | Example from this project |
| VLM | Vision-Language Model — reads images + text, answers in text. | Qwen2-VL-2B, SmolVLM-256M |
| GGUF | llama.cpp's single-file weight format (model + tokenizer, memory-mappable). | SmolVLM-256M-Instruct-Q8_0.gguf — 167MB |
| mmproj | Multimodal projector — the vision-encoder weights file, loaded alongside the text GGUF. | mmproj-Qwen2-VL-2B-Instruct-Q8_0.gguf — 677MB |
| f16 / Q8_0 / Q4_K_M / Q2_K / IQ2_XXS | Weight precisions: 16-bit float, 8-bit, 4-bit (K-quants), 2-bit; IQ = imatrix-calibrated 2-bit. | Q4_K_M: 940MB, 81% AR · IQ2_XXS: 487MB, 53% — the cliff |
| imatrix | Importance matrix — calibration data used to choose what a 2-bit quant should preserve. | computed over wikitext-2 for IQ2_XXS |
| AR | Answer Recall — fraction of ground-truth tokens present in the prediction. The chosen accuracy metric. | GT "JIGSAW PUZZLE" ⊆ "The label says Jigsaw Puzzle" → 1.0 |
| EM | Exact Match — prediction must equal the GT verbatim. Too strict here (0%). | reported as the "strict check" |
| TTFT | Time To First Token — what the user feels before the answer starts. | Qwen @768: 5.8s · SmolVLM @512: 611ms |
| encode / prefill / decode | The three measurable stages: vision tower pass · prompt through decoder · one-token-at-a-time generation. | Qwen: 4124ms / 1485ms / 650ms |
| tok/s | Tokens per second — throughput for prefill/decode. | Q4_K_M decode ~78 tok/s on Mac |
| RSS | Resident Set Size — peak physical RAM used during inference. | Qwen 1.13GB vs SmolVLM 355MB |
| KV cache | Key-Value cache — stored attention state; grows with tokens, reused across turns. | n_ctx 8192 · turn-2 reuses KV |
| n_batch / n_ctx | Tokens per batch (throughput knob) / context window size. | 2048→4096 bought prefill −16% |
| Metal | Apple's GPU API — where all inference actually runs (not CPU, not NPU). | "GPU via Metal, yes; NPU, no" |
| ANE / NPU | Apple Neural Engine / Neural Processing Unit — the phone's AI chip; llama.cpp doesn't use it. | the CoreML/ANE path is the untapped 4x encode lever |
| MTMD | llama.cpp's Multi-Modal (mtmd) inference API — tokenizes images + text into chunks. | chunks=TITITIT = 3 image chunks per SmolVLM photo |
| LoRA | Low-Rank Adaptation — tiny trainable matrices (rank 8) added to frozen weights; fine-tunes for ~1% of full-training cost. | 300 iters on eval-style data → AR 38% → 65% |
| MLX | Apple's ML framework for Apple Silicon — used for LoRA training on the Mac. | lora/mlx-model/ (fp16 conversion) |
| FA | Flash Attention — memory-efficient attention kernel. | on/off tried; no fix for the A16 tower bug |
| EOG | End Of Generation token — the model's "I'm done". Instant EOG = silent failure. | un-strdup'd role → confident instant EOG |
| CACHE_HIT | Debug marker — vision embeddings reused from the pre-computed cache instead of re-encoding. | multi-turn: enc=0ms, TTFT 200ms |
| GT (DRAFT) | Ground Truth — the expected answer per photo. DRAFT = model-written, awaiting human verification. | the one remaining task before submission |
| HEIC | Apple's photo format — the raw captures were converted to JPEG for eval. | sips -s format jpeg |
| JSONL / CSV | Line-delimited JSON / comma-separated values — the raw results formats. | results/device_eval_smolvlm_18.jsonl |
| devicectl | Apple's CLI for installing/launching/pulling files on a physical device. | device copy from … Documents/results.jsonl |