Lab notes and session write-ups from real runs on the home stack — not a product blog. When a run had measured decode speed, the card lists tok/s. Coding stays scripts and repos; Builds is what actually got measured.
Want tok/s as the headline? See sister site toklanes.com.
Local LLM lanes: speed, capacity, and the hybrid that isn't
2026-09-09 · 5080 · B70 · Spark · Halo · Mac
One chart for fit and tok/s — who owns the ≤32GB speed lane vs the 128GB capacity lane.
Read →Quality tied — speed and VRAM decide
2026-09-08 · local LLM · RTX 5080
Gemma 12B vs Qwen 27B IQ4 on the same coding tasks. Same score, different daily-driver feel.
Read →FreeToken on WSL2 — mid MoE yes, Flash-Next no
2026-09-07 → 2026-09-08 · MoE
35B-A3B smoked clean on 16GB. Flash-Next did not. Negative results count.
Read →Same 27B, three quants — only one clears the bar
2026-09-08 · Unsloth · Ollama
IQ4_XS stayed full-GPU. Library Q4 and the YouTube layer recipe spilled.
Read →When the model “thinks,” the editor sees blank
2026-09-02 → 2026-09-08 · VS Code
Answers landed in reasoning fields. content stayed empty. Think-off aliases fixed the UX.
Read →When the research swarm reports back
2026-09-09 · last30days
One word (Intel), many public sources, one usable hardware map. The rollup is the point.
Read →