What happened
Worked: Qwen3.6-35B-A3B NVFP4 with MoE offload — engine logs ~130–150 decode tok/s; wall ~60–90 on short prompts; ~15.7GB VRAM plus host RAM for experts.
Didn’t: Qwen3.8 Flash-Next NVFP4 (~126GB on disk) never got usable KV on 16GB, even with CPU MoE and raised WSL RAM. A 27B NVFP4 FreeToken path hit a WSL/Blackwell load failure — inconclusive next to Ollama.
Cleanup deleted ~192GB of failed experiment caches. Daily path stayed Unsloth IQ4_XS on Ollama.
Takeaway
Negative results belong in Lab notes. Engine success is not the same as every model family fitting. Don’t burn a weekend chasing Flash-Next on 16GB.
Sister site: toklanes.com — measured tok/s review-bench.