2026-09-07 → 2026-09-08 · local LLM · MoE · RTX 5080

FreeToken on WSL2 — mid MoE yes, Flash-Next no

Open-source FreeToken via WSL2 + CUDA on the 5080.

What happened

Worked: Qwen3.6-35B-A3B NVFP4 with MoE offload — engine logs ~130–150 decode tok/s; wall ~60–90 on short prompts; ~15.7GB VRAM plus host RAM for experts.

Didn’t: Qwen3.8 Flash-Next NVFP4 (~126GB on disk) never got usable KV on 16GB, even with CPU MoE and raised WSL RAM. A 27B NVFP4 FreeToken path hit a WSL/Blackwell load failure — inconclusive next to Ollama.

Cleanup deleted ~192GB of failed experiment caches. Daily path stayed Unsloth IQ4_XS on Ollama.

Takeaway

Negative results belong in Lab notes. Engine success is not the same as every model family fitting. Don’t burn a weekend chasing Flash-Next on 16GB.

← All Builds