You run a well-regarded open model locally, and it feels… off. It forgets what you told it two messages ago. It's fast some days and crawls on others. It's tempting to conclude the local model just isn't as smart. Usually, it's not the model — it's four configuration traps, and all four are measurable.
If nothing sets a context size, the server quietly defaults to a few thousand tokens, no matter what the model actually supports. Longer conversations get clipped from the top before the model reads them — which looks exactly like forgetfulness. A model advertising 128k context can be running on 4k without a single warning. This one trap explains most "why did it forget?" complaints. How to confirm and fix it →
The fix for trap #1 — raising the context — sets up trap #2. The context window's memory (the KV cache) is reserved in VRAM up front. Push it too high and the model no longer fits on the GPU, so part of it runs on the CPU. Because generation needs every layer for every token, even a small spill tanks speed — we've measured a drop from ~48 tokens/sec to ~9. It looks like "the model got slow"; it's actually placement. The spill mechanics →
Forums love OLLAMA_KV_CACHE_TYPE=q8_0 as a free memory win. Sometimes it
is. On some GPU/backend combinations it forces the exact CPU spill from trap #2 that the
default setting didn't have — turning an "optimization" into a slowdown. The only way to
know for your card is to A/B it at your real context size.
When it helps and when it hurts →
Once the model is fully on the GPU and the context is right, generation speed is set by physics: each token reads the whole model out of VRAM, so your ceiling is roughly (memory bandwidth ÷ model size). If you're already near it, swapping backends or chasing kernel flags won't help — the real levers are a smaller quantization, a smaller model, or speculative decoding. Knowing you've hit this wall saves you days of pointless tinkering. The four causes, in order →
Every trap above is invisible until you measure — and measurable in minutes: sweep a few context sizes, check GPU placement at each, compare your tokens/sec against your card's bandwidth ceiling. Do that once and "the model feels dumb" turns into a specific, fixable number.