← The GG Blog

Why your local LLM feels dumber than ChatGPT (and how to fix it)

24 July 2026

You run a well-regarded open model locally, and it feels… off. It forgets what you told it two messages ago. It's fast some days and crawls on others. It's tempting to conclude the local model just isn't as smart. Usually, it's not the model — it's four configuration traps, and all four are measurable.

1. Your context is being silently cut to ~4k tokens

If nothing sets a context size, the server quietly defaults to a few thousand tokens, no matter what the model actually supports. Longer conversations get clipped from the top before the model reads them — which looks exactly like forgetfulness. A model advertising 128k context can be running on 4k without a single warning. This one trap explains most "why did it forget?" complaints. How to confirm and fix it →

2. The model is spilling onto your CPU

The fix for trap #1 — raising the context — sets up trap #2. The context window's memory (the KV cache) is reserved in VRAM up front. Push it too high and the model no longer fits on the GPU, so part of it runs on the CPU. Because generation needs every layer for every token, even a small spill tanks speed — we've measured a drop from ~48 tokens/sec to ~9. It looks like "the model got slow"; it's actually placement. The spill mechanics →

3. You copied a KV-cache tweak that backfired on your GPU

Forums love OLLAMA_KV_CACHE_TYPE=q8_0 as a free memory win. Sometimes it is. On some GPU/backend combinations it forces the exact CPU spill from trap #2 that the default setting didn't have — turning an "optimization" into a slowdown. The only way to know for your card is to A/B it at your real context size. When it helps and when it hurts →

4. You've hit the memory-bandwidth wall — and no tweak fixes that

Once the model is fully on the GPU and the context is right, generation speed is set by physics: each token reads the whole model out of VRAM, so your ceiling is roughly (memory bandwidth ÷ model size). If you're already near it, swapping backends or chasing kernel flags won't help — the real levers are a smaller quantization, a smaller model, or speculative decoding. Knowing you've hit this wall saves you days of pointless tinkering. The four causes, in order →

The pattern: measure, don't guess

Every trap above is invisible until you measure — and measurable in minutes: sweep a few context sizes, check GPU placement at each, compare your tokens/sec against your card's bandwidth ceiling. Do that once and "the model feels dumb" turns into a specific, fixable number.

Don't want to run the sweep by hand? ollama-tune measures all four on your actual machine and prints the exact config to fix them. The full report is free.