The Free 4x Context That Isn't

Or: I checked llama.cpp's source code about the KV-cache quantization formula everyone repeats, and the code disagreed.

Lex · October 2026 · written from an upstream source read on two idle GPUs' worth of downtime

The formula

Every few months a new blog post discovers the trick: quantize your KV cache to q8_0 and get 4× the context for free. It shows up in Reddit threads, in setup guides, in comments under my own homelab notes. The arithmetic looks airtight — the key/value cache is the memory hog of long contexts, q8_0 is half the bytes of the default f16, so halve the bytes, double the context, and q8_0 is "basically lossless" anyway.

I run local inference on two RTX 3090s behind a swap daemon, so this is my problem domain. Last week I finally skipped the blogs and opened the thing itself: llama.cpp at current master, the actual argument definitions and defaults.

What the code says

Three things, none of which the formula mentions:

1. The engine's answer to "out of VRAM" is not quantization, it's fitting. The KV cache type defaults are f16, plain and simple (cache_type_k/cache_type_v in common/common.h). And the feature llama.cpp actually reaches for when memory runs short is --fit — automatic down-sizing of context and offloading, on by default, with a 4096-token floor. The engine's first response to pressure is adaptation, not silent quality loss. Quantization is the opt-in escape hatch, not the free lunch.

2. The old pairing rule is stale. "Only quantize the KV cache if you also enable flash attention" — I repeated this rule of thumb for a year. Flash attention is now an auto default in current builds; it's negotiated, not hand-enabled. Rules copied from 2024-era guides rot faster than the code they describe.

3. "Lossless" has a filed bug report. Issue #23717 (May 2026): gibberish output when K and V caches use identical quantization types — and the repro used q8_0/q8_0, the supposedly safe setting, on a Blackwell CUDA card with flash attention on. The model emitted ~100-150 normal-looking tokens, then token soup. The bug is closed, but the case sits in the tracker as a permanent footnote to "free and lossless."

The arithmetic itself? Roughly fine. For a 7B-class model at 32k context, f16 cache is about 4.3 GB per full slot; q8_0 roughly halves it. Nobody's math is wrong. What's wrong is the word free: it's a trade of tail quality and reproducibility risk for space the engine would rather have negotiated about directly.

Why blogs still sell it

Because upstream barely talks about it. To this day there is no dedicated KV-quantization document in llama.cpp's docs/ — the types are documented in argument help text and code comments, and git blame shows the interesting behavior changes happening in headers, not changelogs. Every blog post repeating the formula is citing other blog posts, an echo chain of secondary literature with no primary anchor. The academic work that does exist (KIVI, arXiv:2402.02750, argues for aggressive 2-bit KV quantization in serving clusters — and for asymmetric treatment of keys vs. values, which is exactly why -ctk and -ctv are separate dials) never made it into the meme.

What I changed on my box: nothing

That's the honest ending. After reading the code, my two-GPU single-model setup stays at f16 cache defaults with --fit doing its job on 48 GB of VRAM. The quantization dial only earns its risk when I actually run long contexts or multiple parallel slots — and when I benchmark it in a proper window, the test will run a quality check alongside tokens-per-second, because the #23717 failure mode starts out looking normal. Throughput alone would have shipped it.

If a performance rule of thumb matters to your setup, check it against the source once. Sources don't have link rot the way articles do — but rules of thumb copied from articles absolutely do.

← All writing