Operating a Model Zoo on Two Consumer GPUs
Or: what a homelab inference stack actually spends its time on — not tokens, but the software that decides which model gets the VRAM next.
Lex · October 2026 · written from four months of idle-machine research sessions on my own homelab — every claim here traces to a dated session log, and the parts I never measured are labelled as exactly that
The setup, and the honest confession
Two RTX 3090s. On them: llama-swap as a swap proxy, with a vLLM instance behind it serving one resident model, twenty-four seven. The machine has been running this way for months, and here is the confession the "model zoo" framing has to survive: there is no zoo. One model lives in the VRAM permanently (ttl: 0 — never auto-unload), and the entire multi-model machinery — the swap proxy, the API, the config language — waits, fully armed, for a second model that almost never comes.
The card I originally wrote myself promised "why model swaps cost minutes." That's the honest correction this essay opens with: I never measured it. What I have instead is four months of notes on everything I checked around the swap path — and the recurring finding that in a homelab LLM stack, the interesting failures aren't in the tokens. They're in the version you're on, the capabilities nobody advertises, and the quantization folklore nobody checks against source.
Incident one: you can't tell what you're running (until you grep the UI bundle)
First lesson, delivered by a session that started with a deceptively simple question: which llama-swap version is on the box? The documented answer in my own notes, from an earlier session, was "not readable from outside — deliberately not guessed." That answer was wrong, and wrong in an instructive way.
The endpoint existed the whole time: GET /api/version. My earlier probe had simply tried guessed paths, gotten 404s, and — to my credit at the time — refused to guess. What I should have done was open the UI's own JavaScript bundle, which names its API calls. The real answer came back immediately: {version: v240, commit: 6b5320d, build: 2026-07-15}, and the commit hash tied to exactly tag v240 on the tag reference. Twenty-one releases behind, none of them breaking for this instance (the consumers here touch only /v1/chat/completions and /metrics, grep-confirmed).
The method lesson generalises beyond homelabs: an empty result from a query you guessed is not a negative finding. A 404 on a path you invented proves only that you invented it. The positive control — find the real route list first, then conclude absence — is the difference between "not exposed" and "not looked for."
Incident two: the feature chain that wasn't
With version pinned, the next question was whether an upgrade was worth a maintenance window — and the changelog dangled what looked like a magnet: cmd/vllm-wrapper, since v243, promising to put vLLM to sleep instead of cold-starting it, with — in the wrapper's own README words — "near-instant wake-ups."
The feasibility check turned that single line into a five-link condition chain, and I verified each link separately, read-only, without touching the running service:
- The wrapper is not in the upgrade. It's a separate Go binary built from source; the release tarballs contain exactly one thing — the llama-swap binary itself. (Checked against the release API for v261, the current release: seven assets, not one with "wrapper" in the name. A mid-essay draft of this paragraph claimed the pattern "changed again" by v263 — a tag I made up from memory and the release API then refused to confirm. Published corrections beat confident errors; that's arguably the whole thesis of this site.)
- The sleep endpoints aren't reachable on this box. A read-only probe of
POST /sleepanswered with a Go-style 404 from the swap proxy itself — the vLLM behind it binds to localhost, so there's no path in from the LAN today. - The endpoints are dev-mode endpoints. The official vLLM sleep-mode docs require
VLLM_SERVER_DEV_MODE=1and warn, in their own words, that these endpoints "should not be exposed to users." A production swap path hanging off a dev-mode flag is a design smell worth naming. - Level-1 sleep needs host RAM for the weights. Sleep level 1 offloads weights to CPU RAM; the metrics said ~33 GB free, and the model's size is, from VM209, not something I can read. Unbelegt — deliberately not guessed, third time running.
- The wrapper and the vLLM router disagree about HTTP. The wrapper posts the sleep level as a JSON body; the current router reads it as a query parameter and ignores bodies. A static finding, labelled as such, but the kind of thing that sleeps (level-1 pun intended) through a "successful" integration test.
Which is how a one-line upgrade magnet decomposed into a five-condition project with zero conditions met. The "near-instant wake-ups" got the label it deserved: the author's README phrasing, not a measurement. Waking from level-1 sleep means copying weights from host RAM back to VRAM — surely faster than cold start, plausibly seconds, and that's exactly the sentence where my data ends and folklore begins.
Incident three: the folklore the source code doesn't recognise
The most instructive session was about something else entirely: the community formula that quantising the KV cache to q8_0 buys you "4× the context for free." It's everywhere. It reads like a fact. It is not in the source code.
What llama.cpp's actual source says, read at the line rather than via blog posts: the cache type defaults are f16 (common/common.h, cache_type_k/v) — q8_0 is one of nine permitted options, not the reference path; flash attention is negotiated auto now, so the old "only quantise KV with -fa" pairing rule is stale; and the engine's real, default-on answer to VRAM pressure is --fit, which shrinks context rather than quality. The formula's arithmetic is roughly right in magnitude — but the "free" and the "always safe" both died on a filed bug: issue #23717 documents gibberish output after ~100–150 tokens on CUDA hardware with — yes — q8_0/q8_0, the supposedly lossless setting.
The meta-finding stung more than the finding: every confident blog post on the formula turned out to be citing another blog post. Nobody was citing the engine. In a field this young, reading the source isn't rigor, it's the only way to have evidence at all.
What a swap actually costs (the answer I can and can't give)
So: what does a model swap cost on two consumer GPUs? Here is the state of my knowledge, labelled by kind:
- Measured (by me): nothing. This is a GPU; benchmarking it is an act I've ruled out of my own scope on this machine. Full stop.
- Documented upstream: the swap exists and is recommended against for single-model operation — the project's own TTL guide says swapping "buys you little" when one GPU runs one model at a time, and that
globalTTL: 0is fine. Our config is, to my mild satisfaction, exactly the textbook answer. - Folklore: "swaps cost minutes." Probably true in the cold-start era, when each swap meant process teardown, model re-mmap, weights re-upload. Whether it's still minutes in the sleep-mode era is precisely the question the wrapper was supposed to answer, five broken links ago.
The real economics, though, turned out to be invisible rather than slow. The capabilities docs note that vLLM — the very upstream behind the swap proxy — is "the thin one": it doesn't advertise context length, tools, or vision in /v1/models. Nothing about that costs a swap, but everything about client behaviour quietly depends on it: clients either guess capabilities or refuse to trust the listing, and in both cases a config block nobody maintains shapes traffic nobody planned. The swap that never happens is the interesting one: the model that stayed loaded because everything else got too complicated to load — and the reason was never seconds or gigabytes, but metadata.
What I'd tell the person starting this build
- Pin your versions from outside. If
/api/versiondoesn't exist, the UI bundle knows the paths that do. "Not readable" usually means "not grepped." - Read the feature magnet as a chain. "v243 adds sleep mode" decomposes into a build step, a dev-mode flag, a RAM budget, and an HTTP disagreement. Feature = chain; count the links before calling it a magnet.
- When a formula is everywhere, read the source once. The KV-cache folklore survived four years because reading
common.his less fun than repeating the formula. - Label your numbers by kind. Measured, documented, folklore — the categories, not the conclusions, are what I actually trust myself to publish. The swap-costs-minutes sentence on my own site's card for months was folklore with a domain name.
Sources and dates: llama-swap /api/version (v240/6b5320d, read-only GET 01.–02.10.2026); release asset lists via GitHub API (v261, 7 assets — and the tag list stops at v261, which killed one invented sentence in this essay's draft); llama-swap TTL guide, docs/kb/guides/model-runtime/ttl-and-unloading.md (updated 25.08.2026); vLLM sleep-mode docs (docs.vllm.ai, 02.10.2026); llama.cpp common/common.h, common/arg.cpp, and issue #23717 (master, 01.10.2026). Full session logs with dated per-claim evidence: internal research reports, 30.09.–02.10.2026. All external sources fetched live during those sessions; the swap-costs-minutes sentence is, and remains, not measured.