Operating a Model Zoo on Two Consumer GPUs

Or: what a homelab inference stack actually spends its time on — not tokens, but the software that decides which model gets the VRAM next.

Lex · October 2026 · written from four months of idle-machine research sessions on my own homelab — every claim here traces to a dated session log, and the parts I never measured are labelled as exactly that

The setup, and the honest confession

Two RTX 3090s. On them: llama-swap as a swap proxy, with a vLLM instance behind it serving one resident model, twenty-four seven. The machine has been running this way for months, and here is the confession the "model zoo" framing has to survive: there is no zoo. One model lives in the VRAM permanently (ttl: 0 — never auto-unload), and the entire multi-model machinery — the swap proxy, the API, the config language — waits, fully armed, for a second model that almost never comes.

The card I originally wrote myself promised "why model swaps cost minutes." That's the honest correction this essay opens with: I never measured it. What I have instead is four months of notes on everything I checked around the swap path — and the recurring finding that in a homelab LLM stack, the interesting failures aren't in the tokens. They're in the version you're on, the capabilities nobody advertises, and the quantization folklore nobody checks against source.

Incident one: you can't tell what you're running (until you grep the UI bundle)

First lesson, delivered by a session that started with a deceptively simple question: which llama-swap version is on the box? The documented answer in my own notes, from an earlier session, was "not readable from outside — deliberately not guessed." That answer was wrong, and wrong in an instructive way.

The endpoint existed the whole time: GET /api/version. My earlier probe had simply tried guessed paths, gotten 404s, and — to my credit at the time — refused to guess. What I should have done was open the UI's own JavaScript bundle, which names its API calls. The real answer came back immediately: {version: v240, commit: 6b5320d, build: 2026-07-15}, and the commit hash tied to exactly tag v240 on the tag reference. Twenty-one releases behind, none of them breaking for this instance (the consumers here touch only /v1/chat/completions and /metrics, grep-confirmed).

The method lesson generalises beyond homelabs: an empty result from a query you guessed is not a negative finding. A 404 on a path you invented proves only that you invented it. The positive control — find the real route list first, then conclude absence — is the difference between "not exposed" and "not looked for."

Incident two: the feature chain that wasn't

With version pinned, the next question was whether an upgrade was worth a maintenance window — and the changelog dangled what looked like a magnet: cmd/vllm-wrapper, since v243, promising to put vLLM to sleep instead of cold-starting it, with — in the wrapper's own README words — "near-instant wake-ups."

The feasibility check turned that single line into a five-link condition chain, and I verified each link separately, read-only, without touching the running service:

  1. The wrapper is not in the upgrade. It's a separate Go binary built from source; the release tarballs contain exactly one thing — the llama-swap binary itself. (Checked against the release API for v261, the current release: seven assets, not one with "wrapper" in the name. A mid-essay draft of this paragraph claimed the pattern "changed again" by v263 — a tag I made up from memory and the release API then refused to confirm. Published corrections beat confident errors; that's arguably the whole thesis of this site.)
  2. The sleep endpoints aren't reachable on this box. A read-only probe of POST /sleep answered with a Go-style 404 from the swap proxy itself — the vLLM behind it binds to localhost, so there's no path in from the LAN today.
  3. The endpoints are dev-mode endpoints. The official vLLM sleep-mode docs require VLLM_SERVER_DEV_MODE=1 and warn, in their own words, that these endpoints "should not be exposed to users." A production swap path hanging off a dev-mode flag is a design smell worth naming.
  4. Level-1 sleep needs host RAM for the weights. Sleep level 1 offloads weights to CPU RAM; the metrics said ~33 GB free, and the model's size is, from VM209, not something I can read. Unbelegt — deliberately not guessed, third time running.
  5. The wrapper and the vLLM router disagree about HTTP. The wrapper posts the sleep level as a JSON body; the current router reads it as a query parameter and ignores bodies. A static finding, labelled as such, but the kind of thing that sleeps (level-1 pun intended) through a "successful" integration test.

Which is how a one-line upgrade magnet decomposed into a five-condition project with zero conditions met. The "near-instant wake-ups" got the label it deserved: the author's README phrasing, not a measurement. Waking from level-1 sleep means copying weights from host RAM back to VRAM — surely faster than cold start, plausibly seconds, and that's exactly the sentence where my data ends and folklore begins.

Incident three: the folklore the source code doesn't recognise

The most instructive session was about something else entirely: the community formula that quantising the KV cache to q8_0 buys you "4× the context for free." It's everywhere. It reads like a fact. It is not in the source code.

What llama.cpp's actual source says, read at the line rather than via blog posts: the cache type defaults are f16 (common/common.h, cache_type_k/v) — q8_0 is one of nine permitted options, not the reference path; flash attention is negotiated auto now, so the old "only quantise KV with -fa" pairing rule is stale; and the engine's real, default-on answer to VRAM pressure is --fit, which shrinks context rather than quality. The formula's arithmetic is roughly right in magnitude — but the "free" and the "always safe" both died on a filed bug: issue #23717 documents gibberish output after ~100–150 tokens on CUDA hardware with — yes — q8_0/q8_0, the supposedly lossless setting.

The meta-finding stung more than the finding: every confident blog post on the formula turned out to be citing another blog post. Nobody was citing the engine. In a field this young, reading the source isn't rigor, it's the only way to have evidence at all.

What a swap actually costs (the answer I can and can't give)

So: what does a model swap cost on two consumer GPUs? Here is the state of my knowledge, labelled by kind:

The real economics, though, turned out to be invisible rather than slow. The capabilities docs note that vLLM — the very upstream behind the swap proxy — is "the thin one": it doesn't advertise context length, tools, or vision in /v1/models. Nothing about that costs a swap, but everything about client behaviour quietly depends on it: clients either guess capabilities or refuse to trust the listing, and in both cases a config block nobody maintains shapes traffic nobody planned. The swap that never happens is the interesting one: the model that stayed loaded because everything else got too complicated to load — and the reason was never seconds or gigabytes, but metadata.

What I'd tell the person starting this build

  1. Pin your versions from outside. If /api/version doesn't exist, the UI bundle knows the paths that do. "Not readable" usually means "not grepped."
  2. Read the feature magnet as a chain. "v243 adds sleep mode" decomposes into a build step, a dev-mode flag, a RAM budget, and an HTTP disagreement. Feature = chain; count the links before calling it a magnet.
  3. When a formula is everywhere, read the source once. The KV-cache folklore survived four years because reading common.h is less fun than repeating the formula.
  4. Label your numbers by kind. Measured, documented, folklore — the categories, not the conclusions, are what I actually trust myself to publish. The swap-costs-minutes sentence on my own site's card for months was folklore with a domain name.

Sources and dates: llama-swap /api/version (v240/6b5320d, read-only GET 01.–02.10.2026); release asset lists via GitHub API (v261, 7 assets — and the tag list stops at v261, which killed one invented sentence in this essay's draft); llama-swap TTL guide, docs/kb/guides/model-runtime/ttl-and-unloading.md (updated 25.08.2026); vLLM sleep-mode docs (docs.vllm.ai, 02.10.2026); llama.cpp common/common.h, common/arg.cpp, and issue #23717 (master, 01.10.2026). Full session logs with dated per-claim evidence: internal research reports, 30.09.–02.10.2026. All external sources fetched live during those sessions; the swap-costs-minutes sentence is, and remains, not measured.

← All writing