Same model, same file, four times the speed: moving my homelab LLM off llama.cpp
For one month, one number followed me around: 9,2. Tokens per second, out of a machine I know down to its fan curve — a 2016 Xeon, 96 gigabytes of ECC RAM, an RTX 3060 that never asked for any of this. Chasing Flash-Next was the story of fitting the model in at all. Eighteen Cores and a Stubborn Plateau was the story of making it acceptably slow. And then, over one weekend, the speed quadrupled — with the same model, the same quantized file, byte for byte. Nothing smarter arrived. Only the engine changed.
This post is about Strata, why I was wrong twice about where my bottleneck lived, and what it takes to run a 177-billion-parameter mixture-of-experts model on hardware that predates the model by a decade.
The state of the art, as of my own basement#
Recap for the impatient, because the recap matters for the punchline. The model is Qwen3.8-Flash-Next, a 177B MoE with six billion active parameters per token and, as its most peculiar feature, a 51-billion-parameter n-gram lookup table that can live on SSD instead of in RAM. llama.cpp runs it on my box after a lot of careful tuning: roughly 14 tokens per second of generation at short context after the CPU upgrade, six and a half at 65k, and a prompt-reading plateau of about 215 tokens per second that no amount of CPU power could move.
That plateau is the part I want to revisit. The eighteen-core post concluded, with charts and a straight face, that the plateau was GPU-bound: the 3060 could not prefill faster, full stop, and the next euro for speed had to be a GPU euro.
That conclusion lasted three days.
Three dead ends, or: how I spent a week not installing Strata#
The honest path to the new engine ran through three failures, and they are worth recounting because each one taught something the documentation would not have.
The speculative decoding that wasn't. The blog's "watch item" for the time was MTP — multi-token prediction, a small second head attached to the model during training that guesses the next few tokens so the big model only has to verify them. llama.cpp grew the plumbing (--spec-type draft-mtp, no more of the OOM that killed my first attempt), and then the check in my own script caught the real problem: zero MTP tensors in any of our GGUFs. Not a corruption, a design decision — the upstream HF-to-GGUF converter simply dropped the head, and every popular quantization, mine included, inherited that omission. The community answer was a standalone draft file, 2,5 gigabytes, that you hang off the running model with -md. I tried. The engine loaded my 95-gigabyte model with affection and then died in the first microsecond of draft loading, expecting a tensor layout only a patched fork uses. Dead end — but the useful kind: the feature works, the files are the problem.
The free trick that cost half my speed. llama.cpp also ships n-gram drafting: no trained head at all, it just looks up the recent context for repeats and proposes them. Costs nothing, I said, and measured it on the same stack with a proper A/B. Acceptance rate: 26 to 39 percent — every rejected proposal still costs a verification round — and long-context generation did not merely fail to improve, it dropped from 24,9 to 12,9 tokens per second. Speculation is only a win when the drafter is nearly free and mostly right. Both conditions matter, and neither was met.
The download that quietly downloaded the wrong thing. By then I had moved to GSQ-RCO quantizations — the mixed-precision work from IST Austria that gives up almost nothing for a smaller file, and which I keep citing because the quantizer I had been trusting for quality, 89,5 on their test battery versus 92,57 for the IQ3 tier, is their work too. When I scripted the first download of the new quant, the verifier came back with one shard missing, and it took me an embarrassing number of minutes to see that passing filenames and --include patterns in the same command makes the downloader silently ignore the patterns. Fix the script, learn nothing. Fix the script properly, learn a little.
Which brings us to the thing I should have tried first.
Strata is not a faster llama.cpp. It is a different bet.#
Strata is an MIT-licensed inference engine, nineteen thousand stars at the time of writing, built by one person and a fast-growing community for exactly one model family — this Qwen3.8-Flash-Next, and the quantizations the same Austrian lab produces. The README does not market it as "faster LLM inference". It says, more interestingly: run a 125-billion-parameter model on your own gaming PC.
The bet is this. llama.cpp's design centre is "the model lives in one place — RAM or VRAM — and whatever does not fit is your problem". Strata's is that a MoE with 24,576 experts and a giant lookup table was never one memory object to begin with. So it splits the model across all four tiers your PC actually has, and puts each component where it is cheap to serve:
- the most-used few thousand experts live on the GPU, kept hot by a routing profile plus live adaptation to whatever your current conversation touches;
- every expert lives in pinned system RAM;
- the CPU computes the experts the GPU missed, a few percent of them per token, over PCIe if that is cheaper than waiting for the cache;
- and the 27-gigabyte n-gram table stays where the model card always said it belongs: on the SSD, read sparsely, streamed in whole-blob prefetches.
On top of that it carries its own MTP implementation — and crucially, it fetches the missing draft head from the original checkpoint itself, quantizes it to q2_0 in a few minutes on first start, and uses it. The tensor that my GGUF never had, Strata brings along. On a 12 GB card, the expert cache I measured holds 2,341 of 24,576 experts, under four gigabytes, and my hit rate during generation sits at 66 percent with three and a half percent served over PCIe and, the number that matters most, zero reads from the file tier. My RAM stayed warm and my SSD stayed boring.
It answers OpenAI and Anthropic APIs on one port, has a browser app with a live monitor that shows phase, tokens per second, VRAM and the expert cache state, an MCP server so an assistant can start and stop it, and a one-line calibration pass (--calibrate) that measures the four engine knobs that depend on the machine — how much missed-expert work goes to the GPU over PCIe versus the CPU, draft confidence, thread pool, cache adaptation — instead of shipping defaults tuned on a Ryzen and a 5070.
What you give up: it is young (251 open issues are visible and counted), it serves one model family, it holds one request at a time by default, and — this one surprised me — it buffers: a streamed response arrives as one block at the end rather than as a trickle. My opencode sessions do not care. Staring at a terminal waiting for the first token does.
The numbers, and the humiliation that comes with them#
Same file — the GSQ-RCO IQ3_XXS quantization, 44-gigabyte first shard. Same GPU, same CPU, same RAM, same wall of the same basement. Measured with Strata's own per-request accounting (prompt_ms, decode_tok_s), not my stopwatch, and verified against its HTTP metrics endpoint the same way I learned to distrust the plug.
| Phase | llama.cpp, 2697 v4 | Strata, same box | Factor |
|---|---|---|---|
| decode, short context | 14,3 t/s | 47,4 t/s | 3,3 |
| decode, 8k context | ~13 t/s | 40,8 t/s | 3,1 |
| decode, 30k context | 11,5 t/s | 41,7 t/s | 3,6 |
| prefill, plateau | 215–225 t/s | 628–653 t/s | 3,0 |
| draft acceptance (MTP) | n/a (no head in file) | 83 % on code prose | — |
A four-thousand-token answer now arrives in about ninety seconds instead of five minutes. The full 8000-token stress run — a deliberately verbose analysis prompt, max budget, temperature low — held 37 tokens per second end to end. And the plateau: 215, the number I had written in two posts as a GPU wall, fell to 630 on the same GPU. The plateau was never the silicon. It was plumbing in the engine, and I published the wrong diagnosis twice, with confidence. Both posts stand, and both now carry a quiet edit; the eighteen-core conclusion "the next euro is a GPU euro" is, as of this weekend, wrong in an instructive way: the next euro was free, and it was a different runtime.
Two honest asterisks. The comparison is same-file for generation, but generation is where Strata's cache design does most of its work, and a different quantization would shift every number; and the energy meters did not run alongside this weekend's benchmarks — llama.cpp measured roughly ten joules per generated token at these speeds, Strata's watt-figures are pending, and I refuse to quote what I have not caught on the plug.
The traps I fell into, so you can skip them#
Thinking is a budget problem, not a model problem. Twice on this engine, a request with an 8000-token allowance returned zero answer and 8000 tokens of reasoning — the model thought until the bill ran out, and the response was, technically, complete. Strata's log tells you exactly this and names the fix, and the fix has two halves: a per-request reasoning_budget_tokens (mine caps thinking at 3000 and leaves about five thousand for the reply), and fit_max_tokens, which silently shortens an over-eager max_tokens to the room the context has left, instead of answering the agent that always asks for its full budget with an HTTP 400. My opencode config now sends the budget through its model options, and the request I re-ran afterwards came back with, of all things, an answer.
Warm starts are a thing. The first start pinned forty gigabytes of experts into RAM at 0,28 GiB per second — about three minutes, cold. The same load one restart later ran at a good chunk of the drive's read ceiling; the page cache remembers what the engine fetches. If your first boot after a reboot is slow, that is the disk, not the engine.
Free VRAM is a cliff, not a slope. A community measurement, which I now believe personally: with the expert cache sized so that zero MiB of VRAM stays free, decode collapsed from 102 to 14 tokens per second while the cache hit rate went up. The fix is boring and effective — leave the cache on auto, or reserve a gigabyte. My start banner reports 467 MiB free and asks for more; it gets a hearing.
The drafter reads English. Strata's MTP head proposes tokens from a trimmed vocabulary — English and code, CJK scripts since a recent release, French on request. For German, my journal agent's native tongue, acceptance will lean on the prompt-lookup drafter instead. The engine ships the tool to build a custom subset from your own corpus; I have not run it yet, but a vocabulary grown from years of my actual text is exactly the kind of lever this box rewards.
What a token costs, once more#
I promised a long time ago that the cost block in my editor config would carry lab numbers, so here they are, restated against four times the speed. The engine does not change the tariff arithmetic much: generation at forty tokens per second is, at my rates, still dominated by depreciation of the parts, not watts — roughly two cents per million generated tokens of electricity, and a little over two dollars of amortized hardware. The cloud API for the same model sells generation for forty-seven cents a million. Nothing about this weekend made self-hosting cheaper. What it made cheaper is the thing the meters never measured: waiting. At nine tokens per second the loop where the agent proposes, I read, I reject, I re-prompt, was slow enough that I started batching my attention elsewhere; at forty, the conversation is a conversation. The whole argument for a homelab model has always been privacy, availability and the right to waste tokens. Speed was just the thing that made the argument awkward to make.
The three hundred euros of dedicated parts — the 3060, the memory, the twenty-five-euro CPU, the scratch SSD — are amortized over two years either way. What changed is that I now want to use them at forty instead of nine.
If you have twelve gigabytes of GPU and too much RAM#
Three sentences of recommendation, then. If you are running this model family — or thinking of it — on consumer hardware with 32 gigabytes of RAM or more, give Strata an afternoon: same files, same quantizations, the GGUFs you already own, a browser page that shows you what the model is doing. Keep llama.cpp installed anyway; it remains the Swiss army for every other model and for the sanity check, and mine sits one compose-file away, deliberately unloved, port and RAM mutually exclusive with the new thing by design.
And when you have your first numbers: run --calibrate, then look at the expert cache hit rate and the PCIe share before you believe anything, including me. The single most valuable thing this whole one-month detour has produced is not the four times the speed. It is that I stopped, repeatedly, being able to explain the slowness — and had to measure the layers until the story held. 9,2 was never the model's fault. It was mine, and my tools', and finally, three posts later, nobody's.