Eighteen Cores and a Stubborn Plateau: The Flash-Next CPU Upgrade Bench

Why this post exists#

Chasing Flash-Next ended with a scoreboard that blamed the CPU: 9,2 t/s decode, and — so I wrote — a prefill plateau of ~212 t/s caused by the "CPU compute ceiling" of a 2016 ten-core Xeon. The cure looked cheap: an E5-2697 v4 (18 cores, 36 threads, all-core turbo 2,8 GHz, DDR4-2400 instead of 2133) for about the price of a good dinner, dropping into the same LGA2011-3 socket. It arrived today. This post is the before/after — and a correction of my own diagnosis, because the numbers say the ten-core was only half the story.

What I predicted#

The model's experts (~50 GB of quantized weights) are dequantized and multiplied on the CPU, so doubling the AVX2 throughput (10→18 cores, +40 % clock) should lift decode from 9,2 to roughly 13–15 t/s, and a plateau of 212 t/s prefill at 10 cores should move to ~300 if it really was CPU-bound. Measured after the swap, with a fresh thread sweep (--fit off, -t 14, -tb 16 — the old "t8 wins" rule was made for 10 cores and didn't survive it):

ContextBefore (2630L, t8)After (2697 v4, t14)Δ
decode @ 2k9,2 t/s14,3 t/s+55 %
decode @ 33k7,9 t/s11,5 t/s+46 %
decode @ 65k6,8 t/s8,8 t/s+29 %
prefill plateau212 t/s215–225 t/s±0
vision (mmproj)ok13,6 t/sok

The decode prediction landed exactly. The prefill prediction did not land at all — the plateau didn't move a single token per second.

The plateau was never the CPU#

This is the interesting part. I doubled the cores, raised the batch threads, gave the memory controller a faster ceiling (2133 → 2400) — and prefill sat at ~215. A workload that is unresponsive to a 100 % CPU-compute increase is limited somewhere else: for qwen4exp, prefill runs the attention and the QSA sparse-indexer on the GPU (flash-attention on), and that is the 212 t/s wall. The old blog's "CPU compute ceiling" line — and the note in my production config — were wrong; I have fixed both. Practical consequence: the next euro for speed is a GPU euro (a 32 GB V100 moves both phases), not a CPU euro. The CPU upgrade still pays: decode is the phase that dominates real agent sessions, and +55 % is +55 %.

The same joules, again at the wall#

Same Zigbee plug, same method (disks spun down, page cache warm, model loaded). One honest caveat: instead of the 1 Hz logger I polled the plug via Home Assistant at chosen points inside long, steady phases — so read these as ±10 W spot values, not full-window medians.

PhaseRateNet draw (old → new)Energy per token (old → new)× 0,40 €/kWh (new)
idle baseline—87,5 W → ~101 W——
prefill (54k)215 t/s~105 W → ~135 W0,5 J → 0,6 J7 €-ct / Mtok
decode, fresh turns14,4 t/s~77 W → ~140 W8,3 J → 9,7 J1,08 € / Mtok
decode, 65k context9,2 t/s~122 W → ≤150 W~18 J → ≤16 J≤1,8 € / Mtok

Three observations. First, idle rose by ~13 W — the bigger package costs power even doing nothing; that is the price the 24/7 box pays for the burst performance. Second, fresh-turn decode now costs ~17 % more energy per token while being 55 % faster — energy per answer is roughly flat, but an answer arrives twice as fast, and end-to-end a job holds the box busy for less wall time. Third, and slightly embarrassing: prefill got less efficient per token (same speed, more watts) — a direct picture of the plateau being GPU-bound while the CPU now idles faster between memory feeds.

Gotchas on the bench ride#

  • Cold-first-row: the very first request after the container start measured 3,8 t/s prefill — weights and page tables still faulting in. Rows one to three of any sweep are a warm-up ramp, not data. (The 1 Hz power runs in the old post had the model pre-loaded; this one bit me because the box had rebooted for the swap.)
  • Thread rules don't port. "t8 beats t12" was measured on 10 physical cores. On 18, the optimum sat at -t 14 -tb 16; SMT threads still don't help this AVX2-bound loop, consistent with the old finding.
  • 145 W vs 85 W TDP: no throttling, no reset during ~30 minutes of full sweep on the old board's VRMs — but that is one sample on one cooler, not a guarantee.

Status and next#

The production server now carries -t 14 -tb 16, which should help with agentic use cases.

  1. GPU step (V100 32 GB class) — it is now provably the only lever left that moves both prefill and the TTFT that made this box feel slow.
  2. MTP speculative decoding — upstream PR #28243 is in review; with experts partially on the GPU and 18 CPU cores for the verify-batches, this model's decode could compound on top of the upgrade rather than eat it.

Measured on the same 2016 Unraid box: Xeon E5-2697 v4 (18C/36T), 96 GiB DDR4-2400 ECC, RTX 3060 12 GB. Wall meter: Zigbee plug via Home Assistant.