One month on my homelab LLM ran at nine tokens per second — measured, plotted, and explained across three posts. Then one weekend changed everything: the same model file, byte for byte, on a different engine, at forty-seven. What Strata does differently, which of my published diagnoses did not survive it, and the traps (buffered streams, thinking budgets, VRAM cliffs) between you and the same numbers.