Prefill vs decode, why memory bandwidth caps tokens/s, MoE, quantization and context limits — a crash course with real measurements from a 12 GB GPU and a very patient Xeon.
Prefill vs decode, why memory bandwidth caps tokens/s, MoE, quantization and context limits — a crash course with real measurements from a 12 GB GPU and a very patient Xeon.