Kimi K3 and Mooncake: Moonshot AI Shipped the World's First Open 3T-Class Model on a KVCache-Centric Inference Engine That Gets 525% More Throughput by Treating Cache as the Primary Citizen
Every LLM serving paper optimizes for throughput. Mooncake optimizes for cache. The distinction sounds subtle. It is not. When you make KVCache (key-value cache, the memory structure that stores intermediate attention computations) the first-class citizen of your serving architecture, you stop thinking about GPU clusters as compute nodes and start thinking about them as a heterogeneous memory hierarchy.