Kimi K3's Architecture: Kimi Delta Attention, Attention Residuals, and 896 Experts
Softmax costs O(n squared), so what you can afford depends only on how long the axis is. Kimi K3 drops softmax attention from three sequence layers in every four, and in the same model buys it on the depth axis. We derive the delta rule behind Kimi Delta Attention, work out why its decay carries a floor of exactly -5, unroll a residual connection to see what Attention Residuals replaces, and follow quantile balancing back to the dual it solves.