[Kimi-2] Kimi Linear: Rewriting Long-Context Attention with KDA
A systems-oriented derivation of Kimi Delta Attention, covering channel-wise forgetting, the specialized DPLR transition, chunkwise kernels, the 3:1 KDA/MLA hybrid, controlled evaluations, and the quality-throughput trade-off at 1M context.