Beijing-based Moonshot AI has released Kimi K3, a 2.8 trillion parameter model that the company describes in its technical blog as the world’s first open 3T-class system and the largest open-weight AI model to date. The model features a 1 million token context window, native vision, and activates just 16 of its 896 experts per token, representing roughly 1.8% of the pool. Full weights are scheduled for release on July 27.
Kimi K3 Performance and Benchmarks
Moonshot stated that while K3 still sits behind Anthropic’s Claude Fable 5 and OpenAI’s GPT 5.6 Sol on overall performance, it outperformed every other model in the company’s evaluation suite, including Claude Opus 4.8 and GPT 5.5, across coding and agentic benchmarks. Notably, Arena ranked K3 first in its Frontend Code evaluation with 1,679 points, surpassing Claude Fable 5 in blind developer testing. This represents a 17-place jump from Kimi-k2.6, which previously ranked 18th.
In the Frontend Code Arena, Kimi-K3 ranked first in six of seven domains: Brand & Marketing, Reference-Based Design, and Data & Analytics. Bank of America analysts led by Alex Liu noted that K3 demonstrates that large-scale pre-training combined with architectural work can still deliver step-change gains for flagship Chinese models, despite ongoing compute constraints.
Architectural Innovations
Moonshot claims a 2.5x improvement in scaling efficiency over Kimi K2, attributed to two specific architectural changes: Kimi Delta Attention, a hybrid linear attention scheme, and Attention Residuals, which alter how information moves between layers. The company has implemented quantization-aware training starting at the supervised fine-tuning stage, utilizing MXFP4 weights and MXFP8 activations—a combination selected for broad hardware compatibility.
Hardware and Optimization
The company’s kernel optimization benchmark was conducted on Nvidia’s H200 and a GPGPU from an alternative vendor, which remained unnamed. Additionally, MiniTriton, a Triton-like compiler built by the K3 team, was benchmarked against Triton on an Nvidia L20, the cut-down Ada-based card sold into China under U.S. export rules. Moonshot recommends serving K3 on supernodes of 64 or more accelerators to keep expert-parallel traffic inside one high-bandwidth domain.
In a case study, K3 spent a 48-hour autonomous run designing a simulated inference chip for a nano model. Using open-source EDA tools and the Nangate 45nm library, the design closed timing at 100 MHz within 4mm squared, contained 1.46 million standard cells and an INT4 MAC array, and sustained over 8,700 tokens per second of simulated decode.
Pricing and Market Context
API pricing is set at $0.30 per million cache-hit input tokens, $3 per million on cache misses, and $15 per million output tokens. As Kimi K2 launched a year ago at $0.60 per million input tokens, uncached K3 input costs five times as much. Currently, every published K3 figure is a claim made by Moonshot or drawn from API access; these results cannot be fully verified until the weights are made public on July 27. Furthermore, Anthropic accused Moonshot in February of using 3.4 million Claude exchanges to train its models through distillation, and K3 currently benchmarks within a few points of the models named in that complaint.
| Category | Price (per 1M tokens) |
|---|---|
| Cache-hit Input | $0.30 |
| Cache-miss Input | $3 |
| Output | $15 |
Sources: Bloomberg, Tomshardware.