Moonshot AI released the full open weights for Kimi K3 on July 26, 2026 at 7:30 PM EDT, a full day ahead of its announced July 27 target...
What Moonshot Actually Shipped
Kimi K3 is built on a mixture of experts architecture containing 896 experts in total, but only 16 fire on any given token, bringing the effective active parameter count down to roughly the tens of billions rather than the full 2.8 trillion. The model supports a context window of 1,048,576 tokens, native vision input, and an optional reasoning effort mode that lets developers dial computation up or down depending on task difficulty. Moonshot's API has been live since July 16, 2026, when the model first debuted at the WAIC conference in Shanghai, and it runs on an OpenAI SDK compatible endpoint priced at 3 dollars per million input tokens and 15 dollars per million output tokens.
Interest has clearly outpaced Moonshot's own infrastructure. The company paused new subscriptions on July 19 after its announcement pulled in 7.2 million views, an unusually large reaction even by the standards of a crowded AI news cycle. On quality benchmarks, Kimi K3 does not top every index, ranking third overall on some measures despite carrying the largest parameter count in the field, though it does outperform Claude Opus 4.8 on several individual benchmarks.
The 1.4 Terabyte Problem
Here is the catch behind the headline. Even in the compressed 4-bit MXFP4 format, the Kimi K3 weights take up about 1.4 terabytes of storage. At full 16-bit precision, that figure balloons to 5.6 terabytes. No laptop, workstation, or even a typical small business server rack comes close to handling that footprint, let alone the memory bandwidth needed to serve it at usable speed. Running Kimi K3 requires data center grade GPU clusters built around Blackwell or MI400 class silicon, hardware that only large cloud operators currently deploy at scale.
Together AI and Modal moved quickly to offer day-zero hosting access, giving developers a way to use Kimi K3 without owning the underlying infrastructure themselves. Microsoft is reportedly evaluating the model for Copilot workloads and preparing Azure availability, though that detail remains unconfirmed and was reported by The Information rather than by Microsoft directly. As one industry outlet put it, Kimi K3 hands you the uncomfortable part of the AI race in a single download, meaning the weights are free, but the ability to actually run them at scale is not.
A Pattern, Not a One-Off Release
Kimi K3 is not Moonshot's first open-weight bet. The strategy began with Kimi K2 in July 2025 and continued with K2.5 in January 2026, meaning K3 is the third release in a deliberate, escalating open-weight campaign rather than a surprise decision. Founder Yang Zhilin, a Tsinghua University graduate, has backing from both Alibaba and Tencent, giving Moonshot the capital runway to keep releasing frontier-scale models for free even as licensing terms for K3 itself remained unconfirmed at launch.
Industry analysts see this as part of a broader trend among Chinese AI labs, which are open-sourcing aggressively partly as a response to ongoing US chip restrictions. AI researcher Nathan Lambert has estimated that the performance gap between Chinese and US frontier models has narrowed from six to nine months down to roughly three to five months, a pace that has clearly gotten Washington's attention.
The Distillation Accusation Hanging Over the Release
Whatever the outcome of that dispute, the underlying trend is unlikely to reverse. Open weights lower the barrier to entry for developers everywhere, but a 1.4 terabyte model is a reminder that "open" and "accessible" are not the same thing. For most companies, the practical path to using Kimi K3 will run through a hosted API rather than a self-managed deployment, at least until infrastructure costs come down.
Frequently Asked Questions
What is Kimi K3 and who released it?
Kimi K3 is a 2.8 trillion parameter open-weight AI model released by Moonshot AI, a Beijing based lab founded by Tsinghua graduate Yang Zhilin and backed by Alibaba and Tencent. It uses a mixture of experts architecture with 896 experts, of which only 16 activate per token, and it supports a 1 million token context window with native vision and reasoning capabilities.
Can I run Kimi K3 on my own computer or a single server?
Realistically, no. The full weights require about 1.4 terabytes of storage in 4 bit MXFP4 precision and 5.6 terabytes at full 16 bit precision. Running the model requires data center grade GPU clusters with Blackwell or MI400 class silicon, which is why hosting is currently limited to cloud providers such as Together AI and Modal rather than individual developers or small businesses.
How much does the Kimi K3 API cost compared to other frontier models?
Moonshot prices the Kimi K3 API at 3 dollars per million input tokens and 15 dollars per million output tokens through an OpenAI SDK compatible endpoint. The API has been live since July 16, 2026 following its debut at the WAIC conference in Shanghai, undercutting several Western frontier model providers on price even though K3 ranks third rather than first on major quality indexes.
Why is Kimi K3 controversial in the United States?
On July 22, 2026, the White House Office of Science and Technology Policy Director accused Moonshot of running a distillation operation against Anthropic's Claude Fable 5 using restricted NVIDIA chips routed through Thailand. Moonshot has not responded to the allegation. The accusation adds to broader US concern about Chinese labs closing the performance gap with American frontier models while operating under chip export restrictions.
Does Kimi K3 actually outperform Western models like Claude Opus 4.8?
It depends on the benchmark. Kimi K3 ranks third overall on several major quality indexes despite having the largest parameter count of any released model, but it does beat Claude Opus 4.8 on a number of individual benchmarks. Analysts describe the Chinese to US frontier gap as having narrowed to roughly three to five months, down from six to nine months a year earlier, which is the more important trend than any single leaderboard placement.
Keeping up with which AI models are actually production-ready versus which ones just make headlines is a full-time job. If your team needs help evaluating frontier models, building on top of them, or figuring out what any of this means for your product roadmap, ATXSOFT can help you cut through the noise and make a decision that fits your stack and your budget.

