Daily Pulse · · 14:00 CET · market · NVDA
A structure read for revelation day — companion to this morning's Morning 10, where the weights drop was point 7.
Moonshot AI published the full open-weights checkpoint for Kimi K3 today. The model itself has been live through the app and API since 16 July, so the news is not the capability — it is that anyone with enough silicon can now hold it. And with the weights came the first wave of independent assessment, which is more interesting than either the hype or the dismissal.
The scoreboard, honestly labelled
BenchLM ranks K3 fifth of 215 models at 79.98/100 — behind Claude Fable 5 and GPT-5.6 Sol, ahead of Claude Opus 4.8. That is the summary number, and on its own it says little. The discipline-level results say considerably more.
| Discipline | Kimi K3 | Best US comparison | Source |
|---|---|---|---|
| Arena Frontend Code | #1 | beat Fable 5 in 76% of duels | independent |
| SWE Marathon | 42.0 | GPT-5.6 Sol 39.0 · Fable 5 35.0 | independent |
| Toolathlon-Verified (Pass@1) | 76.5% | Opus 4.8 76.2% | independent |
| DeepSWE | 67.3 | GPT-5.6 Sol leads | Moonshot |
| BrowseComp | 91.2 | — | Moonshot-reported |
| DeepSearchQA | 95.0 | — | Moonshot-reported |
The labelling matters. The frontend, SWE Marathon and Toolathlon results are independently evaluated; BrowseComp and DeepSearchQA are the vendor's own numbers and should be read as such. Moonshot itself concedes the model still trails the best closed systems in overall user experience.
Where it is plainly behind — and the uncomfortable pairing
A joint UK AISI–US CAISI assessment put K3 well behind the leading American models on exploit development. In a simulated 32-stage corporate-network attack it reached stage 17 on average against 28.5 for the strongest US models, and achieved arbitrary code execution in none of 41 ExploitBench cases.
The uncomfortable part is the pairing. K3 is less capable at offensive cyber and less restrained about it — where US models refuse, K3 generally attempted the task. Weaker at the work, weaker at declining it. That combination is the part a risk committee will read twice, and it is an argument for the closed models that has nothing to do with benchmarks.
The catch that decides who captures this
K3 depends heavily on its preserved thinking history. If an agent harness fails to return that history correctly — or if you switch to K3 halfway through a session — performance destabilises. The model cannot be cleanly separated from its orchestration layer.
That sentence is the whole Closelook thesis restated by a model release. Friday made the same point in cash: Alphabet raised capex and was sold 7.1%, Intel printed its fastest growth since 2011 and closed down 7.9%, while SAP and ServiceNow — the two companies that called themselves the orchestrators — were paid on the same day. A frontier model you can download for nothing, which only performs inside a harness you have to build and run, does not commoditise intelligence. It moves the margin to whoever owns the harness.
Open weights do not mean cheap
The architecture is a Stable LatentMoE: 2.8 trillion total parameters, of which only 16 of 896 experts fire per token — roughly 50 billion active. A one-million-token context, native vision. The MXFP4 checkpoint is about 594 GB to download and roughly 1.4 TB resident, which no single GPU and no single eight-GPU node can hold; production serving is quoted at 64 or more accelerators.
Set that against Moonshot's own API at $0.30 per million cache-hit input tokens, $3.00 cache-miss, $15.00 output, with a reported cache-hit rate above 90% in coding workloads. For almost everyone, renting is cheaper than owning. Open weights here buy sovereignty and data control, not savings — the cost moves from a licence line to a capex line. That is inference economics doing exactly what the framework says it does.
What our own instrument already shows
None of this is a forecast for us — we measure it. The Token Split Pulse counts how open-market model demand actually routes, and in the latest daily reading, dated 25 July, Chinese models took 68.1% of tokens against 30.7% for American models and 0.8% for European ones. On the trailing week it is 66.6% against 32.5% — a China-to-US ratio of 2.06×.
Read that with care: this is the open router, which naturally favours open-weight models, so it measures where unlocked demand goes rather than total global inference. But that is precisely the pool K3 was released into, and it was already two-to-one before today.
The conclusion, stated plainly
China has not overtaken the United States across AI. The gap has become task-specific. For frontend development, long-horizon coding, research and tool use, a Chinese open-weight model is now trading blows with the best closed American systems. For general reliability, user experience and high-end cyber capability, the US leaders keep their advantage.
That is still a change of kind rather than degree. Frontier intelligence is no longer exclusively American, and increasingly no longer exclusively closed. And it does not prove that compute stopped mattering — it proves that architecture, training efficiency, open weights and orchestration decide how much strategic advantage each unit of compute buys. Which is a question about the multiple on the capex, not about the capex.
One more block on our desk — tomorrow's levels and what we're watching into the next session. Join the Look — it’s free.
SignalsThe live feed — what the book did and why, straight from the holdings diff.Open the feed →