The Stack Splits
Over two weeks in June the United States restricted its two most capable public models and gated the next. China released its best model in the same weeks, as free, open weights that run on Chinese chips. The contest for the global AI standard is moving from capability to diffusion, where US policy is now the obstacle.
Read the companion deck (PDF)→
The two US restrictions and the GLM-5.2 release
Between June 12 and June 25 the US government took its two most capable public models offline and gated the next one. On June 12 an export-control directive ordered Anthropic to cut off Fable 5 and Mythos 5 for every foreign national, its own staff included, inside the country and outside it. No provider can screen foreign from domestic users in real time across hundreds of millions of accounts, so compliance meant a full withdrawal, and both models went dark for everyone. The stated trigger was a reported jailbreak of Fable 5, and Anthropic answered that one vulnerability cannot justify recalling a model already in hundreds of millions of hands, and that the same test applied across the industry would end frontier deployment altogether. The particular dispute matters less than the precedent it set, that the most capable US models now ship at the government's discretion. Thirteen days later that discretion became a procedure. On June 25 the White House asked OpenAI to hold its GPT-5.6 line, the Sol, Terra, and Luna models, for up to thirty days of review before release, and OpenAI complied while warning against letting the review become the default.
On June 13, within the same two weeks, Zhipu, the Chinese laboratory that sells abroad as z.ai, published GLM-5.2 as open weights under an MIT license, free to download and modify, and made the timing its argument, that when frontier access is cut for reasons unrelated to the model, the case for open weights makes itself. Five days later Elon Musk put a date on the convergence, estimating that Chinese models would reach the level of Anthropic's Fable by the first quarter of 2027, the horizon z.ai already gives itself, so the figure reads as a shared expectation. The people building these systems on both sides now plan for an open Chinese model at the US frontier within roughly two quarters.
The capability gap is five points
On the Artificial Analysis Intelligence Index the June table read Fable 5 at 60, Gemini 3.1 Pro at 57, Opus 4.8 at 56, GPT-5.5 at 55, and GLM-5.2, open and free, at 51. Fable 5 sat at the top of that column, and the June 12 directive pulled it, which left Opus at 56 as the best model most US developers could count on in production and set the gap to the best open model at five points. Read as a ranking, the table still favors the United States. But a ranking assumes every model is available to use, and the June 12 directive made the leader unavailable, so for anyone shipping a product the operative score was the 56 they could deploy, five points above the open alternative. A five-point edge among the models a developer can deploy is thin enough for the ordinary compounding of an installed base to erode, and that compounding is already under way. A team that fine-tunes GLM and builds its evaluation harness and deployment tooling around the model raises its own cost of leaving with every week of use, and multiplied across thousands of teams, that switching cost is how a widely held model becomes the default the next cohort inherits.
The strongest case against reading too much into the June events starts with revenue. The closed US frontier still earns almost all of it, a benchmark rank is not revenue, several of GLM-5.2's published figures are vendor-reported and unverified, and the restrictions could prove temporary. We grant all of it. The live question is the margin, the next model a builder commits to, and today's revenue reflects the diffusion of prior years, when the US models were the ones present and reliable.
The installed base, by the numbers
Hugging Face's Spring 2026 review puts Chinese models at 41% of all downloads, the plurality, with Qwen displacing Llama as the most-downloaded family. Our June 29 pull of the Hugging Face API shows the four largest Chinese open families, Qwen, DeepSeek, GLM, and MiniMax, drawing about 229 million downloads over the trailing month, more than double the roughly 105 million taken by the four largest US families, Gemma, Llama, GPT-OSS, and Phi, with Qwen alone near 180 million, ahead of every US open family put together. Raw download counts flatter whatever automated systems pull most, because a script can fetch a model a million times without a human ever loading it. A derivative resists that inflation, because it is a model some team chose to fine-tune and now maintains. Counted that way, Qwen anchors more than 113,000 fine-tuned and quantized descendants, over 200,000 once every model that tags it is included, a base Hugging Face puts above Google's and Meta's families combined. The same review recorded the labs' share of new model creation falling from roughly 70% before 2022 to about 37%, as independent developers rose to 39%. On the paid side, OpenRouter shows open weights at about a third of all tokens by late 2025, with Chinese models growing fastest.
The stack runs without US hardware
Open weights would matter less if the models still needed US silicon to exist, and that dependency is fraying. Zhipu reports that the GLM-5 family, GLM-5.2 included, trained on roughly 100,000 Huawei Ascend 910B accelerators on Huawei's MindSpore framework, with no Nvidia hardware at any stage. The figure is vendor-reported, and we treat it as such, though it is consistent across several independent accounts and with the scale of Huawei's CloudMatrix build-out. The domestic stack carries a measured cost. Inference on the Huawei hardware is reported at about 17 to 19 tokens per second against roughly 25 to 30 for Nvidia-backed peers, close to a third slower, a penalty we quote as a range because the source is single. Put the three layers together, open weights anyone can hold, training on Chinese accelerators, and distribution through a platform where Chinese models already lead, and the result is a working AI stack that sits outside US jurisdiction at every level. Hugging Face calls the pattern AI sovereignty, a government or firm taking an open model, fine-tuning it on its own data, and running it on hardware it controls. For a buyer in Riyadh, Jakarta, or Sao Paulo, the Chinese option carries no risk that a foreign government revokes it mid-deployment, and the one-third throughput penalty is a cost many will pay to remove that risk.
Next week the series turns to the capital cycle behind all of this, and to how much of it is memory. Gavin Baker puts high-bandwidth memory at 30 to 40% of hyperscaler capital spending by 2027, supplied by a three-firm oligopoly, two of them Korean. We will separate the capital-expenditure figure the market quotes from the capacity it buys, and trace why the AI trade increasingly reprices through Seoul before it reaches the megacaps.
Anthropic's June 12 statement on the Fable 5 and Mythos 5 suspension, with CNBC, Bloomberg, and Al Jazeera reporting; TechCrunch, CNN, and Axios on the June 25 GPT-5.6 hold; z.ai's GLM-5.2 release and model card, June 13; the Artificial Analysis Intelligence Index as of June 27; Hugging Face's State of Open Source, Spring 2026; the Hugging Face Hub API retrieved June 29 for the Second Order download and derivative counts; OpenRouter's 2025 token study; and GLM-5 training-hardware reporting from several outlets.
GLM-5.2's vendor benchmarks are not all independently verified, the GLM-5 training-chip claim is vendor-originated, and the Huawei throughput figures are single-sourced and shown as a range. Download counts are inflated by automated pulls and measure scale. Second Order is an independent research briefing, for discussion only and not investment advice.