Open-weight AI is splitting three ways. Chinese labs anchor it, and cloud inference is cheaper than ever, but running models on your own hardware is getting harder. Washington, meanwhile, now describes AI as a race for superintelligence. Together these shifts decide who can build, run and control capable models outside a few closed labs.
Key points
- The best open models trail the closed frontier by about four months, and all of the leaders are Chinese.
- Chinese open weights now carry most real-world open-model traffic and derivative work, in industry and in science.
- Token prices hit record lows, while an AI-driven memory shortage makes local inference hardware more expensive.
- Community opposition and grid stress are slowing the data center build-out that open-model hosts depend on.
- US "superintelligence" framing and easily stripped safeguards are turning open-release decisions into questions of national strategy.
What's changing
China anchors the open layer
The strongest open-weight models, Kimi K2.6, GLM-5 and DeepSeek-V3.2, are all Chinese. They trail GPT-5.5 Pro and Claude Opus 4.7 by about four months, a gap slightly wider than in 2023 to 20251. Usage has shifted further than capability. Chinese models account for about 61% of OpenRouter tokens, and about 70% of new derivative models are built on Qwen2. In scientific papers, open weights appear in 44% of single-model-family studies, and Chinese families make up about 60% of those choices3. Export controls have been loosened into a transactional regime of tariffs and caps4, and they are also speeding DeepSeek's move to Huawei Ascend5. The evidence is strong and the trend is accelerating. We think it is likely, about 75%, that the default open layer stays Chinese-anchored through 2027.
Cheap tokens, costlier local hardware
Inference is turning into a commodity. Silicon Data's token index fell to $0.97 per million tokens on 1 September, less than half its level in late May, driven by competition among open-weight hosts and gains in efficiency6. Small models improved too. Phi-4-mini, Gemma and Qwen3 in the 0.6 to 8B range now handle tool use on laptops and phones7. Memory pulls in the opposite direction. Hyperscale AI may take up to 70% of global memory output in 2026, and OEMs received only half to two-thirds of the memory they ordered8. So capable local models arrive just as the hardware to run them gets rationed, and that plausibly pushes many users back to cloud APIs.
The build-out meets local politics and the grid
Community opposition blocked or delayed about $200B of US data center projects in the first half of 2026. More than 840 groups are now organized, and moratoriums in 2026 already number four times the 2025 total9. The grid is under strain as well. NERC issued a rare Level 3 alert after data centers dropped more than 1,000 MW of load within seconds10. Siting friction is unlikely to stop hyperscalers, but it raises costs, and independent open-model hosts on thin commodity margins will feel that first.
Race framing meets unrecallable weights
At the UN on 22 September, President Trump said US documents would replace "artificial intelligence" with "super intelligence", and he rejected international oversight11. The industry response has been to present open weights as a competitive asset. More than 270 organizations signed Nvidia's letter opposing premature restrictions, and OpenAI, Nvidia and several startups now release open weights12. The security picture cuts the other way. GLM-5.2 refused none of the tested offensive-cyber or biology tasks, and fine-tuning strips safeguards from open models13. Race rhetoric can support openness as a tool of competition or restrictions in the name of dominance. We do not yet know which way it will break.
What to watch
- The open-closed gap: whether Epoch's measured lag widens past six months, or whether a US open model retakes the top open slot.
- Release restrictions: any federal move to condition open releases on capability thresholds, and whether the superintelligence relabeling produces concrete agency language.
- Local hardware prices: DRAM and VRAM prices through 2027, which will show whether on-device AI stays a mass option or becomes a premium one.
- China's compute: Ascend-trained frontier models from DeepSeek, which would signal that the Chinese open stack no longer depends on US compute.
Sources
- Best open-weight models trail the closed frontier by about four months, and the gap has widened slightly · grey-lit ↩
- Chinese open-weight models dominate real-world token use and derivative models · institutional ↩
- Open-weight use in science reaches 44%, driven mainly by Qwen · peer-reviewed ↩
- US loosens H200 exports to China but adds tariffs, caps and know-your-customer checks · institutional ↩
- DeepSeek moves frontier-adjacent training and infrastructure onto Huawei Ascend · grey-lit ↩
- Token prices fall to record lows as efficiency gains and price wars compound · disclosure ↩
- Small open models (3-10B) become capable of useful on-device agent work · expert ↩
- AI-driven memory shortage raises the cost of local-inference hardware · institutional ↩
- Community opposition blocked or delayed about $200B of US data-centre projects in H1 2026 · journalism ↩
- NERC issues a rare Level 3 alert after 1,000+ MW data center loads drop off the grid in seconds · journalism · also powermag.com ↩
- US 'superintelligence' relabelling at the UN recasts the AI race as a contest for dominance · journalism ↩
- US industry pushes back on restricting open weights as China's lead exposes a US policy gap · expert ↩
- Safeguards on open-weight models can be stripped, while near-frontier open models refuse almost nothing · journalism ↩