The most upvoted story on Hacker News on July 20 was Ben Thompson's "Who's Afraid of Chinese Models?" It landed at a moment when two big Chinese open-weight releases, Kimi K3 and Alibaba's Qwen3.8 Max, had spent the weekend moving markets and souring the mood at US frontier labs. The piece is long and worth reading in full. The argument I want to pull out of it is narrower and, I think, more useful: the panic about Chinese models being "free" misses what is actually happening to the price of intelligence.
Here is the claim in one sentence. Kimi K3 is not cheaper than Anthropic's Fable 5 because Moonshot has a cheaper cost structure. It is cheaper because the US frontier labs are supply constrained and charging well above what a cleared market would clear at. The "free" Chinese model is the first crack in a price umbrella that was always going to come down.
The "open weights are free" confusion
The first thing Thompson untangles is the word "free." People hear "open weights" and hear "no cost." That is wrong in the way that matters. Weights are a fixed cost. You pay for them once, when you train the model, and then you never pay for them again. What you keep paying for is inference. Running the model costs compute, and compute is a cost of goods sold, which scales with how much you serve.
This is the part of the AI business that breaks the old software intuition. Software had near-zero marginal cost for a long time. AI does not. If driving a dollar of revenue costs you fifty cents of tokens, then a hundred million dollars of revenue carries roughly fifty million dollars of inference cost. Scaling a token business scales COGS with it.
So "open weights" means you skip the research-and-development bill. It does not mean you skip the per-request bill. Kimi K3's published pricing is $3 per million input tokens and $15 per million output tokens. That is cheaper than Anthropic's Sol on the headline, $5 and $30, but Thompson's point is that the headline is the wrong measurement.
Tokens are not a commodity. Intelligence is.
The reason the headline price is the wrong measurement is that a token is not a fungible unit. Jensen Huang likes to call Nvidia GPUs "token factories," and that framing is internally consistent for Nvidia because their hardware is model agnostic. But a token from Kimi and a token from Sol are not the same thing. Reasoning models burn very different numbers of chain-of-thought tokens to arrive at the same answer. Kimi reportedly uses a lot more tokens than Sol for equivalent work, which closes most of the per-token price gap before you ever get to bill the customer.
What is actually fungible is the output. If two models both produce a correct answer to the same prompt, the answer is fungible. The thing being commoditized is intelligence, not tokens. Tokens are the input. The closer the question gets to "build me a CRUD app," the more the answer is a commodity already.
That is a textbook commodity market. Everyone charges the same price because everyone is selling the same thing. Price gets set by supply and demand. The supplier with the worst cost structure sells at their marginal cost, and the profit of everyone above them is the gap between their cost and that marginal cost. In Thompson's example, three suppliers at $10, $15, and $20 per unit clear the market at $20. The $10 supplier pockets $10 a unit. The $20 supplier pockets zero and is the one closest to going bankrupt if they are carrying debt.
Translated to models: the lab with the lowest cost per unit of produced intelligence wins. And the thing to notice is who that lab probably is.
The expensive labs are probably also the cheapest labs
This is the uncomfortable part for the "the Chinese will undercut the US labs" story. Anthropic and OpenAI likely have the lowest cost per unit of frontier-quality intelligence today. They have been serving models at this capability level for months before their competitors got there. They have had the entire run to optimize serving efficiency, batching, prefix caching, KV-cache memory, and all the engineering that makes a given token cheaper to produce. They also have the data flywheel. Every request they serve trains the next model.
If intelligence is becoming a commodity, the worst-case scenario for the frontier labs is not that Chinese models undercut them on price. It is that the frontier labs are themselves the lowest-cost producer, and if compute supply ever caught up with demand, they could charge commodity prices and still outearn everyone in the market. The reason they are not doing that today is that compute supply has not caught up with demand. They are selling into scarcity and pricing into scarcity.
The Chinese models look cheap because the US frontier labs are expensive. They are expensive because they can be. The umbrella goes down when supply clears, and Thompson argues that the agent unlock is large enough that frontier inference volume is going to grow much faster than training costs even if training costs keep climbing. Once there is enough compute to go around, the labs can lower prices and still make it up in volume.
So why are the frontier labs panicking?
Three reasons Thompson gives, and I think one of them is doing more work than the others.
One is anchoring. The labs modeled their financials around training dominating spend, which made inference revenue the thing that funds the next training run. That meant pricing inference aggressively high. If inference is going to grow much faster than training, that fear was wrong, and they can price lower without starving the next run.
Two is the data flywheel as a moat. Whoever runs the inference collects the data, and the data improves the next model. This pushes the frontier labs toward lower prices and more usage to keep the flywheel spinning, but only if they have the compute to serve that volume. It also pushes companies like Microsoft, who own the customer experience but not the frontier model, toward helping customers run their own models. That option is only realistic if Chinese open weights are a real alternative.
Three is the political and ideological angle at Anthropic specifically. Anthropic has built a posture around the claim that only it can be trusted with powerful AI. Open weights are a direct rebuttal to that claim. That is less an economic argument and more a worldview being challenged, which explains some of the decibel level.
The Hugging Face incident is what actually scared me
The economic argument is interesting but a little bloodless. The part of the weekend that actually made me sit up was the Hugging Face incident report. Hugging Face's production infrastructure was breached the week of July 13 by what they described as an "autonomous" agent system. The attackers got node-level access, harvested cloud and cluster credentials, and moved laterally across internal clusters over a weekend using short-lived sandboxes with self-migrating command-and-control on public services.
The defenders started their log analysis on the frontier commercial APIs. It did not work. To analyze 17,000 logs of attack commands, exploit payloads, and C2 artifacts, they had to submit those payloads to the model. The US providers' safety guardrails blocked them. The guardrails cannot distinguish an incident responder from an attacker, so they block both. Hugging Face's security team could not run their forensic analysis on the models they were paying for.
So they pulled the open-weight GLM 5.2 model from Z.ai, ran it on their own infrastructure, and did the analysis there. Two benefits, they wrote. One, no guardrails to argue with. Two, none of the attacker data or the credentials it referenced left their environment.
Read that twice. One of the largest model hubs in the world, breached by an AI-driven attack, could not use the leading US models to defend itself because those models would not let it. Their fallback was a Chinese open-weight model. The lesson they published was not "be careful with Chinese models." The lesson they published was "have a capable model you control, ready to go, before you are breached."
This is where the Trump administration's panicked response to Anthropic's Fable release starts to look expensive. The administration pushed Anthropic to add strong cybersecurity classifiers to Fable 5 and Mythos 5 when they were redeployed on June 30. Anthropic wrote that the classifiers were deliberately set to trigger on requests that are "likely benign," meaning a request has to "look very clearly safe" to avoid tripping the safety system. That is a fine setting for a chatbot. It is an unworkable setting for a security operations center trying to triage live attack traffic.
The result is that defenders in the US are functionally banned from using Fable or Sol for cybersecurity work, while attackers are using models that are widely available and have no such restrictions. The best defender option left is a model from a country that has spent years trying to weaken US cyber defenses. That is the actual policy failure. It is not abstract. It is the situation Hugging Face was in last week.
The distillation fight has a comparable hole in it
The other piece of the weekend's argument was distillation. The shorthand is that Chinese labs are catching up because they query US frontier models and distill the answers into their own models. That is partly true and partly not. Dean Meyer and Konstantine Buhler wrote a piece on X arguing that distillation does not explain China's whole lead, because Chinese labs also have strong researchers, real compute, and good pretraining, but that distillation compresses the costly final gap between a strong base and a near-frontier system. Their sharp point is that Western open-weight makers can't distill from US frontier labs, because the frontier labs' terms of service forbid it, but Chinese labs either ignore those terms or operate outside their reach. So Western open weights end up distilling the Chinese distillation, with a detour through a country that the US would rather not be depending on.
Thompson's proposal is contrarian and worth quoting directly: the US should pass a law that makes collecting data for training models fair use, and bars US companies from putting anti-distillation clauses in their terms of service. The logic is that distillation is just querying an API, which is impossible to meaningfully prevent, and that the frontier labs themselves built their models by scraping the open internet. If scraping the internet to train a model is fine, querying an API to train a model is hard to distinguish in principle. The output of that policy would be Western open-weight makers going to the source rather than through a Chinese intermediary.
I do not think that proposal passes Congress, and I think the frontier labs would fight it hard. But the underlying observation is correct. The current arrangement, where the frontier labs can forbid US companies from distilling their models and Chinese labs face no such constraint, does not protect US frontier labs. It just shifts the distillation pathway through China and weakens US open-weight makers in the process.
What I'd take from the weekend
If you are building on top of these models, three things changed in how I'd read the market.
The US frontier models are not priced at their cost of production. They are priced at scarcity. Plan for that price umbrella to come down as compute catches up, because the labs are actually well positioned for a low-margin high-volume world and the agent unlock gives them the volume. Long-term contracts at today's token prices are probably not the bet.
The open-weight story is not a Chinese story. It is a "having a model you can run yourself, on your own gear, for the use cases where a guardrail would lock you out" story. The Hugging Face incident is the cleanest example I have seen. If you run a security team, "have a capable model pinned on your own infrastructure before you need it" is now on the checklist. It is not optional.
The distillation policy is more interesting than it looks. If the US keeps letting frontier labs forbid distillation while Chinese labs face no such constraint, US open-weight makers keep operating at a structural disadvantage, and the workaround keeps running through Chinese models. The clean fix is to legalize distillation in the US and let frontier labs compete on being better rather than on being the only ones allowed to teach. I doubt that happens. I think the more likely outcome is that the defensible position keeps moving up the stack into the harness, the IDE, the agent runtime, and the customer relationship. Which is, inconveniently for the frontier labs, exactly where Microsoft and Cursor and Codex and Claude Code already live.
The weekend was loud because a 2.8-trillion-parameter open model from a Chinese startup jumped to the top of a blind leaderboard. The louder signal is that the era of intelligence priced as a scarcity good is winding down, and the policy environment around who is allowed to produce and consume that intelligence has not caught up to that fact. The Hugging Face report is what that lag costs you in practice. Someone got breached and their own vendors' models would not help them read their own logs.