HomeArticles

Still Reading AI Model Pricing as a Price War? The Rules Changed in Mid-2025

2026-08-25 · Original in Chinese

Across 211 list prices at nine model companies, AI pricing stopped falling in mid-2025 and started sorting into tiers instead of a price war.

At the end of our last post, the Q2 cloud earnings scorecard, we left a question hanging. If demand is this strong, why is the market at the same time arguing that model companies are cutting prices and worrying that nobody can make money selling tokens?

Key takeaways

  • Across 211 list prices at nine model companies, price cutting stopped in mid-2025: China ran 10 increases against 4 cuts, the US 13 against 14.
  • The increases cluster in workhorse models and the cuts in lightweight and last-generation ones. That is tiering, and it is what a maturing market looks like.
  • Price the job, not the million tokens. Sonnet 5 lists 60% below Opus 5 and still costs $4,010 to finish the same problem set.
  • DeepSeek wrote a peak-hour surcharge into its price list, a vendor saying out loud that it does not have enough compute.

The worry is not new. Since last year, the line that token prices only go down and the model companies are in a price war has come back over and over, and lately some people have started calling falling token prices the most dangerous signal in AI. Every version says the same thing: model companies are undercutting each other, turning tokens into a business that gets cheaper and cheaper and never earns anything. If that were true, then every dollar of capital spending is going into a market that keeps getting cheaper.

Here is the picture people point at: a token price index sliding since June while H100 rental prices bounce back to the highs. Two lines going opposite ways, and the cheap-token line is the one that scares everybody.

The token price index versus H100 rent chart from the original post is a third-party compilation and is not reproduced here.

The trouble is that the last few weeks of pricing news point in every direction at once:

  • July 31: OpenAI cut the price of its cheapest model, GPT-5.6 Luna, by 80%. The same day, on its flagship Sol, it opened a fast lane that charges 2x the price for 2.5x the speed.
  • August 6: DeepSeek announced a large increase. On August 16 it went through, with the steepest items up more than tenfold.
  • August 10: Anthropic canceled the Sonnet 5 increase it had scheduled for September 1, locking in the cut it had already made.

Cuts, freezes and hikes, all in the same two weeks. How is anyone supposed to read that?

Rather than guess from headlines, we did the arithmetic. We pulled every API price change at nine Chinese and US model companies from September 2023 to today: 211 list-price observations, restricted to models scoring 20 or better on the Artificial Analysis intelligence index. Here is what came out.

This industry is not cutting prices. It is sorting itself into tiers. The market is worried because it has not looked at the data.

Since mid-2025, model pricing has worked differently

Competition among model companies is brutal, and every one of them says it is playing for share. But whether they actually raised or cut is not something they can talk their way around.

We split every price change into two kinds: a straight repricing of the same model, and a new generation launching above or below the one it replaced. Add both together and you have what the vendors actually did.

All pricing events, same-model repricing plus generational repricing: China (top, five vendors) and the US (bottom, four vendors), before and after July 2025; dark = same-model, light = new generation; red = increase, blue = cut
Figure 1: All pricing events, same-model repricing plus generational repricing: China (top, five vendors) and the US (bottom, four vendors), before and after July 2025; dark = same-model, light = new generation; red = increase, blue = cut Original figure from the August 25 post (dashboard snapshot of that date); labels in Chinese.

Two things to take from that chart. The top half is the five Chinese vendors, the bottom half the four US ones, and the split left to right is July 2025. In the two years on the left, China and the US together made 17 price changes and only 4 were increases, which matches everyone's memory that prices only fell. On the right, the direction changes.

  • China flips from mostly cuts to mostly increases: 10 up, 4 down.
  • The US goes from almost nothing but cuts to increases and cuts side by side: 13 up, 14 down.

So the industry has gone from prices that only fall to prices that go both ways. But China and the US are raising for different reasons, and they are worth taking one at a time.

China: when compute gets tight, the price goes up. Launch a new version on promotion, let the promotion lapse and "return to the regular price," add a surcharge when load runs high. That is the standard playbook. DeepSeek has raised the price of the same model three times in eighteen months. The latest one, effective August 16, splits the list into peak and off-peak, where peak is Beijing office hours, and the steepest items went up more than tenfold. Writing "busy hours cost more" into the price list is a vendor admitting in public that it is short of compute. The others run the same barbell: flagships get more expensive every generation (Kimi up 499%, GLM up 124%, Qwen up 56%) while the lightweight models get cut or frozen. The cheap end buys traffic, the expensive end makes the money.

DeepSeek main models: blended price (USD per million tokens, log) with the pricing events marked, April 2024 to August 2026
Figure 2: DeepSeek main models: blended price (USD per million tokens, log) with the pricing events marked, April 2024 to August 2026 Original figure from the August 25 post (dashboard snapshot of that date); labels in Chinese.

The US: leave the model already in your hands alone, and put the increase on the next generation's workhorse. Almost all 13 US increases came from OpenAI and Google, and they were not repricings of existing models. There is exactly one of those on record. They were new workhorses launching above the ones they replaced.

On blended price, which weights input and output into one number at a fixed ratio so different models can be compared, OpenAI's main line climbed from $3.44 on GPT-5 to $11.25, and Google's Flash went from $0.175 to $3.0. The 14 cuts sit in the lightweight and last-generation models. Over the same stretch GPT-5.6 Luna went to $0.45. The tiering is happening inside a single company.

Anthropic looks like the exception. On paper it cut across the board, Opus from $30 to $10 and Sonnet from $6 to $4, then opened a higher band with Fable 5 at $20. But its models are chatty, and what a finished job actually costs went up, so we would not score that as a cut.

US vendors: launch price of each new generation in the same tier (USD per million tokens, log), July 2023 to August 2026
Figure 3: US vendors: launch price of each new generation in the same tier (USD per million tokens, log), July 2023 to August 2026 Original figure from the August 25 post (dashboard snapshot of that date); labels in Chinese.

Price the job, not the million tokens

Go back to that August 10 Anthropic freeze. Prices locked low, a scheduled increase pulled: generous, on the face of it. Run it again on cost per task and the story flips.

Artificial Analysis puts every model through the same set of problems. Along with an intelligence score, it measures what it actually costs to finish the whole set, including the tokens a model burns thinking to itself. On that ruler, Sonnet 5 lists at $4, 60% below Opus 5, and costs $4,010 to run the set, more than any other model. Opus 5 at its highest setting costs $2,909.

The reason is simple. It talks. On the same problem set Sonnet 5 spent 300 million tokens just answering, 32 times what the most economical model used. However cheap the sticker, multiply it by that kind of volume and you get the biggest bill on the board. Databricks' CTO put it plainly: cheaper per token does not imply cheaper per task.

How chatty each model is: tokens burned running the same benchmark (implied, millions); red = China, blue = US
Figure 4: How chatty each model is: tokens burned running the same benchmark (implied, millions); red = China, blue = US Original figure from the August 25 post (dashboard snapshot of that date); labels in Chinese.

The interesting part is that Chinese vendors have already priced this in. Kimi K2.6 lists at $1.7 and costs $841 to run the set. The next generation, K3, lists at $6 and costs $283. The sticker went up and the customer's bill went down. Count it by list price and that is an increase. Count it by cost per task and it is a cut.

Real work gives the same answer. Cursor is the most widely used AI coding tool, and its own benchmark, CursorBench 3.2, runs real programming tasks, which is the most token-hungry job there is. On the August 13 board the cheapest model costs $0.39 per task and the most expensive $17.3, a 44x spread, far wider than the spread in their list prices. The board only tests coding, and Cursor and xAI have a partnership, so read the direction, not the absolutes. It does not change the conclusion.

CursorBench 3.2: agentic coding score against cost per task, effort sweep per model, snapshot of August 13, 2026
Figure 5: CursorBench 3.2: agentic coding score against cost per task, effort sweep per model, snapshot of August 13, 2026 Original figure from the August 25 post (dashboard snapshot of that date); labels in Chinese.

So dollars per million tokens is getting less and less useful for telling you who is cheap, and it is no use at all for telling you whether this industry is in a price war.

Newest intelligence costs more, last year's costs less: two markets, not a contradiction

Stretch the time frame and two things that look contradictory are happening at once. If you want the last few points of intelligence, it keeps getting more expensive: every flagship generation launches above the one before, and Fable 5 opened a new top band at $20. If all you need is what counted as smart last year, the price is headed toward zero. Hold capability fixed at an intelligence score of 37 and the cost of finishing the same problem set fell from $779 to $25 in six months, down 97%. The cheap end is almost entirely Chinese: DeepSeek V4 Flash scores 51.8 at $0.175.

Intelligence index against blended price (log), snapshot of August 13, 2026: color = launch period, dashed = cumulative price-performance frontier
Figure 6: Intelligence index against blended price (log), snapshot of August 13, 2026: color = launch period, dashed = cumulative price-performance frontier Original figure from the August 25 post (dashboard snapshot of that date); labels in Chinese.

Two directions to read on that chart. The whole frontier keeps pushing up and to the left, which is the same intelligence getting cheaper. And the top of it keeps extending up and to the right, which is the best intelligence selling for more. Semiconductors ran this exact path for decades. The chip gets more expensive, the transistor gets cheaper, both at once, and nobody says the chip industry is in a price war.

The point is that these are two markets serving two different sets of customers.

  • Frontier models sell quality: agent runs that go all day, research-grade problems. The customer wants something nobody else can do and is not price sensitive.
  • Last year's intelligence sells volume: classification, summarization, simple repetitive processing. The customer wants good enough at the lowest price available.

White-collar work is tiered the same way. A senior researcher and a document clerk do different jobs for different money, and a falling clerical wage has never meant the researcher is about to be out of a job.

As for where the volume goes and where the money goes, the gateway data is clear enough. Vercel is a single gateway, so read the direction and not the absolute share.

China models combined share of the gateway: tokens versus dollars versus requests, seven-day average, June 13 to August 13, 2026
Figure 7: China models combined share of the gateway: tokens versus dollars versus requests, seven-day average, June 13 to August 13, 2026 Original figure from the August 25 post (dashboard snapshot of that date); labels in Chinese.

Chinese models are already more than half of tokens and only about 15% of spend. The volume market has gone to the cheap models; the money is still in the quality market. That also explains why the price index everyone is worried about is falling. It is a spend-weighted average price, so if volume shifts toward the cheap end, the average gets dragged down mechanically. Not one vendor has to cut a price for that to happen.

Once you split it into two markets, the worry about whether model companies make money comes down to three plain questions.

  • Is anyone buying share by giving price away? The data says no. Buying share that way looks like cutting the workhorse. Since mid-2025 the workhorses have almost all gone up, and what is still being marked down is lightweight and last-generation models, which were always the volume market.
  • Will today's high prices still hold tomorrow? That is the real pressure. The same capability got 97% cheaper in six months, so the premium you can charge for standing at the frontier has a half-life of a few months. Read the generational increases backwards: vendors have to collect while they still can. A model company's margin lives or dies on staying at the frontier.
  • Is the increase a win, or is it forced? DeepSeek raised prices while short of compute and still losing money, which looks cost-pushed. The peak/off-peak design gives away the purpose: using price to push load into off-peak hours is how you ration a resource you do not have enough of, not how you take share. That kind of increase leaves the margin where it was. It just moves the pressure onto the customer.

One clarification. Everything here is API pricing, what developers and enterprises pay. When OpenAI opened Luna to free users for unlimited chat on August 7, that is the consumer side buying a habit with free access and monetizing it somewhere else. Do not read giving it away as a price war.

Back to the main line: this has all been one argument

From July to now, four posts have been making the same case. The compute post said supply stays short at least through the second half of 2027 and today's price hikes are the early innings. The breakeven post drew the cost line and worked out when AI revenue has a chance to catch up with what data centers cost to build. The 2Q26 cloud earnings post ran the scorecard on the supply side: capacity-constrained and still accelerating. This one fills in the demand side. Model companies are charging by tier, and one of them has written "we do not have enough compute" straight into the price list. Put the four together and it is one picture.

In the last two weeks the two companies that sell compute reported in the same direction. CoreWeave is carrying $104B of signed business it has not worked through yet, $25B of it added since July alone. Nebius's latest capacity auction cleared 15% above its previous high, with short contracts quoted at twice the long-contract rate, and the CEO's line was that they choose when they sell, who they sell to, and on what terms. That is a seller talking his own book, of course. But quoting like that and still getting signatures is the most honest evidence of demand there is.

The next thing worth watching is the experiment DeepSeek has already set up for us. The increase went live August 16. Over the next few weeks we find out whether volume holds. If it holds, demand can carry a price increase, and tight compute ties all the way through from supply to demand. If it drops, we change our view accordingly.

One last thought, the thing that stuck with us most from writing this.

It is easy to take one piece of a market and make it the whole story. A spend-weighted price index rolls over and suddenly the whole industry is cutting prices; one vendor cuts and the price war has begun. But the AI market is growing up, and growing up means sorting into layers. The models are tiering, with the frontier selling quality and the cheap end taking volume. The pricing is tiering, with flagships opening high and old models heading toward zero. The customers are tiering too, free on the consumer side and paid API on the business side.

Sort into layers first, then draw conclusions. You end up much closer to what the AI market actually looks like.

That is also how we work this series. AI is a puzzle and pricing is one piece of it. Supply, cost, demand, each worked out on its own and then fitted back into the same picture, which is the only way an answer holds up. The pricing events, task costs and gateway shares in this post live on the FinSight data site.

In the next post we take on the next piece. The market says AI revenue growth is coming in below expectations. First we check that against the numbers, to see whether it really is below expectations or whether the ruler is drawn wrong. Then we lay out the history of AI's growth and look at what drove each step up. Get those two straight and you can answer what everyone actually wants to know: can AI ARR keep climbing from here.