China vs US AI Models: September 2026’s Line-Up, Prices and Token Resets

China vs US AI Models: September 2026's Line-Up, Prices and Token Resets

Twenty-seven AI models shipped in the first eighteen days of September 2026, from thirteen labs. The China vs US AI models race is no longer a narrative — it is a price list. In one month the spread between the cheapest frontier-grade output token and the most expensive hit roughly 42 times: DeepSeek V4.1 Flash at US$1.20 per million output tokens, GPT-6 Astra and Claude Mythos 5.1 at US$50. If you are choosing a model — or a subscription — for your team this quarter, the deciding facts are all on the table now. Here is the line-up, the prices, and the thing almost nobody explains properly: what a “token reset” actually means for the ChatGPT and Anthropic plans you may already be paying for.

The China vs US AI models line-up, September 2026

ModelLabReleasedPrice /1M tokens (input / output)
Claude Fable 5.1 · Mythos 5.1Anthropic (US)1 September$10 / $50 (Mythos); Fable 5.1 cache reads cut $1.00 → $0.25
Gemini 3.8 Flash (+ Cyber variant)Google (US)2 September$0.75 / $3.75 · 1M-token context
Muse Spark 1.3Meta (US)2 September$1.25 / $4.25 · open-weights roadmap
GPT-6 AstraOpenAI (US)3 September$10 / $50 — 2.5× the GPT-5.6 Sol rate it replaces
Qwen3.8-Max-0902Alibaba (China)2 September$2 / $6 · plus open-weight Qwen3.8 2.4T-A95B
DeepSeek V4.1 FlashDeepSeek (China)10 September$0.30 / $1.20 peak — off-peak billed at 50%
Kimi K2.8 PreviewMoonshot AI (China)11 SeptemberNo published rate
Grok 4.7xAI (US)21 September$6 output

Three reads off that table. The US frontier got more expensive — GPT-6 Astra’s launch rate is 2.5 times the model it replaced, and it tops independent coding benchmarks with that price tag. The Chinese frontier got cheaper again — DeepSeek V4.1 Flash’s peak output rate is 1/42nd of Astra’s, and its off-peak pricing halves it further. And the open-weights camp crossed a line worth marking: per independent tracker reporting, a Kimi open-weights model reached the highest position ever recorded for a non-proprietary model on the Artificial Analysis Intelligence Index, and Abu Dhabi’s IFM shipped an Apache-licensed six-model fleet from 0.9B to 375B parameters the same week OpenAI launched its flagship.

Benchmark honesty. Some early launch-table numbers circulating in September — notably portions of GPT-6 Astra’s benchmark sheet — remain unverified by third parties, and one tracked figure (Astra at 62.7% on ARC-AGI-3) comes from the vendor’s own harness. Where trackers and vendors disagree, we have said which is which. There is also one unreleased model both camps are now watching: Google has announced Gemini 3.5 Pro without a date.

What “token reset” actually means for ChatGPT and Anthropic plan users

If your team is on ChatGPT or Claude subscriptions, you have probably hit a wall with a countdown attached and wondered what exactly refills when. The mechanics differ per vendor, and the differences matter.

Claude: a five-hour clock plus a weekly cap you do not control

Anthropic meters usage inside a rolling five-hour session that starts with your first message. Pro buys “at least five times” the free allowance per session; Max 5x and Max 20x multiply that again. On top sits a weekly cap with a fixed reset time assigned to your account — not the day you subscribed, not your billing date. You can see the next reset in Settings → Usage, but you cannot move it. Anthropic’s own help text is explicit that the reset day stays the same regardless of when you start using Claude or when your subscription began.

Three September changes are worth knowing. Claude Code shares the chat allowance and had its limits doubled in May 2026, with a permanent 25% weekly increase added on 14 September. The new Fable 5.1 is only on Max plans, and it can consume up to half your weekly allowance on its own. And if you hit the wall, the documented escape is usage credits — prepaid, billed at standard API rates, covering chat and Claude Code, with the explicit caveat that they do not change reset timing. In other words: Anthropic will sell you overflow at API rates, which quietly makes the subscription a floor, not a ceiling.

ChatGPT: unlimited text, metered everything else

Since 6 August 2026, plain text chat is unlimited on every ChatGPT plan, including Plus. What still meters: reasoning effort on the GPT-5.6 family (fall back to lighter thinking when exhausted), file uploads (80 per rolling three hours), image generation, voice, and a monthly Deep Research allowance. Codex, the coding agent, runs its own five-hour windows per model — published ranges on Plus run from roughly 10 to 100 messages on the largest model to 250 to 2,000 on the fastest — with weekly caps on top of that. One more sign of demand outstripping supply: OpenAI paused new sign-ups to the top Pro 20x tier in September.

Gemini: compute budgets, not message counts

Google switched to compute-based limits refreshing every five hours inside a weekly ceiling, publishing only multipliers — AI Plus at 2x standard, AI Pro at 4x, AI Ultra at 5x–20x. The concrete published numbers are context windows: 32K free, 128K on Plus, 1M on Pro and Ultra.

Plan (Sep 2026)PriceWhat refills, whenWhat to expect when you hit it
Claude Pro$20/mo ($17 annual)~45-message sessions every 5h + account-assigned weekly resetLocked out until the window rolls; credits can bridge at API rates
Claude Max 5x / 20x$100 / $2005x / 20x Pro, larger weekly capSame reset clocks, bigger buckets; Fable 5.1 included
ChatGPT Plus$20/moUnlimited text; metered thinking effort, uploads, Codex windowsFalls back to lighter reasoning or smaller models
ChatGPT Pro 20x$200/mo20x Plus, shared across surfacesWeekly cap after heavy days; new sign-ups paused in September
Gemini AI Plus / Pro / Ultra$7.99 / $19.99 / $99.99+Compute budget every 5h inside a weekly ceilingFalls back to Flash; heavier features drain faster

The one-sentence version: a reset is a refill schedule, not a refund, and it is almost never your billing date. Plan around the clock you cannot move, or pay credits at API rates — at which point ask yourself why the subscription is there at all.

Token plans vs subscriptions: the comparison that actually matters for teams

That last sentence is the fork in the road. Consumer subscriptions meter in opaque compute units with vendor-controlled reset clocks. Token plans — API access, prepaid, priced per million tokens — meter in a published unit with no reset clock at all. For an individual, a $20 subscription is fine. For a team of five running AI inside real work, the subscription math gets strange fast: per-seat pricing, shared pools that drain unpredictably, and an overflow path priced at the very API rates you skipped.

Token plans flip the trade: you pay exactly what you use, at a rate card you can audit, billable to the company, with costs attributable per person and per project. The trade-off is that you carry the routing and the overuse risk — which is precisely the gap gateway products (Tencent Cloud’s CMR among them, covered in our previous piece) have moved to close.

And then there is the open-weights route, which China’s labs keep feeding. DeepSeek, Qwen and Kimi’s open models carry no rate card and no reset because you run them yourself — the token price becomes your own GPU cost. For Malaysian enterprises with data-residency sensitivities and steady, predictable workloads, an open-weight model on your own infrastructure is the only option with literally no meter running. For everyone else, the honest comparison is total cost per completed task — and on that measure, September’s data says the cheap model frequently wins, because a model that reasons four times longer can erase a four-times-cheaper rate card.

Dr Henry Tye’s take

Dr Henry Tye, founder of BigDomain and operator of its commercial LLM gateway, has watched the September wave from the routing layer, where every model’s price card lands first. His view, as he puts it: “The China camp is shipping ninety percent of the capability at five percent of the rate card, and the US camp is shipping capability priced for enterprises that have already committed. Malaysian businesses should stop asking which camp wins and start asking what each task is worth. Route the routine work to the efficient models, cache everything cacheable, and put the premium models only where the premium outcome is real. The teams that get the routing layer right this quarter will spend a fraction of what their competitors spend next year.”

His own operation runs day-to-day work on DeepSeek’s efficient tier, with premium models reserved for the tasks that justify them — the exact pattern the September price spread rewards.

What to expect next

  • Gemini 3.5 Pro is the one announced-but-unreleased model the trackers are watching; its arrival will pressure both camps’ mid-tier pricing.
  • Expect the price floor to keep falling. DeepSeek’s off-peak discounting and GLM’s sub-ten-cent input rates are price-war behaviour, and September’s 42× spread is not a stable equilibrium.
  • Expect subscriptions to keep adding caps in disguise. Unlimited text with metered reasoning is the pattern; watch for the same structure appearing across the industry.
  • Expect gateways to become default infrastructure. More models, more rate cards, more reset clocks — someone in your architecture has to hold the routing table.

About BD Media

BD Media is the editorial desk of Big Domain. We cover the dates, rules and market shifts Malaysian businesses have to plan around, name our sources so you can check them, and mark what we could not verify instead of filling the gap.

Sources

  • Capital & Compute, AI model releases by month with launch prices — capitalandcompute.net and its September 2026 roundup. The 27-releases/13-labs count; per-model launch dates and rate cards from lab announcements (Claude Fable 5.1/Mythos 5.1 $10/$50; Gemini 3.8 Flash $0.75/$3.75; Muse Spark 1.3 $1.25/$4.25; GPT-6 Astra $10/$50 at 2.5× GPT-5.6 Sol; DeepSeek V4.1 Flash $0.30/$1.20 peak; Qwen3.8-Max-0902 $2/$6; Kimi K2.8 Preview no rate; Grok 4.7 $6); the 42× spread; Fable 5.1’s cache-read cut; the IFM K2 Horizon Apache-2.0 fleet.
  • Flowtivity, US vs China AI models compared — flowtivity.ai (updated 4 September 2026). Astra’s ARC-AGI-3 figure and its unverified-portion caveat; Kimi’s open-weights Artificial Analysis ranking; DeepSeek V4 Pro’s SWE-bench tie at ~1/34th input price; GLM 5.3 Flash at $0.071/M input; Astra’s 1.05M context; Gemini 3.8 Flash’s 1M context and March 2026 cutoff.
  • Jobbit, AI coding usage limits in 2026 — jobbit.uk. Claude’s 5-hour window and account-assigned weekly cap; Claude Code doubling in May 2026 and the permanent +25% weekly change on 14 September 2026; Codex per-model 5-hour ranges on Plus (Sol 10–100, Terra 25–200, Luna 250–2,000); ChatGPT Pro 20x sign-up pause since 10 September 2026.
  • UseRightAI, AI plan usage limits — userightai.com (verified 2 September 2026). ChatGPT unlimited text since 6 August 2026 with metered thinking effort and uploads; Gemini’s compute-based 5-hour refresh with published multipliers and context windows; Fable 5.1 on Max only, up to 50% of weekly limits.
  • AI Models Compared, Claude usage-limit reset mechanics — aimodelscompared.com. Anthropic’s verbatim reset documentation: sessions reset every five hours, the weekly reset at a fixed account-assigned time, and usage credits prepaid at standard API rates that do not change reset timing.
  • BD Media, “Tencent CMR: One Gateway for Every AI Model Your Enterprise Uses” — media.bigdomain.my. The gateway layer this comparison points to.

Not verified: any benchmark figure still resting solely on a vendor’s own harness (notably parts of the GPT-6 Astra launch table); Moonshot’s pricing for Kimi K2.8 Preview (no published rate at research time); and ringgit conversions, deliberately omitted — rate cards move weekly and a fixed conversion would mislead more than it informs.