US and China AI competition represented by a futuristic AI chip, data centers, and competing AI models with contrasting blue and red technology networks.

US and China AI Competition in 2026: Chips, Models and What It Costs You

The US and China AI competition is now close enough that price, not capability, decides which model most teams pick. American labs still lead on the hardest reasoning work, but Chinese models sell for a fraction of the price and ship their weights publicly. This guide is for creators, marketers and small business owners choosing where to spend an AI budget. Skip it if you want a policy essay — this one sticks to verified prices, dated rule changes and what to do next. Every number below comes from a vendor page or a government filing, and where no official figure exists, AI Era says so.

Key takeaways

  • Kimi K3, released 16 July 2026 by Moonshot AI, costs $3 per million input tokens and $15 per million output tokens — the same output rate as Claude Sonnet 4.6 and half the price of GPT-5.6 Sol.
  • The US approved H200 chip sales to China in January 2026, then China refused to buy them. Brookings reported no H200 units sold to Chinese firms as of mid-May 2026.
  • DeepSeek V4 Flash costs $0.44 per million input tokens at peak and $0.22 off-peak, roughly a tenth of GPT-5.6 Terra’s $2.00.
  • Open weights are now the main Chinese advantage. Moonshot posted K3’s weights publicly on 17 July 2026; no US frontier lab did the same in 2026.
  • A joint US–UK assessment placed Kimi K3 about six months behind US models on cyber capability — narrow, but not parity.

What the US and China AI competition actually is

Two things are being contested in the US and China AI competition, and most articles blur them. The first is compute: who can make and buy the chips that train large models. The second is distribution: whose models developers actually run.

The US controls the first. Nvidia, AMD and TSMC sit inside the American export-control net, and the Bureau of Industry and Security (BIS) decides who gets served. China is closing the second. Its labs price aggressively and release weights that anyone can download and host.

United StatesChina
Leading labsOpenAI, Google DeepMind, AnthropicDeepSeek, Alibaba (Qwen), Moonshot AI
Example frontier modelGPT-5.6 Sol, Claude Opus 5Kimi K3, DeepSeek V4 Pro
Weights released publiclyNo (frontier tier)Yes (Kimi K3, 17 July 2026)
Top-model output price$25–$30 / 1M tokens$3.96–$15 / 1M tokens
Chip supplyDomestic and alliedRestricted; ~41% domestic in 2025 (Brookings)
Grid additions, 2025540+ GW added (Brookings)

Compute here means the processing power used to train and run a model. Weights are the trained numbers inside a model; releasing them lets anyone run it on their own hardware.

Pricing: what each side charges, and the costs nobody lists

These are current published API rates per million tokens.

ModelVendorInputOutput
GPT-5.6 SolOpenAI (US)$5.00$30.00
Claude Opus 5Anthropic (US)$5.00$25.00
Claude Sonnet 5Anthropic (US)$2.00$10.00
Gemini 3.7 FlashGoogle (US)$0.75$3.75
Kimi K3Moonshot AI (China)$3.00$15.00
DeepSeek V4 ProDeepSeek (China)$1.32$3.96
DeepSeek V4 FlashDeepSeek (China)$0.44$1.32

DeepSeek figures are peak rates; off-peak is exactly half. Kimi K3’s cache-hit input drops to $0.30.

Four cost factors that rarely appear in comparison posts:

  1. Token inflation. Anthropic’s docs state Claude 4.7 and later use a newer tokenizer that produces about 30% more tokens than earlier models. Same text, bigger bill.
  2. Introductory rates that expire. Google lists Gemini 3.7 Flash at $0.75 input “through December 31, 2026,” rising to $1.50 on 1 January 2027. Budgets built on today’s rate will double.
  3. Cache and batch multipliers. Anthropic charges 1.25x for five-minute cache writes and 2x for one-hour writes, then 0.1x on hits. Batch processing cuts 50% at OpenAI and Anthropic.
  4. Clock-based billing. DeepSeek’s price changes by time of day. A scheduled job at the wrong hour costs double.

One honest gap: Alibaba’s public Model Studio documentation lists Qwen models such as qwen3.7-max and qwen3.6-flash but does not publish a plain USD per-million-token table. AI Era will not quote a Qwen price until Alibaba publishes one.

Where each side leads in the US and China AI competition

Chips and energy. The US holds the advanced-node supply chain. China holds power. Brookings reports China added over 540 GW of generating capacity in 2025 and that Huawei’s Ascend 950PR line is expected to scale to 750,000 units in 2026. Cheap electricity is why Chinese inference prices are sustainable, a point we expand on in our AI hardware trends map.

Model capability. US frontier models still lead on long-horizon reasoning and agentic tasks. The gap is months, not years — the joint US–UK read on Kimi K3 put it roughly six months behind on cyber capability.

Openness and price. This is China’s clearest lead. Downloadable weights plus low API rates make Chinese models the default for cost-sensitive builders, which is a distribution advantage that export controls do not touch.

Capital. US hyperscalers plan around $650 billion of spending in 2026, per Brookings. That buys training runs Chinese labs cannot yet match, though it also raises questions we cover in our look at whether the AI bubble bursts in 2026.

Comparing the models directly

ModelOpen weightsBest forWeak spot
GPT-5.6 SolNoComplex agent workflowsHighest output price here
Claude Opus 5NoLong documents, careful writingNewer tokenizer inflates counts
Gemini 3.7 FlashNoHigh-volume, low-cost tasksPrice doubles in 2027
Kimi K3YesSelf-hosting, codingSix-month capability lag
DeepSeek V4 FlashVariesBulk classification, draftsPeak/off-peak billing complexity

Verdicts. Sol wins when a mistake is expensive. Sonnet 5 is the balanced default at $2/$10. Kimi K3 is the pick if you need to own the model. DeepSeek V4 Flash wins on raw cost per task, and nothing else comes close.

Who should use Chinese models — and who should skip them

Use them if you:

  • Run high-volume, low-risk work: tagging, summarising, first drafts
  • Want to self-host on your own servers for privacy or cost control
  • Are testing many prompts and need cheap iteration
  • Build tools where a six-month capability lag does not matter

Skip them if you:

  • Handle regulated client data with contractual location requirements
  • Sell to US government or defence-adjacent buyers
  • Need the strongest available agentic reasoning today
  • Cannot absorb sudden policy changes that restrict a vendor

That last point is real. Officials have publicly floated restricting Chinese AI models in the US, and the US and China AI competition makes that policy risk hard to price.

Kimi K3 vs GPT-5.6 vs DeepSeek V4, by the job

Job: ship a feature on a small budget. DeepSeek V4 Flash. At $0.44 per million input tokens peak, you can afford to be wasteful with prompts. Pair it with our best AI tools worth paying for in 2026 list.

Job: run a model you fully control. Kimi K3. Public weights mean no vendor can change your terms mid-project. Read our Moonshot AI and Kimi K3 breakdown before committing hardware.

Job: automate something that touches money or clients. GPT-5.6 Sol or Claude Opus 5. The premium buys reliability, and a single bad output costs more than the token difference.

Qwen sits between these; our Qwen AI overview covers what it does well.

Recent changes and what has already expired

  • 13 January 2026: Commerce announced case-by-case licence review for H200 and MI325X exports to China. Under Secretary Jeffrey Kessler said permitting H200 sales “under controlled conditions will strengthen the American technology ecosystem.”
  • 15 January 2026: The rule took effect. It covers chips under 21,000 TPP and 6,500 GB/s DRAM bandwidth, and caps China-bound shipments at 50% of the same product shipped for US end use.
  • May 2026: Chinese authorities barred domestic AI firms from buying H200s. Brookings found none had been sold by mid-May.
  • 1 June 2026: BIS clarified that licence requirements apply to any company headquartered or parented in China, wherever its subsidiary sits.
  • 16–17 July 2026: Moonshot released Kimi K3 and published its weights.

Already dead, still searched: the Biden-era AI Diffusion Rule was rescinded in May 2025 — do not plan around it. The H20 chip era is over; H200 is the current licensed ceiling. And the 2025 arrangement giving the US government 15% of certain China chip sales has been overtaken by the January 2026 rule and its 50% cap.

How to stop using a model or cut your spend

API access is pay-as-you-go, so “cancelling” means stopping spend, not clicking one button. Do it in this order:

  1. Revoke or delete the API key in the vendor’s console. This halts all charges immediately.
  2. Turn off auto-recharge or auto top-up before your balance triggers another payment.
  3. Downgrade rather than delete if you may return — most consoles keep usage history.
  4. Export your prompt logs and fine-tuning data first; deletion is often permanent.
  5. Check the vendor’s own billing page for refund terms on prepaid credit before topping up again.

For consumer subscriptions, cancel inside the account settings of the app itself. Our Google AI Pro price and limits guide walks through one example in detail.

FAQ

Who is winning the US and China AI competition in 2026? Neither outright. The US leads on frontier capability, chip supply and capital, while China leads on price and open weights. A joint US–UK assessment placed Kimi K3 about six months behind US models, so the gap is months, not years.

Are Chinese AI models safe for business use? It depends entirely on your data. For public marketing content and internal drafts, the risk is low. For regulated or client-confidential material, use a US vendor or self-host open weights on infrastructure you control and audit yourself.

Can China buy Nvidia chips now? Legally, some can. Commerce has allowed case-by-case H200 licences since 15 January 2026, capped at 50% of US-bound volume. But Chinese authorities blocked domestic AI firms from buying H200s in May 2026, so approved supply sits unsold.

Why are Chinese AI models so much cheaper? Lower energy costs, efficiency-focused training methods, and a deliberate strategy of winning developers through price rather than benchmarks. DeepSeek V4 Flash runs at $0.44 per million input tokens at peak and $0.22 off-peak, roughly a tenth of comparable US rates.

Does open-weight release mean the model is free? No. You can download and run Kimi K3’s weights without a licence fee, but you pay for GPUs, hosting and electricity, which often exceeds API cost at low volume. Hosted access still bills the published per-token rate.

Verdict

The US and China AI competition has stopped being a single race and become a split market: America sells capability, China sells price and control. For most creators and small teams, the practical answer is to use a cheap Chinese or open-weight model for volume work and a US frontier model for anything that touches money. Next step: price one real workflow at both rates using the vendor pages linked above before you commit a budget.


About AI Era

AI Era reviews AI tools for creators, marketers and small businesses. We take pricing only from vendor documentation and policy facts only from primary government filings. When a figure is not published officially, we say so instead of estimating. Nothing here is sponsored.

Sources: Commerce/BIS press release · Federal Register rule · Kimi K3 pricing · DeepSeek API pricing · Gemini API pricing · Claude pricing · OpenAI API pricing

Similar Posts