Two frontier models shipped inside 48 hours in July, and the scoreboard is the least interesting part of it. xAI answered the pricing question with one aggressively cheap sheet. OpenAI answered the same question with a menu of three. Put the two price lists side by side – a comparison neither launch made – and the lowest input rate on the page belongs to the tier nobody called a price war.

Grok 4.5 and GPT-5.6 Arrive Within Two Days of Each Other
xAI introduced Grok 4.5 on July 8, with developer API access first and public rollout on July 9. It is a mixture-of-experts model built on the company’s 1.5-trillion-parameter V9 foundation, trained on tens of thousands of Nvidia GB300 GPUs inside the Colossus complex in Memphis, and it ships with a 500,000-token context window.
Bloomberg reported the model was developed with Cursor – whose parent Anysphere agreed to an all-stock merger with SpaceX valued around $60 billion – and trained on trillions of tokens of Cursor interaction data, making it the first frontier model built directly on a coding IDE’s usage corpus.
OpenAI began its public, global rollout of GPT-5.6 the following day, after a limited preview, coordinated with the US government, that reached only trusted partners. Sol is the flagship for complex coding, research, biology and cybersecurity, with a maximum-reasoning setting and an “Ultra” mode that coordinates multiple subagents on a single task. Terra is the balanced production tier; Luna is built for speed and volume. The family moves to a 1.5-million-token context window, and OpenAI says ChatGPT, Codex and wider API access follow over the coming weeks.
Anthropic restored broad access to its newest models in the same window. All four of the rate cards below were on the table in front of buyers inside a single business week.
| Model / tier | Input | Output | Context |
|---|---|---|---|
| Grok 4.5 | $2.00 | $6.00 | 500,000 |
| GPT-5.6 Sol | $5.00 | $30.00 | 1,500,000 |
| GPT-5.6 Terra | $2.50 | $15.00 | 1,500,000 |
| GPT-5.6 Luna | $1.00 | $6.00 | 1,500,000 |
The Bill Stopped Following the Rate Card
Both launches make the same argument, and they make it in opposite directions. xAI reports that Grok 4.5 resolves a SWE-Bench Pro task in 15,954 output tokens where a rival model needs 67,020 – a 4.2x gap on the company’s own figures. Stacked on cheaper tokens, that saving shows up in the cost of a finished task, not on the rate card. Analysts at The Decoder made the consequence explicit: the model is cheap enough that benchmark gaps may not matter for a large class of buyers.
OpenAI runs the same arithmetic backwards. Coordinating several subagents on one request multiplies token consumption, so on the hardest settings the lowest per-token rate can quietly produce the highest per-task bill. One vendor advertises the effect; the other’s product design leaves buyers to account for it. Neither is disputing that the effect decides the invoice.
That tokens-per-task is displacing benchmark rank as the operative number is already visible across the reasoning economy. What the July pair adds is an asymmetry in who does the arithmetic. xAI published a per-task token count. OpenAI published a rate card and left the per-task cost to the buyer. Only one of those two numbers can be compared across vendors, and it is not the one printed on the price sheet.
One Price, or a Menu
xAI collapsed the decision into a single sheet: one rate, one model, and the efficiency claim doing the persuading. OpenAI split the same decision into three products and handed the choice to the buyer. Sol holds the frontier, Terra is positioned near GPT-5.5 performance at roughly half the cost, and Luna is built for volume.

Terra is the tier aimed at the middle of the market, and it is the one competitors were expected to answer. The more revealing row is the bottom one. Luna lists at $1 per million input tokens and $6 per million output. Grok 4.5 lists at $2 and $6. On output they are identical; on input, the volume tier of the menu is half the price of the launch that was read as a price war.
The two products are not equivalent, and the comparison should not be stretched past what it supports. They sit at different capability tiers, carry different context windows, and Grok’s per-task efficiency claim has no published Luna counterpart. The rate cards are comparable; the products are not. But a buyer sizing a high-volume, low-complexity workload reads rate cards first, and on that page the aggressive number came from the tier that was never framed as aggressive.
This reshapes what a buyer is actually doing. The problem becomes selection rather than ranking – which tier, for which workload, at which context length – and three overlapping tiers add real selection risk, because standardizing on the wrong one means overpaying or under-serving. Cloud compute matured into instance families along the same path.
The One Yardstick Both Sides Reached For
The benchmark picture is mixed, and xAI barely disputes it. On the company’s own numbers, Grok 4.5 beats Opus 4.8 on DeepSWE 1.0 (62% against 55.75%) and roughly ties GPT-5.5 on Terminal-Bench 2.1, while losing DeepSWE 1.1 by six points and SWE-Bench Pro by 4.5. Artificial Analysis placed it fourth overall, and Elon Musk’s own framing – that it “competes with last year’s Claude Opus” – was unusually modest for a frontier launch.
Third-party launch testing put Sol Ultra at 91.9% on Terminal-Bench 2.1, with plain Sol at 88.8%, GPT-5.5 at 88.0% and Gemini 3.1 Pro Preview at 70.7%. Those are launch figures, not independently audited results.
Terminal-Bench 2.1 is the one measurement both launches reached for, which makes it the closest thing to a shared yardstick in the week – and both sets of numbers arrived through a launch rather than through independent replication. There is a separate evidence problem on the xAI side: reporting indicates an earlier Cursor codebase snapshot leaked into the training data, which inflates any Cursor-derived score, though independent suites are unaffected.
So the ranking is contested and the pricing is published. A buyer can verify only one of those today, which goes some way toward explaining why the pricing did the work.
What Sits Underneath Both Price Sheets
There are two reasons to doubt these rates survive contact with volume. The first is subsidy: pricing this aggressive may not be self-funding, and whether it holds through the Anysphere merger closing and real utilization is unknown. The second is compute. A 1.5-million-token context and subagent orchestration both raise computation per query, feeding the same accelerator and memory squeeze the rest of the industry is fighting over. Per-token prices are falling and per-query compute is climbing, and this pair of launches leans into both.
Both pressures point at the same layer. The cheaper tokens become, the more of the remaining margin sits with whoever supplies the compute underneath – which is why frontier labs are negotiating directly for manufacturing capacity, and why a lab cut off from the leading edge is designing its own inference silicon. On the buyer side the effect is blunter: a mid-tier at roughly half the previous cost lowers the price of putting AI behind every seat.
Buyers can produce the missing number themselves before either vendor does. Run the same finished job on both, count the invoice instead of the rate, and the ranking that matters appears. Whether independent evaluations reproduce either benchmark set, and whether the $2/$6 rates and Terra’s half-price position survive real volume, will show up in that figure first – and until a vendor publishes it, the buyer is the only party holding it.
Where the gap between rate and bill actually comes from
The wedge between a published rate and a delivered invoice is not mysterious, and by 2026 it has three named components. Internal reasoning tokens – the model’s own working, which the customer never reads – are billed at the higher output rate on essentially every reasoning model, so a task that thinks harder costs more without producing a longer answer. Tool-enabled requests carry roughly three to seven hundred extra input tokens each before any user content arrives.

And cached input now runs about 90% below list at both major vendors, which means two buyers on identical rate cards can face effective prices that differ by an order of magnitude depending on how repetitive their prompts are.
| Model | Input | Output |
|---|---|---|
| DeepSeek R1 | $0.55 | $2.19 |
| o4-mini | $1.10 | $4.40 |
| Claude Sonnet 4.x | $3.00 | $15.00 |
| Claude Opus 4.6 | $5.00 | $25.00 |
| o1 | $15.00 | $60.00 |
Those three together explain why a rate card stopped predicting a bill, and they cut in opposite directions. The first two inflate the invoice invisibly; the third deflates it invisibly. A buyer comparing vendors on published input price is comparing the one number that all three of these effects sit on top of.
The comparison across the frontier makes the point sharper than any single sheet. DeepSeek R1 lists around $0.55 per million input and $2.19 per million output, roughly 96% below OpenAI’s o1 at $15 and $60 – a gap wide enough that it looks like the argument is over. It is not, because o4-mini at $1.10 and $4.40 performs close enough on most reasoning work, which puts the real spread nearer 2x than 27x once the task rather than the token is held constant.
Claude Sonnet at $3 and $15 and Opus at $5 and $25 sit above both. Ranking those five by list price produces one order; ranking them by cost to finish a job produces another, and nobody publishes the second.
Same Diagnosis, Opposite Remedies
Two labs priced within a day of each other and agreed on the premise: a rate card no longer tells a buyer what a job will cost. They disagreed on the remedy, one sheet against three tiers, and the cheapest published input rate ended up belonging to the menu’s volume tier rather than to the launch that was read as a price war. Until per-task cost is published on terms that can be compared across vendors, the number buyers most need is the one neither company supplies.
The Packaging Is the Story, Not the Model
The durable move in July was the packaging, not the model. Segmenting intelligence into named, separately priced tiers is how a market stops competing on a single score and starts competing on fit, and once a buyer is choosing a tier instead of a leader, the leaderboard becomes a marketing input rather than a purchase criterion. The evidence currently tilts toward that being what outlasts this cycle: these scores will be superseded within months, and the pricing structure will not.
Rollout Status for Both Launches
Not fully. GPT-5.6 launched as a limited preview to trusted partners through the API and Codex, then began a global rollout on July 9, with OpenAI saying ChatGPT, Codex and wider API access follow in the weeks after launch. Grok 4.5 reached developer API access on July 8 and public rollout the next day.
Sources
- x.ai — xAI’s official Grok 4.5 announcement (pricing, benchmarks, V9/GB300 specs) – direct fetch returned HTTP 403 on every attempt; date and figures carried from the original author’s citation, not independently re-confirmed here. (2026-07-08)
- openai.com — OpenAI’s official GPT-5.6/Sol-Terra-Luna announcement – direct fetch returned HTTP 403 on every attempt; date and figures carried from the original author’s citation, not independently re-confirmed here. (2026-07-08)
- artificialanalysis.ai — Artificial Analysis leaderboard – fetched and confirmed as the site the draft names, but the current homepage does not show a Grok 4.5 entry; the “fourth overall” placement was not independently re-confirmed in this pass and the date is carried from the original author’s citation. (2026-07-09)
- inference.net — Inference.net cross-vendor LLM API pricing comparison – confirms reasoning/thinking tokens billed at the output rate and roughly 80-90% prompt-caching discounts; the specific DeepSeek R1/o1/o4-mini rows the draft cites were not located in this fetch’s extracted view (the page covers 30+ current models and may have dropped older/deprecated rows since the original citation). (2026-02-21)
View all sources
- aipricing.guru — AI Pricing Guru’s GPT-5.6 tier rate page, fetched and quoted verbatim; the page is continuously updated (first published 2026-04-03) and now reflects a 2026-07-30 price cut, so its current Terra/Luna numbers ($2.00/$12, $0.20/$1.20) differ from the draft’s launch-day figures ($2.50/$15, $1/$6), which match the pre-cut pricing reported elsewhere at launch. (2026-08-06)
- bloomberg.com — Bloomberg report on GPT-5.6’s rollout, Cursor training-data origin and the Anysphere-SpaceX merger context – direct fetch returned HTTP 403 on every attempt; date carried from the original author’s citation. (2026-07-08)
- venturebeat.com — VentureBeat report on the three-model launch and the US-government-coordinated limited preview – direct fetch returned HTTP 429 (rate-limited) on repeated attempts; date carried from the original author’s citation. (2026-07-08)
- the-decoder.com — The Decoder’s pricing-over-benchmarks analysis – the citation is the bare homepage URL rather than a specific article link, and the current homepage does not surface the original piece; date carried from the original author’s citation. (2026-07-09)
- datacamp.com — DataCamp explainer on Sol/Terra/Luna tier roles, fetched and quoted verbatim: “Terra has competitive performance with GPT-5.5 while being about 2x cheaper.” Page schema metadata dates the article to 2026-06-26. (2026-06-26)
- edenai.co — Eden AI benchmark roundup, fetched and quoted verbatim: Sol Ultra 91.9%, Sol 88.8%, GPT-5.5 88.0% on Terminal-Bench 2.1, matching the draft exactly. Byline shows “last updated on July 11, 2026.” (2026-07-11)
- aitoolsreview.co.uk — AI Tools Review piece on GPT-5.6 tier pricing and use cases, fetched and quoted verbatim; the page states a 1,050,000-token context window and does not compare it to Gemini, which differs from the draft’s 1.5-million-token figure (sourced to OpenAI’s own launch material, C3) – cited here for its general tier-differentiation framing, not the context-window number. (2026-08-01)
- techmymoney.com — TechMyMoney report confirming the July 9, 2026 public rollout of Sol/Terra/Luna beyond the limited partner preview, fetched and quoted verbatim. (2026-07-08)
This article is for informational and educational purposes only and does not constitute investment, financial, or legal advice.