GPT-6 Astra pricing, $10 input and $50 output per 1M tokens

Every GPT-6 Astra rate from the official OpenAI pricing page and model card, the long context surcharge over 272K tokens, and three worked jobs priced on Standard, Batch and Fast mode.

Every price read from the official OpenAI pricing page and model card on 6 September 2026, cross-checked the same day against OpenRouter and LiteLLM, which supply the listing timestamp and registry figures quoted in the methodology. Prices change, so treat this as a dated snapshot.

GPT-6 Astra costs $10.00 per 1M input tokens and $50.00 per 1M output tokens on the Standard tier, with cached input at $1.00 and cache writes at $12.50. Batch and Flex halve every one of those rates to $5.00, $0.50, $6.25 and $25.00, and Fast mode doubles them to $20.00, $2.00, $25.00 and $100.00.

$10 / $50input and output per 1M, Standard
$5 / $25Batch and Flex, half price
2x in, 1.5x outover 272K input tokens
1,050,000token context window

There is a second price column that most quick comparisons miss. Once a prompt goes past 272K input tokens, the official model card says the whole request is billed at 2x the input and cache rates and 1.5x the output rate, so Standard becomes $20.00 in and $75.00 out. That makes GPT-6 Astra 2.5x the price of GPT-5.6 Sol on both input and output, and the long context surcharge stacks on top of that.

GPT-6 Astra price tables, all four tiers

All four rows come from the developers.openai.com pricing page, which groups each tier into a short context block (up to 272K input tokens) and a long context block (over 272K). USD per 1M tokens.

gpt-6-astra by service tier, USD per 1M tokens, official OpenAI pricing page read 6 September 2026

TierContextInputCached inputCache writesOutput
StandardUp to 272K$10.00$1.00$12.50$50.00
StandardOver 272K$20.00$2.00$25.00$75.00
BatchUp to 272K$5.00$0.50$6.25$25.00
BatchOver 272K$10.00$1.00$12.50$37.50
FlexUp to 272K$5.00$0.50$6.25$25.00
FlexOver 272K$10.00$1.00$12.50$37.50
Fast modeUp to 272K$20.00$2.00$25.00$100.00
Fast modeOver 272K$40.00$4.00$50.00$150.00

Three multipliers explain the whole grid, and the model card states each of them. Cache writes bill at 1.25x the uncached input rate, Batch and Flex at 50 percent of Standard, Fast mode at 2x the applicable rates. Two footnotes on the pricing page matter before you budget. Fast mode is unavailable for GPT-6 Astra with EU data residency, and regional processing endpoints carry a 10 percent uplift for models released on or after March 5, 2026 that are eligible for data residency.

The model card itself lists a 1,050,000 token context window, 128,000 max output tokens, an Apr 30, 2026 knowledge cutoff, text and image input with text output, reasoning token support, and a reasoning.effort parameter with levels low, medium, high, xhigh and max. The card describes the model in one sentence, "our most capable model, built for the hardest end-to-end work", and this page repeats nothing beyond that about capability.

Three worked cost examples

Each figure below was computed in Python from the table above. Example one is a single agentic task that reads 60,000 tokens of code and tool output and writes 8,000 tokens back.

One task, 60,000 input and 8,000 output tokens, no cache hits, USD

TierInput costOutput costTotal
Standard$0.60$0.40$1.00
Batch or Flex$0.30$0.20$0.50
Fast mode$1.20$0.80$2.00

Example two is the same task after the context has grown to 400,000 input tokens, which is where the long context block takes over. The model card wording matters here. Prompts with more than 272K input tokens are priced at the higher rates "for the full request", not just for the tokens beyond the threshold, so all 400,000 input tokens bill at $20.00 and the 8,000 output tokens bill at $75.00.

One task, 400,000 input and 8,000 output tokens, long context rates on the full request, USD

TierInput costOutput costTotal
Standard$8.00$0.60$8.60
Batch or Flex$4.00$0.30$4.30
Fast mode$16.00$1.20$17.20

Had the short context rates applied, the Standard request would have been $4.40, so crossing the line costs 1.955x on this shape of task. Put another way, 272,000 input tokens at the short rate cost $2.72 while 400,000 at the long rate cost $8.00, a 47 percent bigger prompt for a 2.94x bigger input bill.

Example three is a month of production use. Take 5,000 of the short tasks from example one on Standard. With no caching the month costs $5,000.00. If 80 percent of each prompt is a shared system prompt and toolset served from the cache, then 48,000 tokens bill at $1.00 and 12,000 at $10.00, output is unchanged at $0.40, and each task drops to $0.568.

The month becomes $2,840.00, a saving of $2,160.00 or 43.2 percent. Writing those 48,000 shared tokens into the cache once costs $0.60 at the $12.50 cache write rate, against $0.48 as plain input. The same month on Batch is $2,500.00 without caching and $1,420.00 with it. Our prompt caching break-even calculator runs this arithmetic for any hit rate.

GPT-6 Astra against the GPT-5.6 models

The pricing page lists Astra on the same Standard table as the three GPT-5.6 general models (a separate Cyber table lists gpt-5.6-cyber), which makes the ladder easy to read. The ratios below are input and output rate divided by the sibling's rate.

Standard tier, up to 272K input tokens, USD per 1M tokens, official OpenAI pricing page read 6 September 2026

ModelInputCached inputCache writesOutputAstra multiple60k in, 8k out task
gpt-6-astra$10.00$1.00$12.50$50.001x$1.00
gpt-5.6-sol$4.00$0.40$5.00$20.002.5x in, 2.5x out$0.40
gpt-5.6-terra$2.00$0.20$2.50$12.005x in, 4.17x out$0.22
gpt-5.6-luna$0.20$0.02$0.25$1.2050x in, 41.67x out$0.02

Two things on that table are not obvious from the numbers alone. The pricing page notes that GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026, so the 2.5x gap is measured against a promotional rate. And every GPT-5.6 row uses the same 272K threshold and the same long context rule, so the ladder holds at both context sizes. Sol's long context Standard rate is $8.00 in and $30.00 out against Astra's $20.00 and $75.00.

Sources and cross-checks

Primary source is the OpenAI developer pricing page and the gpt-6-astra model card, both saved as text on 6 September 2026. Every rate in the tables above is quoted from the four tier tables on that page. Two secondary sources were read the same day. The OpenRouter models API lists openai/gpt-6-astra and openai/gpt-6-astra:batch, both with a created timestamp of 1788552838, which is 4 September 2026 20:13 UTC, the OpenRouter listing date.

Its per token prices, multiplied by one million, match the official Standard and Batch short context rows exactly, and its override block at 272,000 minimum prompt tokens matches the long context row. The LiteLLM registry key gpt-6-astra matches every Standard, Batch, Flex and Fast mode (which it calls priority) rate it carries (it has no batch cache or batch long context fields), though it records max_input_tokens as 922,000 rather than the 1,050,000 context window on the card, which is consistent with 1,050,000 minus 128,000 max output.

One disagreement. OpenRouter lists openai/gpt-5.6-sol at $2.00 input and $10.00 output, while the official Standard table and LiteLLM both list $4.00 and $20.00. The OpenRouter figure is exactly the official Batch and Flex rate for Sol, so a calculator that reads OpenRouter alone understates Standard tier Sol by 2x and reports the Astra to Sol gap as 5x instead of 2.5x. The tables on this page use the official figure. We track this class of error on the OpenAI calculator price accuracy page.

Evidence files for this page are published at gpt-6-astra-pricing-claims.json (every figure with its source quote) and gpt-6-astra-pricing-calc.txt (the script output for the worked examples).

Frequently asked questions

How much does GPT-6 Astra cost per million tokens?

Standard tier, up to 272K input tokens, is $10.00 input, $1.00 cached input, $12.50 cache writes and $50.00 output per 1M. Over 272K the full request bills at $20.00, $2.00, $25.00 and $75.00.

What does GPT-6 Astra cost on the Batch API?

Batch and Flex are both 50 percent of Standard, so $5.00 input, $0.50 cached input, $6.25 cache writes and $25.00 output for short context, and $10.00, $1.00, $12.50 and $37.50 over 272K. Our OpenAI Batch API guide covers the queue mechanics.

How does GPT-6 Astra pricing compare with GPT-5.6 Sol?

Astra is 2.5x Sol on both input and output at Standard rates, $10.00 against $4.00 and $50.00 against $20.00. The 60,000 input, 8,000 output task above costs $1.00 on Astra and $0.40 on Sol.

What is the GPT-6 Astra context window and max output?

The model card lists a 1,050,000 token context window, 128,000 max output tokens and an Apr 30, 2026 knowledge cutoff. Price doubles on input and cache rates and rises 1.5x on output once a prompt exceeds 272K input tokens.

Why does OpenRouter show a lower GPT-5.6 Sol price than OpenAI?

On 6 September 2026 OpenRouter listed gpt-5.6-sol at $2.00 and $10.00, which equals the official Batch and Flex rate rather than the $4.00 and $20.00 Standard rate that OpenAI and LiteLLM both publish. We don't know why. Budget from the official figure.

KickLLM Margin Studio · Offline analysis app

Know your AI costs. Now plan your margin.

Turn your usage CSV into a cost breakdown, test revenue and growth assumptions, and export a report for your next pricing decision. Margin Studio runs locally with rates you supply.

Get Margin Studio — $39 Try the interactive preview → One-time purchase · Downloadable ZIP
By the same builder: GitHub · theluckystrike BeLikeNative · Grammar AI EarlyThunder · Dev Blog Bug Bounty Reality Zovo · AI Dev Tools