Updated for GPT-6.1 Sol · released 29 Sep 2026

What does your OpenAI request actually cost?

Almost everyone leaves service_tier on the default. It is the single biggest lever on an OpenAI bill that nobody touches. Compose your real cost here — tier, caching, reasoning effort and the 272K long-context cliff — across GPT-6.1 Sol, GPT-6 Astra and GPT-6 Luna.

1. What are you doing?

This sets a starting configuration. Everything stays editable.

2. Tune the request

Defaults are typical for an API-backed app.



×3
Regional (EU) processing endpoint +10%

Cost

Every line below is billed separately on the usage dashboard.

$0 / call
Uncached input
$0
Cache read
$0
Output
$0
Thinking tokens
0
Per day
$0
Per 30 days
$0
Blended rate
$0 / M

Cache writes: $0 per day (set the count above; each write bills at 1.25× the input rate).

Same workload, every tier

Cheapest first.

TierRateSpeedPer callPer 30 days

Same workload, every model

At the tier you selected. Sorted on your inputs, not on a benchmark score.

ModelPrice / 1MPer callPer 30 days

Monthly cost at a glance

Same numbers, drawn to scale.

Two levers, and almost everyone only pulls one

An OpenAI request has a headline rate — $2 / $10 per million tokens for GPT-6.1 Sol — and almost every cost calculator stops there. But the headline rate is only what you pay on the Standard tier, at short context, with no caching. Three multipliers sit on top of it, and they are where the money actually is.

The rule of thumb: set the tier to the slowest one your workload tolerates, then cache the stable part of the prompt, then keep effort at the lowest level that still produces a correct answer. In that order.

Lever 1 — the service tier

service_tier is a per-request parameter, not a plan you buy. Batch and Flex run at half the standard rate. Fast runs at double. Ultrafast runs at six times — and on GPT-6 Astra that is six times the speed for exactly six times the price, published as $300 per million output tokens against a standard $50.

The failure mode is not picking a bad tier. It is never picking one at all: overnight enrichment, backfills, eval runs and embeddings refreshes sit on Standard at full price for no reason, because the default is Standard and nobody revisits it.

Lever 2 — the 272K cliff

Above 272,000 input tokens the whole request is re-priced: input and cached input double, output rises 1.5×. The premium applies to every token in the request, not just the tokens past the line. A 271,999-token prompt and a 272,001-token prompt are billed at rates that differ by 100% on the input side.

That makes prompt trimming a first-class cost lever, not a housekeeping chore. If an agent loop drifts past the line, the fix is to shorten the history — not to buy a faster tier.

Lever 3 — reasoning effort

Reasoning tokens are billed as output tokens, and the output rate is five times the input rate on every current model. Raising effort therefore multiplies the expensive half of your bill while leaving the cheap half untouched.

The multipliers in the slider above are planning assumptions, not vendor-published figures — OpenAI does not publish a thinking-token multiplier per effort level. Measure your own traffic, then set them to match.

Where each tier belongs

TierRateReach for it when
batch0.5×Overnight enrichment, backfills, eval runs, embeddings refresh
flex0.5×Non-production traffic, unpredictable spikes, retryable jobs
standard1×Interactive traffic you cannot predict
fast2×Latency-critical requests — bought per request, never as a default
ultrafast6×Bulk generation where wall-clock time is the constraint. Astra only, no EU endpoint

Two mistakes that cost real money

  1. Leaving the tier on the default forever. Standard is the correct choice for interactive traffic and the wrong one for everything else. Sort your workloads by whether a human is waiting, then move the ones that are not onto Batch or Flex.
  2. Paying for a faster tier when the prompt is the problem. If a request is slow because it carries 300K tokens of history, Fast mode makes an expensive request expensive and quick. Trim it below the cliff instead.

Frequently asked

Is Ultrafast worth six times the price?

Only when wall-clock time is the actual constraint. OpenAI publishes up to 8× faster token generation in Codex and up to 6× in the API — which is exactly the price multiple, so you are buying speed at par, not at a premium or a discount. It is also Astra-only today, runs at deliberately low default rate limits, and has no EU endpoint.

Do these prices include prompt caching?

They include it as a lever you control. A cache read on GPT-6.1 Sol costs $0.10 per million against $2.00 uncached — a 95% discount — while a cache write costs $2.50, or 1.25× the input rate. Set the hit rate and write count above and the calculator composes both.

How current are the numbers?

The data file behind this page records a verification date, shown in the footer. Rates in this market move monthly, and service tiers were still changing through the second half of 2026 — always confirm against the vendor's own pricing page before you commit a budget.