The 272K cliff: GPT-6.1 Sol long-context pricing
One token over 272,000 and the entire request re-prices — not just the tokens past the line. Here is the band table and the arithmetic.
Short version: below the threshold GPT-6.1 Sol bills at
$2 input / $10 output per million tokens. Above it, the same
request bills at $4 / $15. The premium is not prorated.
The two bands
| Input length | Input /1M | Cached input /1M | Cache write /1M | Output /1M |
|---|---|---|---|---|
| ≤ 272,000 tokens | $2.00 | $0.10 | $2.50 | $10.00 |
| > 272,000 tokens | $4.00 | $0.20 | $5.00 | $15.00 |
Input, cached input and cache writes double; output rises by 1.5×. The same 2× / 1.5× rule applies across the GPT-6 family, and to GPT-5.6 Sol and Terra.
Why it is a cliff and not a slope
Most tiered pricing charges a premium rate only on the tokens above the threshold. OpenAI does not. Crossing the line re-prices every token in the request, which means the cost curve is discontinuous — a step, not a ramp.
Take a request with 271,000 input tokens and 1,000 output tokens on Standard:
| Request | Input cost | Output cost | Total |
|---|---|---|---|
| 271,000 in / 1,000 out | $0.5420 | $0.0100 | $0.5520 |
| 272,001 in / 1,000 out | $1.0880 | $0.0150 | $1.1030 |
Adding 1,001 tokens — about 0.4% more input — takes the request from
$0.55 to $1.10. The input line exactly doubles. If that request
runs 2,000 times a day, the difference is roughly $33,000 a month for a
prompt that grew by a thousandth.
The one place caching narrows the gap
Cached input is the only line where the absolute increase is small: $0.10 to
$0.20 per million. It is still a doubling, but a cache-heavy workload with a
modest uncached tail feels the cliff far less than a workload that ships raw context.
Concretely, 100,000 cached tokens plus 10,000 ordinary input tokens cost
$0.03 below the line and $0.06 above it — the ratio holds, but
the absolute number stays small.
What to do about it
- Measure input length per request, not per account. Averages hide the cliff completely. Look at the 95th percentile — that is the traffic paying double.
- Trim before you upgrade. If an agent loop drifts over the line, summarising or dropping history is cheaper than buying a faster tier.
- Cache the stable prefix. A cache read is 95% off the input rate on this model, and the write that unlocks it costs 1.25× input — so a cached prefix is cheaper from the second call onwards.
- Watch the ceiling. GPT-6.1 Sol documents a 1,050,000-token context window with a 922,000-token maximum input. Long before that, you are paying double.
The threshold is per request. Splitting one 400K-token job into two 200K-token requests keeps both below the line — same tokens, half the input rate.
Work out your own number
The composer on the home page flags the cliff automatically: set input tokens above 272,000 and it shows the monthly cost before and after trimming, alongside the tier and cache levers. It runs entirely in your browser — nothing is uploaded.