A service tier
Ultrafast is a new service tier — the setting that decides how your request is processed. It is not a new model, and it is not a piece of silicon you buy.
Independent explainer · built from OpenAI’s public announcement of 18 August 2026
OpenAI previewed Ultrafast — a new service tier, not a new model and not a new chip — that runs GPT‑5.6 Sol at up to 14× faster than Standard processing, launching first in the OpenAI API. It is powered by Cerebras and generates up to 750 output tokens per second. This page turns those two numbers into something you can watch, press and time.
Before you scroll: how fast do you think it is? Set both dials, then reveal.
01 The announcement
Five facts carry the whole story. Everything else on this page is an attempt to make them physical.
Ultrafast is a new service tier — the setting that decides how your request is processed. It is not a new model, and it is not a piece of silicon you buy.
The model on the other end is the same frontier model, described by OpenAI as its most intelligent model. Nothing is swapped for something smaller.
Against Standard processing as the baseline. The qualifier matters: the source says up to, and this page never drops it.
The next step in OpenAI’s partnership with Cerebras for ultra‑low‑latency inference on its platform — now supporting OpenAI’s most intelligent model.
Available today in a limited preview to a select group of customers, launching first in the OpenAI API. Access expands as capacity grows.
Until now, getting real‑time speed typically meant choosing a smaller or more specialized model. OpenAI frames Ultrafast as progress in a new direction: more useful work per second.
Read this as: the announcement is about removing a trade-off, not about raising a benchmark. Frontier intelligence and real-time speed used to be a choice; the claim is that on this tier they no longer are.
02 Station 02 · Showdown
OpenAI’s own demonstration puts Ultrafast and Standard side by side, building a working 3D warehouse simulator from the same text prompt. Press start and watch the same answer arrive twice.
Build a working 3D warehouse simulator.
Our wording — the announcement says only that both tiers built the
simulator “from the same text prompt”.
Stepped presentation. Motion is reduced, so the same run is shown as timed snapshots. Step through them — each card is the state of both lanes at that moment.
Take away: the gap is not a nicer progress bar. At Standard pace you are waiting for an answer; at Ultrafast pace the answer arrives inside the moment you asked the question — which is what makes it usable mid‑conversation, mid‑outage, or mid‑checkout.
03 Station 03 · The misconception
It is the most common misreading of this announcement, and it is an easy one to make — “powered by Cerebras” sits right next to a hardware company’s name. Three different layers get collapsed into one. Pull them apart and the whole thing clicks.
Standard selected. The request goes to the same GPT‑5.6 Sol model on OpenAI’s standard processing path. This is the 1× baseline that the “up to 14×” figure is measured against.
Take away: your application is the same, the request is the same, and — the part people miss — the model is the same. What the tier changes is the bottom band, and the announcement names only Ultrafast’s: powered by Cerebras. That is where up to 14× and up to 750 output tokens per second come from. What Standard processing runs on is not stated, so this page leaves that band unnamed.
04 Magnitude
“Up to 14×” is a handful of printed characters. Here it is three other ways, all built from the same two stated figures.
One second of output, drawn to scale.
The Standard bar is 1⁄14th of the Ultrafast bar. If it looks like a rounding error, that is the point — that sliver is the whole of ordinary generation speed at this scale.
How many Standard seconds fit inside one Ultrafast second.
Fourteen ticks of the ordinary clock, collapsed into one. This is the up to 14× figure as a shape rather than a label.
Every dot is one output token. The lit dots are what Standard manages in the same second.
750 tokens is roughly a long email, or a short function with its tests — produced while you are still lifting your finger off the key.
05 Station 04 · Workload dial
Speed only matters in proportion to the work. Set a response length and watch both tiers report the same answer at two very different moments.
Or jump to a shape of work:
Take away: the multiple is constant, but the felt difference is not. At 180 tokens both tiers feel quick. Somewhere past a thousand tokens, Standard becomes a wait you plan around — and that is exactly the range where interactive products live.
06 Station 05 · Reflex
A human reaction to a visual signal is a real, measurable interval. Here is yours — and here is what the stated ceiling would put on the page inside it.
Mouse, touch or keyboard — Enter and Space both work. Three rounds; your best time counts.
Arm the pad to take your measurement.
Take away: whatever your own number turns out to be, at the stated ceiling the answer is already a paragraph deep before you have consciously registered that anything happened. That is what the announcement is pointing at when it talks about products changing once the model can “keep pace with the person using it”.
07 Where the value showed up
The announcement names five scenarios it calls encouraging. Each one has a different clock running underneath it — and each miniature below draws the clock that scenario is fighting.
When a critical system fails: analyze application logs, recent code changes and engineer reports to identify the likely cause and help prepare a fix while the outage is still unfolding.
The clock Every minute of the outage is the cost.
Analyze market signals, assess transactions and identify suspicious activity while conditions are still changing.
The clock The answer expires as it is written.
Resolve complex customer issues in real time without interrupting the conversation, even when finding the answer requires multiple steps or systems.
The clock A pause in speech is heard as a fault.
Answer product questions, check inventory, personalize recommendations and resolve checkout issues while the shopper is still deciding — before hesitation becomes an abandoned cart.
The clock Attention leaves before the answer lands.
Turn research that previously took an overnight run into an interactive working session — test an idea, examine the results, adjust the approach and run another experiment without breaking flow.
The clock One idea per night, or one per coffee.
Take away: none of these are about producing more text. They are about arriving before the window shuts — while the outage is live, while the market moves, while the caller is still on the line, while the cart is still open, while the thought is still warm.
08 Station 06 · Triage
Five calls. For each one, decide whether it needs frontier intelligence at real‑time speed, or whether it is fine on Standard processing. Feedback comes straight from the announcement.
Which tier fits?
Take away: the tier is a judgement about the window, not about the difficulty. A hard question with a patient reader is a Standard job. An ordinary question with a shrinking window is where the announcement points Ultrafast.
09 The loop that changed shape
The clearest before-and-after in the announcement is not a benchmark. It is a working rhythm.
1 iteration per day
Several iterations per day
Take away: the interesting unit is not tokens per second, it is attempts per day. When the wait falls below the length of a thought, the loop stops being a schedule and starts being a conversation.
10 Dogfooding
Inside OpenAI, a group of developers has been testing GPT‑5.6 Sol on Ultrafast mode to understand which workflows benefit from frontier intelligence that can answer in real time.
When an alert fires, engineers need to build an accurate picture while the system and the evidence are still changing.
All in a fraction of the time, with the intelligence of Sol. It shortens the delay between observing a signal, testing a hypothesis and choosing the next action.
Engineers remain responsible for judgement and deployment. The announcement is explicit about this: the speed is in the analysis, not in the decision.
The research team uses Ultrafast to rapidly search knowledge sources, query data, and quickly gather, organize and summarize information across connected tools.
The overnight‑batch habit gives way to a loop tight enough to run several times before the day is out.
See the loop that changed shape for that before‑and‑after drawn out.
11 The preview cohort
OpenAI says it has been testing GPT‑5.6 Sol on Ultrafast mode with an initial group of companies, in real production environments, to learn where an order‑of‑magnitude change in speed creates the most value.
Rendered as plain text, not logos. No usage volumes, contract terms or pricing are stated in the source, and none are invented here.
“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”
12 Where this actually is
This is the part most easily skimmed past, so here it is at full size.
In the OpenAI API. Launching first there.
Not a general release. Limited preview, select customers.
Not a new model. GPT‑5.6 Sol, unchanged.
Not an OpenAI chip. Powered by Cerebras.
Take away: every figure on this page carries an “up to” and sits behind a limited preview. Those two qualifiers are load‑bearing, and dropping either one would misstate what was announced.
13 Your run
Nothing attempted yet — every station is still open.