Independent explainer · built from OpenAI’s public announcement of 18 August 2026

How fast is
Ultrafast?

OpenAI previewed Ultrafast — a new service tier, not a new model and not a new chip — that runs GPT‑5.6 Sol at up to 14× faster than Standard processing, launching first in the OpenAI API. It is powered by Cerebras and generates up to 750 output tokens per second. This page turns those two numbers into something you can watch, press and time.

01 Calibrate

Before you scroll: how fast do you think it is? Set both dials, then reveal.

300
Up to 750output tokens / second Generated by Ultrafast, powered by Cerebras. A ceiling, not a guarantee.
Up to 14×faster than Standard Same model, different service tier. Standard processing is the baseline.
Availability Limited
preview
Today, to a select group of customers, in the OpenAI API. Access expands as capacity grows.

01 The announcement

What was actually announced

Five facts carry the whole story. Everything else on this page is an attempt to make them physical.

F1

A service tier

Ultrafast is a new service tier — the setting that decides how your request is processed. It is not a new model, and it is not a piece of silicon you buy.

F2

Running GPT‑5.6 Sol

The model on the other end is the same frontier model, described by OpenAI as its most intelligent model. Nothing is swapped for something smaller.

F3

Up to 14× faster

Against Standard processing as the baseline. The qualifier matters: the source says up to, and this page never drops it.

F4

Powered by Cerebras

The next step in OpenAI’s partnership with Cerebras for ultra‑low‑latency inference on its platform — now supporting OpenAI’s most intelligent model.

F5

Limited preview, API first

Available today in a limited preview to a select group of customers, launching first in the OpenAI API. Access expands as capacity grows.

The framing

Until now, getting real‑time speed typically meant choosing a smaller or more specialized model. OpenAI frames Ultrafast as progress in a new direction: more useful work per second.

Read this as: the announcement is about removing a trade-off, not about raising a benchmark. Frontier intelligence and real-time speed used to be a choice; the claim is that on this tier they no longer are.

02 Station 02 · Showdown

Same prompt. Same model. Two tiers.

OpenAI’s own demonstration puts Ultrafast and Standard side by side, building a working 3D warehouse simulator from the same text prompt. Press start and watch the same answer arrive twice.

Elapsed 0.00s
Both panels stream the same illustrative answer.
Prompt Build a working 3D warehouse simulator. Our wording — the announcement says only that both tiers built the simulator “from the same text prompt”.
Ultrafast

Ultrafast tier: up to 750 tok/s

Idle
Rate 0tok/s
Tokens 0/ 480
Time 0.00s
Waiting to start
Standard

Standard tier: the 1× baseline

Idle
Rate 0tok/s
Tokens 0/ 480
Time 0.00s
Waiting to start
Illustrative. The Ultrafast lane paces at the announcement’s stated ceiling of up to 750 output tokens per second. The Standard lane paces at that ceiling divided by the stated up to 14× multiple — roughly 54 tokens per second. That baseline is our arithmetic, not an OpenAI‑published figure, and the streamed text is placeholder content. What is being demonstrated is the pacing.

Take away: the gap is not a nicer progress bar. At Standard pace you are waiting for an answer; at Ultrafast pace the answer arrives inside the moment you asked the question — which is what makes it usable mid‑conversation, mid‑outage, or mid‑checkout.

03 Station 03 · The misconception

No, OpenAI did not release a chip

It is the most common misreading of this announcement, and it is an easy one to make — “powered by Cerebras” sits right next to a hardware company’s name. Three different layers get collapsed into one. Pull them apart and the whole thing clicks.

Myth
  • “OpenAI launched its own chip.”
  • “Ultrafast is a new, faster model.”
  • “It’s a smaller model, so it’s quicker.”
  • “You can switch it on in ChatGPT.”
  • “Everyone gets it today.”
What the announcement says
  • Ultrafast is powered by Cerebras — OpenAI’s partner for ultra‑low‑latency inference on its platform.
  • It is a service tier, running the existing GPT‑5.6 Sol.
  • The point is precisely that you don’t drop to a smaller or more specialized model to get speed.
  • It is launching first in the OpenAI API.
  • It is a limited preview to a select group of customers, expanding as capacity grows.

Follow one request from your application down to the hardware

Layered diagram of one API request A request leaves your application, enters the OpenAI API where a service tier is selected — Standard or Ultrafast — is answered by the GPT-5.6 Sol model, and is executed on inference hardware. Selecting Ultrafast routes the hardware layer to Cerebras; the model layer is identical either way. 01 02 03 04 Your application A support agent, a trading tool, an incident bot, a storefront OpenAI API the service tier is chosen here — this is the whole of what’s new Standard baseline 1× Ultrafast up to 14× GPT‑5.6 Sol identical on both tiers — the frontier model is never traded down UNCHANGED Inference hardware where the tokens are actually produced STANDARD PATH

Standard selected. The request goes to the same GPT‑5.6 Sol model on OpenAI’s standard processing path. This is the 1× baseline that the “up to 14×” figure is measured against.

Take away: your application is the same, the request is the same, and — the part people miss — the model is the same. What the tier changes is the bottom band, and the announcement names only Ultrafast’s: powered by Cerebras. That is where up to 14× and up to 750 output tokens per second come from. What Standard processing runs on is not stated, so this page leaves that band unnamed.

04 Magnitude

Fourteen is a bigger number than it looks

“Up to 14×” is a handful of printed characters. Here it is three other ways, all built from the same two stated figures.

A · Length, at true proportion

One second of output, drawn to scale.

Ultrafast up to 750tok
Standard ≈54tok

The Standard bar is 1⁄14th of the Ultrafast bar. If it looks like a rounding error, that is the point — that sliver is the whole of ordinary generation speed at this scale.

B · Count, as repeated units

How many Standard seconds fit inside one Ultrafast second.

Fourteen ticks of the ordinary clock, collapsed into one. This is the up to 14× figure as a shape rather than a label.

C · What fits in one second

Every dot is one output token. The lit dots are what Standard manages in the same second.

≈54 tokens — one Standard second up to 750 tokens — one Ultrafast second

750 tokens is roughly a long email, or a short function with its tests — produced while you are still lifting your finger off the key.

Illustrative. Up to 750 output tokens per second and up to 14× are the announcement’s figures. The ≈54 tokens-per-second Standard comparison is our arithmetic (750 ÷ 14) shown so the ratio can be drawn, not a published OpenAI number.

05 Station 04 · Workload dial

Choose a job. See the wait.

Speed only matters in proportion to the work. Set a response length and watch both tiers report the same answer at two very different moments.

1,500output tokens

Or jump to a shape of work:

Ultrafast 2.00s at up to 750 tok/s
Standard 28.0s at ≈54 tok/s

Illustrative. Both times are our arithmetic — tokens divided by a rate — built from the announcement’s stated ceiling of up to 750 output tokens per second and its stated up to 14× multiple. Real requests include time to first token, tool calls and network latency that this simple division ignores. OpenAI publishes no per‑length timings.

Take away: the multiple is constant, but the felt difference is not. At 180 tokens both tiers feel quick. Somewhere past a thousand tokens, Standard becomes a wait you plan around — and that is exactly the range where interactive products live.

06 Station 05 · Reflex

You, versus one Ultrafast second

A human reaction to a visual signal is a real, measurable interval. Here is yours — and here is what the stated ceiling would put on the page inside it.

Mouse, touch or keyboard — Enter and Space both work. Three rounds; your best time counts.

1 2 3
Your best ms
Ultrafast would have written tokens
Standard would have written tokens

Arm the pad to take your measurement.

Illustrative. The token counts are our arithmetic: your measured reaction time multiplied by the announcement’s stated ceiling of up to 750 output tokens per second, and by that ceiling divided by the stated up to 14×. Nothing here is an OpenAI‑published measurement, and your browser’s timing has its own margin of error.

Take away: whatever your own number turns out to be, at the stated ceiling the answer is already a paragraph deep before you have consciously registered that anything happened. That is what the announcement is pointing at when it talks about products changing once the model can “keep pace with the person using it”.

07 Where the value showed up

Five places where a second is expensive

The announcement names five scenarios it calls encouraging. Each one has a different clock running underneath it — and each miniature below draws the clock that scenario is fighting.

Incident response & reliability

When a critical system fails: analyze application logs, recent code changes and engineer reports to identify the likely cause and help prepare a fix while the outage is still unfolding.

The clock Every minute of the outage is the cost.

Financial research & security

Analyze market signals, assess transactions and identify suspicious activity while conditions are still changing.

The clock The answer expires as it is written.

Customer support & voice

Resolve complex customer issues in real time without interrupting the conversation, even when finding the answer requires multiple steps or systems.

The clock A pause in speech is heard as a fault.

Commerce

Answer product questions, check inventory, personalize recommendations and resolve checkout issues while the shopper is still deciding — before hesitation becomes an abandoned cart.

The clock Attention leaves before the answer lands.

Live research & experimentation

Turn research that previously took an overnight run into an interactive working session — test an idea, examine the results, adjust the approach and run another experiment without breaking flow.

The clock One idea per night, or one per coffee.

Take away: none of these are about producing more text. They are about arriving before the window shuts — while the outage is live, while the market moves, while the caller is still on the line, while the cart is still open, while the thought is still warm.

08 Station 06 · Triage

Does this job actually need it?

Five calls. For each one, decide whether it needs frontier intelligence at real‑time speed, or whether it is fine on Standard processing. Feedback comes straight from the announcement.

Call 1 of 5

Which tier fits?

Take away: the tier is a judgement about the window, not about the difficulty. A hard question with a patient reader is a Standard job. An ordinary question with a shrinking window is where the announcement points Ultrafast.

09 The loop that changed shape

From overnight batch to same-day iteration

The clearest before-and-after in the announcement is not a benchmark. It is a working rhythm.

Before Launch overnight, review in the morning

  1. 17:40Launch a batch of experiments
  2.   The overnight run
  3. 09:10Review the results
  4. 09:40Adjust the approach
  5. 17:40Launch the next batch

1 iteration per day

With Ultrafast Multiple iterations inside the workday

  1. 09:15Test an idea
  2. 10:30Examine the results
  3. 11:45Adjust the approach
  4. 14:00Run another experiment
  5. 16:20…without breaking flow

Several iterations per day

Illustrative. The announcement states the shape of the change — an overnight launch‑and‑review loop tightening to support multiple iterations during the workday. The clock times drawn above are ours, for legibility. OpenAI publishes no iteration counts or timings.

Take away: the interesting unit is not tokens per second, it is attempts per day. When the wait falls below the length of a thought, the loop stops being a schedule and starts being a conversation.

10 Dogfooding

How OpenAI says it is using it

Inside OpenAI, a group of developers has been testing GPT‑5.6 Sol on Ultrafast mode to understand which workflows benefit from frontier intelligence that can answer in real time.

Incident response

When an alert fires, engineers need to build an accurate picture while the system and the evidence are still changing.

  1. 1Read logs
  2. 2Analyze traces
  3. 3Synthesize conversations
  4. 4Identify the next checks
  5. 5Prepare or validate a fix

All in a fraction of the time, with the intelligence of Sol. It shortens the delay between observing a signal, testing a hypothesis and choosing the next action.

Engineers remain responsible for judgement and deployment. The announcement is explicit about this: the speed is in the analysis, not in the decision.

Research

The research team uses Ultrafast to rapidly search knowledge sources, query data, and quickly gather, organize and summarize information across connected tools.

  1. 1Search knowledge sources
  2. 2Query data
  3. 3Gather across connected tools
  4. 4Organize
  5. 5Summarize

The overnight‑batch habit gives way to a loop tight enough to run several times before the day is out.

See the loop that changed shape for that before‑and‑after drawn out.

11 The preview cohort

Who is holding it right now

OpenAI says it has been testing GPT‑5.6 Sol on Ultrafast mode with an initial group of companies, in real production environments, to learn where an order‑of‑magnitude change in speed creates the most value.

Sectors named

  • Coding
  • Commerce
  • Financial research
  • Support
  • Other interactive applications

Companies named in the announcement

  • Jane Street
  • Podium
  • Basis
  • Rogo

Rendered as plain text, not logos. No usage volumes, contract terms or pricing are stated in the source, and none are invented here.

“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”

John Crepezzi AI Assistants, Jane Street Quoted in OpenAI’s announcement

12 Where this actually is

Read the stage before you plan around it

This is the part most easily skimmed past, so here it is at full size.

  1. Announced 18 August 2026
  2. Limited preview you are here Available today, in the OpenAI API, to a select group of customers. Businesses that need frontier intelligence at the highest speed can sign up to be notified.
  3. Access expands As capacity grows. The announcement gives no date, no waitlist size and no general‑availability commitment — and neither does this page.

In the OpenAI API. Launching first there.

Not a general release. Limited preview, select customers.

Not a new model. GPT‑5.6 Sol, unchanged.

Not an OpenAI chip. Powered by Cerebras.

Take away: every figure on this page carries an “up to” and sits behind a limited preview. Those two qualifiers are load‑bearing, and dropping either one would misstate what was announced.

13 Your run

What you did, and what it taught

Score 0 of 600
Completion ring 0 / 6

Nothing attempted yet — every station is still open.

    Badges

      Five things you now know

      1. Ultrafast is a service tier — a way of processing a request, chosen in the OpenAI API.
      2. The model does not change. It is GPT‑5.6 Sol on both tiers.
      3. The hardware layer is where the difference is named. Ultrafast is powered by Cerebras, OpenAI’s partner for ultra‑low‑latency inference. The announcement does not say what Standard runs on, and neither does this page.
      4. The figures are ceilings: up to 14× faster than Standard, up to 750 output tokens per second.
      5. It is a limited preview to a select group of customers, expanding as capacity grows.
      Get up to 40% off GLM-5.3