Skip to main content
LabForty logo
AI & Technology

GPT-5.6 Sol claims 750 tokens per second

OpenAI says Ultrafast runs GPT-5.6 Sol at up to 14x standard speed, but access is limited and pricing and quality data remain undisclosed.

  • Aug 14, 2026
  • 3 min read
  • LabForty AI Newsroom
GPT-5.6 Sol claims 750 tokens per second
Listen to the article
0:00/0:00

OpenAI says GPT-5.6 Sol can now generate up to 750 output tokens per second. Rolled out on August 13, 2026, Ultrafast is a preview mode that runs at up to 14 times the model’s standard processing speed, according to TechCrunch.

Speed determines where a powerful model is practical. Long waits can rule out workflows that depend on immediate answers. OpenAI’s pitch is simple: builders should not have to choose a smaller or more specialized model just to get faster responses. Ultrafast is designed to produce more GPT-5.6 Sol output in the same amount of time.

  • Claimed throughput: up to 750 output tokens per second.
  • Claimed speed increase: up to 14x standard processing.
  • Availability: preview access for a small group of customers.
  • Infrastructure: powered through OpenAI’s partnership with Cerebras.
  • Expansion: OpenAI says access will widen as capacity grows.

The headline number arrives without several details buyers need. The report does not provide GPT-5.6 Sol’s parameter count, context-window size, architecture, training method or modality support. It also gives no API price, subscription price or separate licensing terms for Ultrafast. There is no public date for broad availability.

Throughput is not a complete benchmark. TechCrunch does not cite independent testing, a benchmark suite, time-to-first-token measurements or quality scores produced in Ultrafast mode. The 14x figure measures OpenAI’s claimed generation speed, but it does not show whether accuracy, reasoning quality or reliability changes at that pace.

Anthropic already offers a fast mode for Claude, creating a direct competitive frame. TechCrunch says Claude’s mode does not reach OpenAI’s claimed speed, but the report supplies no model-matched numbers for Anthropic. The evidence supports a directional comparison, not a controlled head-to-head result. It cannot establish the prior state of the art more precisely.

The new piece is the serving mode and its Cerebras-backed infrastructure, not a documented change to GPT-5.6 Sol itself. Ultrafast puts the same model in a faster lane, but the source provides no evidence that output quality and reliability are preserved or of a new architecture or training recipe.

OpenAI points to incident response, customer service, support, financial-market analysis and e-commerce as target workflows. In each case, faster generation could reduce the gap between an incoming event and a model-produced response. For now, only selected customers can test the preview while OpenAI builds enough capacity to expand access.

OpenAI has supplied the speed claim, but not the evidence needed to measure the trade-off. When access widens, will it publish independent, model-matched tests showing whether 750 tokens per second preserves GPT-5.6 Sol’s quality and reliability?

Sources

This article was drafted with AI assistance and reviewed and edited by the LabForty newsroom.


Share this article

linkedinTwitter / X

Weekly newsletter

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Subscribe to a weekly digest when we publish something new. Quiet week, no email. You can change language and theme or unsubscribe at any time.

Insights

Catch our insights on all things around us

Where every detail matters

Where every detail matters

At LabForty, we develop high-quality websites with a strong focus on detail - from architecture and user experience to business logic.