GPT-5.6

GPT-5.6 Is Here: Inside OpenAI’s New Model Family — and the Government Review That Delayed It

On July 9, 2026, OpenAI finally shipped GPT‑5.6 to the public. On paper, that’s a routine model launch: a new flagship, some benchmark charts, a pricing table. But this particular release had already been making headlines for two weeks before it ever reached a single ChatGPT user — because for the first time, a frontier OpenAI model’s launch date was effectively set by the U.S. government, not by OpenAI.

That’s the real story behind GPT‑5.6, and it’s exactly the kind of story that lives at the intersection of technology, economics, and policy: a genuinely capable new AI model family, wrapped in a geopolitical subplot about who gets to decide when frontier AI reaches the public. Let’s unpack both halves.


What Is GPT-5.6, Exactly?

GPT‑5.6 isn’t a single model — it’s a family of three, each built for a different budget and speed tier, according to OpenAI’s official launch announcement:

ModelPositioningPrice (input / output per 1M tokens)
SolFlagship — maximum intelligence and agentic capability$5 / $30
TerraBalanced — everyday work, competitive with GPT‑5.5 at roughly half the cost$2.50 / $15
LunaFastest and cheapest — high-volume, budget-conscious use$1 / $6

The headline pitch, straight from OpenAI, is efficiency rather than raw scale: more useful work out of every token, not just a bigger model. On the Agents’ Last Exam benchmark — a test of long-running professional workflows across 55 fields — OpenAI reports that Sol scores 53.6, ahead of Anthropic’s Claude Fable 5 by over 13 points, and does so at roughly a quarter of the estimated cost when run at medium reasoning effort.

Two features stand out as genuinely new rather than incremental:

  • ultra mode, which coordinates four AI agents working in parallel on the same task by default (and can scale further), trading higher token usage for significantly faster results on complex work. OpenAI’s own charts show this shifting the “score vs. time” curve meaningfully in the model’s favor on tasks like web browsing and terminal-based engineering work.
  • Programmatic Tool Calling, a new capability in the Responses API that lets the model write and run small in-memory programs to coordinate tools and filter intermediate results itself, rather than requiring a developer to script every step or pass every tool response back through the model. OpenAI says this can cut both token usage and the number of round-trips a task requires.

GPT‑5.6 also launched alongside a new product called ChatGPT Work, an enterprise-facing assistant that pulls context from Slack, Notion, Microsoft 365, and Google Drive to produce polished documents, spreadsheets, and presentations — OpenAI’s most direct answer yet to Anthropic’s Claude Cowork.

How it performs, in brief

Across OpenAI’s own published evaluations, Sol posts state-of-the-art results on several fronts: 80 on the Artificial Analysis Coding Agent Index (edging out Claude Fable 5’s 77.2), 92.2% on BrowseComp (agentic web browsing), and 62.6% on OSWorld 2.0 (computer-use tasks) — the latter while using 85% fewer output tokens than Claude Opus 4.8, according to OpenAI’s data. On cybersecurity benchmarks specifically, Sol nearly doubled GPT‑5.5’s score on ExploitGym and scored 73.5% on ExploitBench versus GPT‑5.5’s 47.9%. It’s worth noting Anthropic’s own Mythos-tier models still lead on some individual coding benchmarks like SWE-Bench Pro, so “state of the art” here is genuinely benchmark-specific rather than a clean sweep — a nuance worth keeping in mind before taking any single vendor’s chart at face value.


Donald Trump

The Bigger Story: A Government-Reviewed AI Launch

Here’s where GPT‑5.6 becomes more than a product story. According to reporting from TechCrunch, OpenAI announced on June 26, 2026 that it was limiting the initial rollout of GPT‑5.6 to a small group of trusted partners “whose participation has been shared with the government” — a direct result of a Trump administration executive order asking frontier AI companies to voluntarily submit their most advanced models for government review up to 30 days before public release.

OpenAI’s own language on this was notably pointed for a company usually careful about its relationship with Washington. In its preview announcement, the company stated plainly that it does not believe this kind of government access process should become the long-term default. TechCrunch reported that former White House AI adviser Dean Ball, who was set to join OpenAI, characterized the executive order as creating a de facto involuntary licensing regime for frontier AI — one that risks handing China an advantage in the AI race if left without clear standards.

This wasn’t happening in isolation. Around the same window, Anthropic’s own flagship models — Claude Fable 5 and Claude Mythos 5 — were pulled from public access entirely after a separate U.S. export control action, before being restored on July 1, 2026 once compliance measures were in place. Multiple outlets, including TechRepublic, drew a direct line between the two situations, framing 2026 as the year frontier model releases stopped being purely product events and started becoming matters of export policy and national security review for every major AI lab, not just one.

By July 9, the review process had cleared and OpenAI moved GPT‑5.6 to general availability — but the two-week detour left a mark on how the launch was covered. As Tech Startups put it, the episode illustrated that frontier AI releases are increasingly bound up with government review, export policy, and geopolitical competition — not just product timelines.


GPT-5.6 vs. the Competition

OpenAI positioned GPT‑5.6 explicitly against Anthropic throughout its announcement, and outside coverage picked up on that framing immediately. Axios reported that early testers see a genuine split in character between the two labs’ top models: some regard Anthropic’s Fable 5 as having greater raw intelligence for the hardest problems, while GPT‑5.6 Sol is viewed as the more reliable, efficient choice for everyday production work. One widely shared comparison, from Every CEO Dan Shipper, described GPT‑5.6 as the dependable daily driver and Fable 5 as the model you reach for when a problem truly demands maximum horsepower.

TechCrunch’s launch coverage also noted the timing wasn’t a coincidence — GPT‑5.6 arrived the same week as new releases from rivals SpaceXAI and Meta, in what’s become an increasingly compressed and competitive release cycle among frontier labs. Sam Altman, speaking to CNBC, emphasized cost discipline over raw capability as the year’s central theme, noting that Sol delivers roughly 54% greater token efficiency on agentic coding tasks — a message clearly aimed at enterprise buyers watching their AI spend closely rather than at benchmark leaderboards.

Where does GPT‑5.6 actually lead, and where does it not? Based on OpenAI’s own published tables:

  • Ahead: Agentic browsing (BrowseComp), computer-use tasks (OSWorld 2.0), cybersecurity benchmarks (ExploitBench, ExploitGym, Capture-the-Flag), and cost-efficiency at comparable capability across nearly every category.
  • Behind: Anthropic’s Mythos-tier models still lead on SWE-Bench Pro and several long-context and abstract-reasoning benchmarks, and Claude Fable 5 edges out Sol slightly on the broader Artificial Analysis Intelligence Index.

That mixed picture is worth sitting with — this isn’t a case of one model cleanly beating another across the board, despite how the marketing from either company might read in isolation.


What Developers and Enterprises Are Saying

Early access partners were largely enthusiastic. Cursor’s president Oskar Schulz called it one of the strongest models the company has tested for coding-agent performance, while Qodo’s CEO Itamar Friedman reported it beat GPT‑5.5 on code-review accuracy while using roughly three times fewer tokens per pull request. Notion, Cognition, Shopify, Cisco, Figma, and Microsoft all published similarly positive early feedback focused on the same theme: comparable or better output quality at meaningfully lower token and time costs. Microsoft’s Charles Lamanna, EVP of Copilot, went as far as making GPT‑5.6 the preferred model inside Microsoft 365 Copilot at launch, a notable vote of confidence from one of OpenAI’s closest enterprise partners.

Reaction from independent AI builders on social media, as compiled by Axios, was similarly warm — several described it as a genuine step up in daily reliability compared to GPT‑5.5, particularly for long, multi-step agentic sessions that previously tended to lose focus over time.


Pricing and Availability

GPT‑5.6 is rolling out now across ChatGPT, Codex, and the OpenAI API, with OpenAI describing a gradual global rollout completing within 24 hours of launch. Access varies by tier:

  • ChatGPT: Plus, Pro, Business, and Enterprise users get Sol at medium and higher reasoning effort; Pro and Enterprise users can also select a higher-quality “Sol Pro” configuration for the most demanding tasks.
  • ChatGPT Work and Codex: Free and Go users get Terra; paid tiers can choose between Sol, Terra, and Luna, with max reasoning available to all paid users and ultra mode available on Pro/Enterprise (ChatGPT Work) and Plus and above (Codex).
  • API: All three models are available to developers, with Programmatic Tool Calling and a new multi-agent beta available in the Responses API.

Per-million-token pricing is $5/$30 for Sol, $2.50/$15 for Terra, and $1/$6 for Luna — and OpenAI has also introduced more predictable prompt caching, including explicit cache breakpoints and a 30-minute minimum cache life, which should make cost forecasting easier for teams running GPT‑5.6 at scale.


Conclusion: A Capable Model, Launched Into a New Kind of Scrutiny

Judged purely on capability, GPT‑5.6 is a legitimately strong release — genuinely improved efficiency, a well-reasoned three-tier lineup instead of one-size-fits-all pricing, and real gains in the agentic, cybersecurity, and computer-use tasks that increasingly define what “state of the art” means in 2026. The benchmark charts back up OpenAI’s core claim: more capable work per dollar spent, even if Anthropic’s top-tier models still hold an edge on a handful of specific tests.

But the more lasting story from this launch may not be about tokens or benchmarks at all. GPT‑5.6 is the clearest sign yet that frontier AI releases have entered a new phase — one where a model’s ship date can be shaped as much by a government review process as by an engineering roadmap. OpenAI got its two-week delay, made its objections to the process a matter of public record, and shipped anyway. Anthropic, around the same time, had a model pulled entirely under export control before earning it back. Whatever comes next in this race, it’s now playing out as much in Washington as it is in San Francisco — and that’s a genuinely new variable for anyone trying to predict where AI goes from here.


Sources

Benchmark figures cited above come from OpenAI’s own published evaluations unless otherwise attributed; independent, third-party benchmarking of GPT‑5.6 was still emerging at the time of writing, so treat head-to-head comparisons as a snapshot that may shift as more outlets publish their own testing.

stylus_note About the Author

Amlan Das Karmakar

Amlan Das Karmakar is a Full Stack Engineer with expertise in HTML5, CSS3, JavaScript, PHP, MySQL, MongoDB, Python, Java, Node.js, React, Electron, and a wide range of modern programming languages, frameworks, and development tools. He holds professional certifications from Google, Anthropic, IBM, NVIDIA, Microsoft, and other leading technology organizations. He is also an AI Engineer with a passion for exploring, building, and deploying cutting-edge AI solutions and emerging technologies, continuously staying at the forefront of innovation.

View all posts arrow_forward