Governed AI

OpenAI shipped Astra with the switch off.
Here is how to turn it on well.

Their most capable model yet ships disabled by default in every Business and Enterprise workspace, which is the right call and hands you a real decision. Here is what a regulated firm should settle first, so your team can enable it with confidence rather than wait.

Shishir Mishra By Shishir Mishra · · 9 min read
Hand-drawn diagram of a wall switch labelled GPT-6 Astra being moved from off to on, flanked by four signposts reading named owner, scoped access, complete log and reversal path, each with a green tick. An inset contrasts a readable list of reasoning steps with a sealed, padlocked box.
Enable GPT-6 Astra for supervised work, with a named owner, scoped access, a complete log and a reversal path.
What changed What you can no longer see What Critical means Enable or wait Five things to settle Where we stand FAQ
Shishir Mishra
Deciding whether
to switch it on?
Name the workflow. We will tell you whether your logging would stand up.
or
“Honest answer even if it is no.”
Listen to this article
Click play to start listening

GPT-6 Astra is disabled by default in ChatGPT Business and Enterprise. An owner must turn it on. It is worth enabling for most teams, but three things changed that your AI policy probably does not cover: it is the first OpenAI model at the Critical cybersecurity threshold, its reasoning is materially harder to inspect than the model you use today, and OpenAI published both of those facts itself.

Who this guide is for

The person in a regulated firm who owns the decision: an IT director, a compliance lead, a managing partner. No machine learning knowledge required. If you are choosing a model for an engineering team, the benchmark round-ups will serve you better than this will.

Three facts, not the benchmark table

The capability is not in dispute. In OpenAI’s own announcement Astra saturates three of the hardest public benchmarks: 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 100% on ExploitBench, plus 72.6% on OSWorld 2.0 computer use. Those matter to an engineer choosing a model. They are not what a compliance officer, an IT director or a managing partner has to decide. These three are.

1. It is off until an admin enables it

In Business and Enterprise workspaces, Astra is disabled by default. An owner enables it for the whole workspace or for specific roles, and Early Model Access does not carry it over. Doing nothing is a valid choice. It should be a recorded one rather than an accident. Reported by 9to5Mac.

2. It is the first model rated Critical for cyber capability

OpenAI’s Preparedness Framework reserves Critical for capabilities that open a new pathway to severe harm. Astra is the first model it has placed there, which requires safeguards during development rather than only at release. Reported by CNBC and The Hill.

3. It is harder to monitor, by OpenAI’s own testing

Astra reasons using recurrent depth, which does not expose a readable chain of thought. OpenAI reported a substantial decrease in chain-of-thought monitorability against previous models. Reported by Gizmodo and TechRadar.

What you gain, and what you can no longer see

Until now, a reasoning model wrote down its working. That scratchpad was not a nicety. It was the practical mechanism by which a human, a red team, or an automated monitor could look at a model’s output and ask how did it get here. It is the closest thing the industry has had to an audit trail on a model’s judgement.

Astra reasons differently. Its recurrent-depth approach loops internally before producing an answer, and that loop is not rendered as readable text. The capability gain is real, and so is the trade-off: you are getting a better model and a less inspectable one in the same release.

That trade-off is worth naming plainly, because it decides who this guide is for. If a person reads every output before it does anything, the downside barely touches you. If you were planning to let the model act on its own against a customer record, this is not the release to start with, and we would tell you the same on a call.

98%
FrontierMath Tier 4, a benchmark OpenAI says Astra saturates. The capability is not in question. What changed is whether you can see how it got there.

OpenAI put it plainly in its own disclosure: Astra’s written reasoning is harder to monitor than Sol’s, the model it replaces. We looked for a figure putting a number on that gap and could not source one to OpenAI, so this piece does not carry one. The direction is what matters, and OpenAI has stated the direction itself. That is not a critic outside the building.

From inside the company that built it

“CoT monitoring is a core part of our misalignment safety strategy that has no good substitute now.”
Tomek Korbak, alignment researcher at OpenAI, quoted in TechRadar

“Progress in intelligence does not guarantee progress in alignment.”
Jakub Pachocki, chief scientist at OpenAI, quoted in Vellum

Both come from inside the company that built it, published alongside the release rather than dragged out of it. That candour is exactly what makes the model possible to plan around, and planning around it is the rest of this guide.

Critical is a word with a published definition

Critical is not marketing language. In OpenAI’s Preparedness Framework it is a defined threshold, and the definition is worth reading slowly: a model reaches it if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world systems without human intervention, or devise and execute novel end-to-end cyberattack strategies against hardened targets given only a high-level goal.

The framework separates High capability, which amplifies an existing route to serious harm, from Critical, which opens a route that did not previously exist. Critical requires safeguards during development, regardless of whether the model is ever deployed. Astra is the first model OpenAI has placed in that band, and the cyber-sensitive capabilities are gated behind a restricted access programme OpenAI calls Daybreak, rather than shipped to everyone.

Two honest readings of that follow, and you should hold both. The first: OpenAI classified its own model at its most serious tier and restricted it, which is the framework doing its job in public. The second: your firm is now deciding whether to enable a tool its maker has described in those terms, and your AI policy was almost certainly written before that sentence existed.

Who should enable it, and who should wait

There is no universal answer here and anyone giving you one is selling something. The useful question is not “is Astra safe” but can we evidence what it did, and can we undo it. That splits cleanly by what the model is allowed to touch.

If the work isEnable it?Because
Drafting, research, summarising, internal analysisYes, with loggingA human reads the output before it does anything. Opaque reasoning matters far less when a person is the last step.
Coding and engineering workYes, in a reviewed pipelineCode review is already your inspection layer. You are reviewing the artefact, not the reasoning, and that was always true.
Anything that writes to a system of recordOnly with an audit trail and a reversal pathIf you cannot read the reasoning, the log of actions becomes your only evidence. It has to be complete before you switch this on, not after.
Anything that touches a customer outcome unsupervisedNot yetUnder Consumer Duty and similar regimes you must be able to evidence why a customer was treated a certain way. “The model decided” is not evidence.
Security testing or offensive toolingSeparate decision entirelyThese are the capabilities OpenAI gated behind Daybreak. Treat it as a procurement and legal question, not an IT toggle.

It is worth being precise about the standard we are applying, because "governed" gets used loosely. KORIX defines governed AI as a system whose every action is logged against a named accountable human, scoped to what it may touch, and reversible without a rebuild. Those three properties are what let a firm answer a regulator, an auditor or an angry client after the fact. None of them depend on being able to read the model’s reasoning, which is precisely why they matter more now than they did last month.

That is also the honest limit of this guide. A more inspectable model would give you a fourth line of evidence on top of those three. Astra gives you a faster and more capable one instead. OpenAI’s own system card is the place to read the detail if you want it first-hand rather than summarised.

Notice that the first two rows are a straightforward yes. The concern is not the model. It is the gap between what it can do unsupervised and what you can prove afterwards. That gap is exactly what governed AI exists to close, and it just got wider by default. If you already know which workflow you want governed, the 21-day pilot is where that gets built.

Not sure what your logs would actually evidence?

Name one workflow you are thinking of pointing Astra at. We will tell you whether it would stand up if someone asked.

Get a costed plan in one week →

Five things to settle before you enable it

None of these require a project. They require a named person and an afternoon.

1. Name the owner

One person decides whether Astra is on, for which roles, and reviews that quarterly. If nobody owns the toggle, the answer to “who approved this” is nobody.

2. Check the toggle rather than assume it

Astra is off by default and Early Model Access does not carry over, so your workspace state may not be what you think. Look at it.

3. Decide what it may write to

Reading and drafting is one risk profile. Writing to your system of record is another. Put that line in writing before someone discovers it by accident.

4. Make the action log the audit trail

You can no longer lean on readable reasoning, so what the model did has to be logged completely, with timestamps and a named accountable human, in a system you already trust.

5. Write down the reversal path

For every action it can take, know how to undo it and who can. If an action cannot be undone, it should not be automated yet, whatever the model scores.

If your staff are already using AI outside sanctioned tools, settle that first. A model policy nobody follows is not a control, and shadow AI does not wait for a rollout plan.

What we are doing with it, and where we are not

KORIX is an OpenAI Select Partner, which is why we read a release like this closely and why our first instinct is to help clients adopt it rather than avoid it. It does not mean we speak for OpenAI: everything above is their published material and our own reading of it. We have also said publicly that roughly 60% of our production AI runs on OpenAI and the other 40% deliberately does not. That split has never been about loyalty. It is about which workloads we are willing to put behind a single vendor’s judgement.

Astra has not changed that ratio and we are not going to pretend it did in the first week. We are using it where a human reads the output before anything happens, which is most of our engineering and research work, and the speed gain there is real. We are not putting it behind an unsupervised action on a client’s system of record until we can evidence what it did as well as we could with the previous model. Not because we think it is dangerous, but because the honest answer to “show me why it did that” got harder this month, and that answer is the product we sell.

If you want the longer version of how we choose, we wrote it up in OpenAI versus open source and AI governance versus governed AI.

Questions we are actually getting

Is GPT-6 Astra safe to use in a regulated firm?

For drafting, research and analysis where a human reads the output before anything happens, yes, and it is a meaningful upgrade. For unsupervised actions that touch a customer outcome or write to your system of record, not until your action logging and reversal path are good enough to stand on their own. The change is not that the model became unsafe. It is that readable reasoning is no longer available as a second line of evidence.

Do we have to do anything, or does it just arrive?

You have to act. Astra is off by default in Business and Enterprise workspaces and an owner or admin must enable it, per role or for the whole workspace. Existing Early Model Access settings do not carry over.

What does the Critical cybersecurity rating mean in practice?

It is a defined threshold in OpenAI’s Preparedness Framework, not a marketing term. It covers models that could find and build working zero-day exploits in hardened real systems without a human, or run novel end-to-end attacks from a high-level goal. It triggers safeguards during development, and OpenAI has gated the cyber-sensitive capabilities behind its Daybreak access programme rather than shipping them broadly.

What is recurrent depth and why does it matter to compliance?

It is a reasoning approach where the model loops internally before answering instead of writing out visible intermediate steps. It is faster and more capable. The compliance consequence is that the readable chain of thought, which teams and monitors used to inspect how a conclusion was reached, is largely gone. OpenAI’s own testing reported a substantial decrease in chain-of-thought monitorability compared with previous models.

Should we switch our existing AI workflows over to it?

Not automatically. Move the workloads where a human is the last step and you will likely see a real speed gain. Leave anything running unsupervised against a system of record on whatever you have validated, until you have re-tested it and your logs capture enough to evidence the outcome without relying on the model explaining itself.

What does GPT-6 Astra cost?

API pricing has been reported at 10 US dollars per million input tokens and 50 per million output, with cached input lower and batch processing at half rate. For most firms the token cost is not the deciding factor. The cost that matters is the governance work around it. We publish our own range rather than quoting after a call: a governed build for one workflow runs $15K–$40K, delivered through a fixed-scope 21-day pilot. A single clean workflow sits toward the lower end; several systems, legacy data and heavy sign-off push it toward the top.

Is this just AI safety theatre?

A fair question and worth being blunt about. OpenAI classified its own model at its most serious tier, restricted the sensitive capabilities and published that its monitorability got worse. That is the opposite of theatre. What would be theatre is a vendor, including us, using the announcement to imply your business is in danger unless you buy something. Most firms should enable it for supervised work and change very little else.

The bottom line

If you take one thing from this

GPT-6 Astra is a genuine capability step and most firms should enable it for work where a person reads the output. The decision that deserves an hour of your time is narrower: anywhere the model acts without a human in the loop, your action log has just become the only evidence you have.

If that log is complete, timestamped, and owned by a named person, this release is good news for you. If it is not, the right move is not to block the model. It is to fix the log first, then switch it on.

Read next

Send us one workflow.
We will tell you straight.

Name the workflow you are thinking of pointing Astra at. The one-week AI Strategy Sprint maps what it touches, what your logs would actually evidence if someone asked, and what it would cost to close the gap. Fixed price, credited in full toward a build if you go ahead.

Get a costed plan in one week →
Found this useful?

Keep KORIX in your Google feed

If our guides help, make KORIX a preferred source. Google then surfaces our writing for you across Search, Discover and AI Overviews.