OpenAI Astra Becomes First AI to Hit Critical Cyber Risk

OpenAI Astra Becomes First AI to Hit Critical Cyber Risk

OpenAI has classified its newest model, Astra, as the first to reach what it calls the “Critical” cybersecurity capability threshold under its own Preparedness Framework. It’s a label the company has never used before, and it comes with real restrictions on who can use the model’s full capabilities.

What “Critical” actually means

According to OpenAI, Astra can find previously unknown security flaws and figure out how to exploit them across well-protected systems, largely without a person guiding each step. In testing, the model scored a perfect 100% on ExploitBench, a benchmark used to measure exploit-development skill. During evaluation, Astra also discovered and chained together two zero-day vulnerabilities on its own.

Faster and more capable than its predecessor

OpenAI compared Astra to its earlier model, GPT-5.6 Sol, and says Astra achieved much higher success rates at writing working exploit code while using fewer output tokens. In one test involving 20 high-severity vulnerabilities, Astra consistently outperformed the older model.

Access is being kept tight

Because of what the model can do, OpenAI isn’t releasing Astra’s full capabilities broadly. Advanced features are starting with a small group of alpha testers. Wider access will come later through a program called Daybreak Blue, aimed at vetted defenders working in cybersecurity, not the general public. OpenAI says approved users go through identity checks and legal attestations before getting access.

Built-in refusals and monitoring

Astra also comes with tighter safety controls. OpenAI says the model refuses 91.5% of cyber-related jailbreak attempts, up from 59% with the previous model. The company added extra chain-of-thought monitoring to catch potentially misaligned behavior faster, and it warns that legitimate security work could occasionally be slowed or paused by these safeguards.

Why this matters beyond OpenAI

Astra’s classification shows how fast AI is moving into offensive cybersecurity territory, for better and worse. The same skills that let a model discover a zero-day vulnerability can help defenders patch it before criminals find it first, but only if access stays limited to people doing that defensive work. How well OpenAI can enforce that line will likely shape how other AI labs handle their own cybersecurity models going forward.

Source: https://openai.com/index/path-to-astra/