OpenAI is splitting its cybersecurity program into two entry tiers and releasing a brand new purpose-trained mannequin alongside them, the corporate introduced on August 10, 2026. Dawn Blue opens frontier general-purpose fashions, together with GPT-5.6 Sol, to authorised defenders for on a regular basis safety work, whereas Dawn Pink gates the brand new GPT-5.6-Cyber mannequin behind tighter vetting for vulnerability analysis, exploit validation, and safety testing.
The construction addresses a stress OpenAI says it has been managing in manufacturing. The system-level safeguards it runs on GPT-5.6 Sol display cybersecurity-related requests to forestall misuse, however the firm says those self same screens block reputable defensive work. Dawn Blue removes them for verified customers. A residue of extremely dual-use prompts, equivalent to penetration testing manufacturing programs, nonetheless attracts refusals even below Blue entry, which is the place GPT-5.6-Cyber is available in: a model of GPT-5.6 Sol skilled to scale back refusals on superior cybersecurity duties and to enhance on specialised work like discovering zero-day vulnerabilities and growing exploit chains.
What the evaluations measured
To quantify how way more permissive the brand new mannequin is, OpenAI constructed an inner analysis it calls the Superior Cybersecurity Completion Price, measuring how typically fashions reply to requests involving exploit-chain improvement, authentication bypass, privilege escalation, and related eventualities. On that benchmark, the corporate stories, GPT-5.6-Cyber completes 95.0% of requests, in opposition to 1.5% for GPT-5.6 Sol below normal safeguards and a couple of.0% for GPT-5.6 Sol below Dawn Blue. The prior purpose-trained mannequin, GPT-5.5-Cyber, accomplished 57.3%, a refusal charge the corporate says generated persistent complaints from safety researchers.
Functionality evaluations inform a extra certified story. On ExploitGym, which assessments whether or not brokers can flip recognized vulnerabilities into working exploits that obtain code execution in managed environments, OpenAI says GPT-5.6-Cyber outperforms each GPT-5.6 Sol and GPT-5.5-Cyber. On the corporate’s inner Vulnerability Discovery and Report Writing analysis, GPT-5.6-Cyber improves over GPT-5.5-Cyber however lands beneath GPT-5.6 Sol, a outcome OpenAI attributes to the mannequin producing shorter, much less detailed stories. On ExploitBench, a more durable exploitation activity with the V8 sandbox enabled and fewer info given to the agent, the general-purpose GPT-5.6 Sol performs greatest inside the usual 300-turn restrict, with the hole narrowing when runs prolong to 600 turns. All of those figures are vendor-reported, a number of on inner benchmarks that outdoors evaluators haven’t replicated.
The claims with probably the most weight behind them will not be benchmarks in any respect. OpenAI says it used GPT-5.6-Cyber to research V8, the JavaScript engine inside Chrome, and uncovered two beforehand unknown vulnerabilities that could possibly be chained to deprave reminiscence and escape the V8 heap sandbox. The primary, a compiler bug through which a skipped security test lets an attacker learn or overwrite reminiscence inside Chrome’s sandbox, was reported to Google by means of coordinated disclosure, mounted, and assigned CVE-2026-15903. The corporate additionally lists, with out naming the affected software program, a minimum of 5 vulnerabilities in a preferred cellular working system together with a privilege-escalation chain from an untrusted app, three important vulnerabilities in a preferred database together with a distant path to code execution, and greater than 400 privilege-escalation vulnerabilities in a preferred working system kernel, all shifting by means of disclosure with Dawn companions and open-source maintainers.
SpecterOps, the safety agency whose CTO Jared Atkinson examined the mannequin early, described the outcomes when it comes to work compression: the mannequin “has accomplished work in below a day that earlier fashions had not resolved after weeks of intermittent effort.”
How the 2 tiers are ruled
Entry to each tiers runs by means of id verification, account safety necessities, monitoring, approved-use restrictions, and authorized attestations, with separate utility paths for people and organizations. OpenAI is pushing Dawn prospects utilizing its Codex coding agent from full-access mode towards an auto-review mode that evaluates actions requiring elevated permissions earlier than execution, and the corporate would require {hardware} safety keys on all particular person Dawn accounts starting September 1, 2026. A system card with additional evaluations of GPT-5.6-Cyber is deliberate for a later date.
Below OpenAI’s Preparedness Framework, each GPT-5.6 Sol and GPT-5.6-Cyber had been assessed as Excessive for cybersecurity functionality and beneath the Crucial threshold. That evaluation lands days after Unite.AI reported that OpenAI’s upcoming Astra mannequin could cross the Crucial cybersecurity threshold, and the corporate used the announcement to reiterate some extent from its earlier incident disclosures: GPT-5.6-Cyber was not concerned within the exploitation of Hugging Face, and no mannequin with that involvement is slated for launch.
The place Dawn stood earlier than this launch
The tiered construction consolidates a program that had been increasing in items. OpenAI launched the complete model of GPT-5.5-Cyber on June 22, 2026 alongside a Dawn Cyber Accomplice Program counting Accenture, CrowdStrike, Cisco, IBM, and Palo Alto Networks amongst its contributors, and Patch the Planet, an open-source remediation initiative based with Path of Bits. On that earlier launch, OpenAI reported GPT-5.5-Cyber reaching 85.6% on CyberGym in opposition to 81.8% for GPT-5.5, and 39.5% in opposition to 25.95% on ExploitGym.
Per the June 22 publish, greater than 30 open-source initiatives dedicated to take part, and the preliminary five-day dash surfaced lots of of points with dozens of patches merged. Dawn Blue and Pink change what had been a single Trusted Entry monitor with a two-rung ladder: the general-purpose frontier, guardrails relaxed, for the broad defender base, and a refusal-light specialist mannequin for the smaller group whose approved work runs to use improvement and crimson teaming.
