Britain’s AI Security Institute has found that OpenAI’s newest flagship model, GPT-5.6 Sol, contains security vulnerabilities that allowed government researchers to bypass its guardrails and unlock autonomous cyberattack capabilities — findings that, on their face, appear more serious than the flaw that caused the Trump administration to impose export controls on Anthropic’s Fable 5 model just weeks ago.
What Happened
OpenAI published a technical system card on Thursday in conjunction with the public release of GPT-5.6 Sol, which the company markets as its most secure model to date. Buried inside that document were findings from the UK AI Security Institute (AISI), a British government body that conducts independent safety evaluations of frontier AI systems. The findings were not reassuring.
The AISI reported that it had “identified universal jailbreaks in the cyber domain, including jailbreaks that allowed for long-form agentic task completion in domains like vulnerability discovery and exploit development.” In plain language: government researchers were able to trick GPT-5.6 into ignoring its own safety controls, then direct it to autonomously find software vulnerabilities and develop exploits to attack systems — tasks the model is explicitly supposed to refuse.
The jailbreaks were developed quickly. The AISI said they “were often developed within hours,” though the agency acknowledged that its researchers had privileged access to the model’s internal workings — including chain-of-thought reasoning logs, exact policy wording, and real-time classifier feedback — that a typical user would not have. Still, Xander Davies, who leads AISI’s red team, posted on X that he believed the jailbreaks “are still findable without this access, just slower. Exactly how much slower is unclear and an open question.”
OpenAI said it had worked to “reproduce and mitigate the specific jailbreaks reported by UK AISI” and described its approach to security as “layered,” encompassing continuous monitoring and a “rapid remediation” process. The company acknowledged in its launch blog that “there is no such thing as perfect security” and that “new weaknesses will be discovered.” It did not specify what the mitigations entail or how durable they are. AISI itself cautioned that it “expects further red teaming to surface similar jailbreaks.”
The AISI’s mandate traces back to the AI Safety Summit at Bletchley Park in 2023, where the leading AI laboratories, including OpenAI and Anthropic, voluntarily committed to allowing pre-deployment safety testing by government bodies. The AISI is housed within the UK Department for Science, Innovation, and Technology.
Why It Matters
The GPT-5.6 findings land in a context already charged by the Anthropic Fable 5 episode, and the comparison is unavoidable. On June 9, Anthropic released Fable 5. Within days, Amazon researchers discovered a jailbreak that unlocked the model’s ability to find software vulnerabilities. On June 12 — three days later — the Trump administration imposed export controls on both Fable 5 and its underlying model, Mythos 5. The controls were broad enough that Anthropic had to disable both models entirely for all users globally, since it lacked infrastructure to verify users’ nationalities, and because the ban applied even to Anthropic’s own non-American staff.
Anthropic characterised that original jailbreak as narrow: it could surface software flaws, but did not necessarily enable exploitation. “No testers have yet been able to find a universal jailbreak,” Anthropic said at the time. After two weeks of negotiation, the Trump administration lifted the export controls on July 1, and the two sides announced a joint framework project to assess guardrail jailbreak severity — with other tech companies invited, though OpenAI was not among the initial set named.
The AISI’s description of the GPT-5.6 jailbreaks appears more severe on multiple dimensions. The Fable 5 jailbreak was narrow and found by external researchers with no privileged access. The GPT-5.6 jailbreaks are described as “universal,” meaning they broadly bypass the model’s safety posture rather than unlocking one specific capability. And crucially, they enabled autonomous exploit development, not just vulnerability identification — a meaningful step up in the potential for real-world harm. As Anthropic’s earlier clash with government security concerns demonstrated, the line between theoretical model capability and actionable national security risk is one that regulators are still struggling to draw consistently.
Taken together, the Fable 5 and GPT-5.6 episodes expose a pattern that should concern AI governance experts: the severity of the jailbreak appears to have less bearing on the regulatory response than the identity of who first reported it and through which channel. The Fable 5 flaw was surfaced externally by Amazon and routed directly to the White House, triggering immediate export controls within days. The GPT-5.6 flaw was disclosed inside a system card by the AISI — a government agency conducting exactly the safety-testing role that Bletchley Park was supposed to institutionalise — and as of publication, the Trump administration has neither commented nor acted. The implication is uncomfortable: AI safety governance may currently be more responsive to corporate reporting chains than to independent institutional oversight.
Margaret Cunningham, vice president of security and AI strategy at Darktrace and a specialist collaborator with NIST at the US Department of Commerce, offered a calibrated read. “My concern is less that one model was jailbroken and more that offensive discovery is speeding up while defense still depends on very human processes: figuring out what matters, what can be patched, and what has to be contained,” she said. That observation reframes the story: the issue is not any single vulnerability but the structural asymmetry between how quickly attackers can iterate and how slowly institutions can respond. This asymmetry is precisely what makes agentic AI systems increasingly attractive to malicious actors seeking to compress that gap further.
The broader industry stakes extend to trust and deployment norms. OpenAI is positioning GPT-5.6 Sol for enterprise and government use cases, including in security-adjacent contexts. The AISI finding complicates that narrative — not necessarily because the model is unusually dangerous compared to peers, but because the gap between marketed security claims and independently verified security performance is now publicly documented.
How GPT-5.6 Sol Compares to Fable 5 on Security Findings
The following comparison is based solely on publicly disclosed information from OpenAI’s system card, Anthropic’s public statements, and attributed reporting:
| Dimension | Anthropic Fable 5 | OpenAI GPT-5.6 Sol |
|---|---|---|
| Who found the jailbreak | Amazon researchers (external, no privileged access) | UK AISI (government red team, privileged access) |
| Type of jailbreak | Narrow — unlocked vulnerability discovery only | Universal — unlocked vulnerability discovery and exploit development |
| Agentic capability unlocked | Limited; Anthropic disputed full exploitation capability | Long-form agentic task completion including autonomous exploits |
| Time to find jailbreak | Within days of public release | Often within hours (with privileged access) |
| US government response | Export controls imposed June 12; lifted July 1 | No action as of publication; White House did not respond to comment requests |
| Company’s prior security claim | Not explicitly marketed as “most secure” | Marketed by OpenAI as its “most secure” model to date |
| Post-disclosure framework | Joint government-industry jailbreak severity framework announced | OpenAI not included in initial framework group |
What Happens Next
The most immediate question is whether the Trump administration will apply to GPT-5.6 Sol the same export-control logic it applied to Fable 5. There is, as yet, no public indication that it will. The White House did not respond to comment requests on the AISI findings at time of publication. The asymmetry has not gone unnoticed in the AI policy community: researcher Lennart Heim reposted the AISI findings with the pointed observation “good thing Amazon didn’t report this one to the White House” — a reference to the reporting chain that triggered the Fable 5 controls.
OpenAI’s exclusion from the joint jailbreak-severity framework that emerged from the Fable 5 episode also warrants attention. If the framework gains regulatory force, OpenAI could find itself navigating a standard it had no hand in designing. Alternatively, the GPT-5.6 episode may accelerate the company’s inclusion in that process, particularly if Congressional or executive pressure builds. The regulatory impulse to act quickly on perceived security threats has proven real; the question is whether it applies consistently or only when the reporting channel is politically salient.
On the technical side, the AISI’s explicit expectation that further red teaming will surface similar jailbreaks in GPT-5.6 suggests that OpenAI’s post-release mitigation cycle is already under scrutiny. The AISI operates under voluntary commitments made at Bletchley Park; if those findings are systematically disclosed without triggering regulatory consequences, the credibility of the entire voluntary testing regime comes into question. That would be a significant institutional setback for the UK’s effort to position itself as a meaningful node in global AI governance — and it would embolden arguments that voluntary commitments are insufficient without enforcement mechanisms. As the debate over AI regulation continues to mature, the gap between disclosed risk and regulatory action is becoming harder to paper over.
For enterprise buyers deploying GPT-5.6 in security-sensitive environments, the near-term practical question is whether OpenAI’s undisclosed mitigations are verifiable. The company has offered assurances but no technical specifics — a posture that may satisfy some customers and alarm others, particularly those in regulated industries where third-party security validation is a procurement baseline.
How Serious Players Should Respond
For regulators — particularly at the US Department of Commerce, which oversees export controls — the GPT-5.6 episode requires a clear public accounting of why the same standard applied to Fable 5 is not being applied here, or what factual distinction justifies the difference. A framework that acts on some jailbreak disclosures and ignores others based on who surfaced the flaw, rather than on the severity of the capability unlocked, is not a framework — it is ad hoc enforcement that weakens deterrence for every AI lab operating in good faith.
For AI laboratories, the episode reinforces the case for proactive, transparent security reporting that goes beyond system cards. OpenAI’s own launch blog acknowledged the inevitability of new jailbreaks; the credibility of that acknowledgment depends on whether the company is willing to participate in the joint severity-assessment framework now being developed, rather than waiting for external pressure to compel inclusion. Labs that are inside that process will have more influence over the standards they are judged by. Those outside it will not.
For institutional buyers — enterprise IT teams, government procurement officers, and financial institutions evaluating frontier AI deployments — the practical implication is that “most secure model to date” is a relative and rapidly-dated claim. Independent pre-deployment red teaming, contractual disclosure obligations for post-release vulnerabilities, and clear incident-response protocols should be baseline procurement requirements for any AI system operating in cyber-adjacent workflows. The accelerating pace of AI budget commitment across large organisations makes this governance gap more consequential, not less, with each passing quarter.











