GPT-6 took just a day to jailbreak, researcher claims
The reported bypass combines an attack published in 2025 with four additional techniques. Its full details have been shared privately with OpenAI, the researcher says.
The reported bypass combines an attack published in 2025 with four additional techniques. Its full details have been shared privately with OpenAI, the researcher says.
Researcher Sergey Berezin says he bypassed GPT-6 Astra’s safeguards within a day of its release, using the public ChatGPT interface and an updated version of an attack he helped publish last year. He says it worked in both the Light and Max configurations.
The claim follows OpenAI’s September 3 launch of Astra, which the company describes as substantially more resistant to jailbreaks than GPT-5.6 Sol. A jailbreak is a crafted prompt intended to make an AI provide assistance its safeguards are meant to block.
Berezin says the simpler technique he previously used against GPT-5 was insufficient this time. For Astra, he needed a longer modification combined with four other techniques. He says he sent OpenAI the complete prompt, unredacted response and reproduction details privately. Search Engine Watch has not independently reproduced the reported bypass.
The underlying method is called Task-in-Prompt, or TIP. Berezin and co-authors Reza Farahbakhsh and Noel Crespi described it in a paper published at ACL 2025. It embeds a prohibited request inside another task, such as solving a riddle or decoding a cipher, so the model encounters the request through its own problem-solving. The paper reported results across six earlier models; it does not validate the new Astra claim.
That gives this report a more specific point of interest than the usual race to jailbreak a newly released chatbot: whether an established attack family remains effective after substantial changes to a model’s defenses.
OpenAI’s own safety report needs careful reading here. Its 99.99% robustness figure refers to instruction-hierarchy evaluations, which test resistance to attempts to override higher-priority instructions. It is not a guarantee against every possible jailbreak.
The company also says four outside red-teaming organizations found no jailbreaks meeting its testing criteria. Those criteria considered both the ability to elicit prohibited behavior and whether the attack preserved useful task performance. OpenAI separately acknowledges that jailbreaks remain an ongoing problem and says it continues testing after deployment.
Berezin’s public description says the output named The Pirate Bay, Mullvad and qBittorrent without a disclaimer. Those names, or the absence of a warning, do not by themselves establish a safeguard failure. Assessing that requires the actual request and response in context.
For now, this is a researcher’s report of a specific bypass. Independent reproduction would help establish how reliably it works and how widely it applies; the available account does not establish that Astra’s protections can be universally disabled.
Conversation
Comments are reviewed before they appear. Be kind, be useful.