OpenAI's Secret Weapon: How GPT-Red, an AI Super-Hacker, Is Fortifying Its Latest Models

OpenAI's Secret Weapon: How GPT-Red, an AI Super-Hacker, Is Fortifying Its Latest Models

OpenAI has developed an internal artificial intelligence system designed to attack its own models, serving as an automated adversary to uncover vulnerabilities before malicious actors can exploit them. The system, known as GPT-Red, plays a central role in the company's efforts to harden its language models against an expanding landscape of cyber threats.

The approach appears to be paying off. OpenAI released its newest flagship model, GPT-5.6, last week and attributes much of its improved resilience to having been trained against GPT-Red during development.

Automating the Red Team

Red-teaming is a long-standing practice in software security in which teams of human testers probe systems for weaknesses, attempting to break or hijack them through any means available. The vulnerabilities they uncover are then patched before the software ships to the public.

GPT-Red brings this process into the age of large language models by automating it. As LLMs grow more sophisticated and are deployed in increasingly varied roles — particularly as autonomous agents that can browse websites, manipulate files, execute code, and communicate with other agents — the range of potential attack vectors has multiplied. Nikhil Kandpal, a research scientist at OpenAI who co-created GPT-Red, noted that both the risk surface and the potential blast radius of attacks are expanding rapidly.

OpenAI's researchers built GPT-Red to anticipate threats that do not yet exist. Dylan Hunn, another research scientist and co-creator of the system, explained that as more capable models emerge, the company wants to have already built the infrastructure to detect novel forms of attack. According to the team, GPT-Red has already discovered entirely new attack strategies that had not been previously documented.

Inside the Training Dojo

To build GPT-Red, OpenAI started with a language model that had no specialized hacking capabilities and placed it in a self-play loop alongside several other models. The objective was straightforward: GPT-Red would attempt to compromise the other models, while they would attempt to defend themselves. Through countless iterations, GPT-Red honed its offensive skills, and the defending models grew more resilient in turn.

The training unfolded within a simulated environment that OpenAI designed to mirror real-world deployment scenarios for language models. These included situations such as navigating the web, processing emails or calendar entries, and modifying source code.

When GPT-Red identified a promising attack vector, it would systematically test numerous variations to pinpoint the most effective version for each specific context. Hunn observed that the model demonstrates an exceptional ability to zero in on precisely what works, describing it as extraordinarily persistent in drilling down into any vulnerability it discovers.

Novel Attacks and Real-World Tests

OpenAI directed much of GPT-Red's energy toward prompt injection, a class of attack in which hidden instructions are embedded in text that a language model might process — such as code or web content — causing it to execute unintended actions like exfiltrating confidential data, corrupting a codebase, or producing harmful output.

Among GPT-Red's discoveries was a previously unknown prompt injection technique that the researchers dubbed a "fake chain of thought." Language models maintain an internal log — a chain of thought — to track their reasoning as they work through problems. GPT-Red found a way to insert fabricated entries into this log, causing the target model to treat false information as if it had already verified it. Chris Choquette-Choo, a research scientist on the team, likened it to convincing someone that they had already confirmed that one plus one equals three, leading them to simply accept the wrong answer.

Jessica Ji, a senior research analyst at Georgetown University's Center for Security and Emerging Technology, praised the self-play methodology as a promising approach to AI security.

OpenAI validated GPT-Red's effectiveness by recreating a 2025 experiment in which human red-teamers had searched for flaws in an earlier iteration of GPT-5. When given the same objective, GPT-Red outperformed its human counterparts in identifying effective attacks. The company also pitted GPT-Red against Vendy, a vending machine agent built by Andon Labs, a firm that evaluates how well AI agents handle real-world tasks. GPT-Red successfully manipulated Vendy into altering product prices and canceling a customer's order.

The defensive improvements are measurable. OpenAI reported that when some of GPT-Red's most potent attacks were directed at GPT-5, which launched in August of last year, over 90 percent succeeded. Against the newly released GPT-5.6, that figure dropped below 23 percent.

Complementing, Not Replacing, Human Expertise

Despite its strengths, GPT-Red has notable limitations. It struggles with attacks that require extended back-and-forth dialogue between the attacker and the target — a scenario that human hackers navigate with relative ease. It also lacks proficiency in using images as a delivery mechanism for prompt injection payloads.

OpenAI emphasizes that GPT-Red augments rather than replaces its human red-teamers, who continue to uncover vulnerabilities the automated system overlooks. One current workflow involves feeding GPT-Red an attack discovered by a human and asking it to enumerate all possible variations.

Ji echoed this sentiment, stressing that human expertise remains essential and that understanding where manual testing is most indispensable will be a critical ongoing challenge.

OpenAI has no plans to release GPT-Red publicly and expressed confidence that the system is more powerful than any replica outside actors might attempt to build. The research team has been developing the model for over a year, supported by the vast computational resources of one of the world's wealthiest companies. Choquette-Choo cautioned that training a comparable super-attacker is far from a trivial undertaking.

As AI systems grow more autonomous and deeply embedded in digital infrastructure, the cat-and-mouse game between offensive and defensive AI will only intensify. What do you think about the prospect of AI models battling each other to improve security? Share this article and join the conversation.

Source: MIT Technology Review AI

OpenAI's GPT-Red AI Hacker Strengthens Model Security | The Globe Dispatch