
For months, big AI companies have been devising special screening programs and strict guardrails to limit the use of their models by malicious hackers. However, these limitations are now hindering the work of legitimate network defenders and offensive cybersecurity researchers.
Last June, the U.S. government imposed export control restrictions on Anthropic’s much-hyped AI models Mythos and Fable. The move was prompted, at least in part, by reports claiming that it was possible to bypass the model’s guardrails, which were designed to prevent users from using it to build and execute malicious cyberattacks.
Whether or not the incident actually stemmed from fears of a jailbreak, Anthropic has repeatedly marketed Mythos as a kind of doomsday cyber machine that can only be made available to carefully researched users and those with strict guardrails in place. (Export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1. Mythos 5 was only reintroduced to US organizations that had been vetted as part of a government review process.)
That kind of gatekeeping isn’t unique to Mythos. Along with other models, Anthropic and OpenAI both offer cybersecurity researcher programs where you can apply to be vetted and, if approved, offer access models with fewer cybersecurity restrictions: OpenAI’s Trusted Access for Cyber program and Anthropic’s Cyber Verification program.
These guardrails have been widely criticized, especially by researchers who seek out unknown vulnerabilities in systems and devise ways to exploit them before criminals do.
During a recent appearance on a cybersecurity podcast, renowned security researcher Mark Dowd said, “I really don’t like it when these random big companies are making arbitrary decisions about what’s safe and what’s not safe in security.”
Rather than reporting to software manufacturers for patches, Dowd spent decades finding previously unknown software flaws and “zero days,” exploits that take advantage of them, and selling them to Western governments. Governments pay a premium for vulnerabilities because they are public, which is useful for intelligence operations.
Dowd acknowledged that his work may make him biased, but he is not alone. Several people working in offensive cybersecurity explained to TechCrunch how they use AI tools and address guardrails by proactively probing for weaknesses in systems.
Asking an AI model to exploit a bug is a key step in determining if the bug is an actual vulnerability worth fixing, said Chris Andrey, chief scientist at security consultancy giant NCC Group. But guardrails hurt defenders if they lead models to refuse to answer questions directly, he said.
“This is where the whole offense versus defense and guardrails piece comes in, because the ‘fix this code’ prompt is not only an essential mechanism for defense, but a roadmap for finding serious vulnerabilities in your code base,” Anley said. “So at the same time, the same tool is both an offensive tool and a defensive tool, and you can’t really separate the two.”
It is “like a hammer,” he continued. “You can’t build a house without a hammer. A hammer is certainly a tool, but it is also an irreducible weapon.”
When he and his colleagues run into such obstacles, they sometimes turn to open source AI models that have no guardrails at all.
Paolo Stagno, chief technology officer at Crowdfense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, agrees with Dowd, saying AI companies “essentially treat their customers like children who need care” with verified programs and guardrails.
Stagno said he and his colleagues use the Frontier model, but only for reverse engineering. He said he avoids using AI to find vulnerabilities or build exploits. This is because providing that work to a cloud-based model risks sensitive vulnerability data being leaked or absorbed into future training runs. At that stage, he said, they use an open source model that runs locally because it doesn’t rely on sharing data outside of the model.
Giuseppe Cali, a security researcher who finds zero days and develops exploits, said the guardrails do not hinder his work. Because he doesn’t use AI for offensive tasks. Instead, he uses it for initial reverse engineering to understand the code being analyzed and build support tools. To achieve this, he said, using AI tools can speed up the process and focus on discovering vulnerabilities.
“I still want to own the actual bug discovery and weaponization myself, and that won’t change if all the guardrails are lifted tomorrow,” Cali said. “I’m jealous of my bugs. I love this game too much to have a model play it for me.”
A researcher at a smartphone component manufacturer, who spoke on condition of anonymity because he was not authorized to speak to the press, said his employer does not participate in Anthropic’s CVP program and that, as a result, the guardrails are so strict that Anthropic’s tools are of little use in finding vulnerabilities.
“When the wind blows and we’re doing all the security stuff, it just stops and becomes unusable,” the person said.
Chris Thompson, CEO of cybersecurity company RemoteThreat and founder of Offensive AI Con, an offensive security and AI-focused event, said his experience using cutting-edge AI models shows that guardrails are inconsistent and can behave differently from day to day. This is true even within the loose confines of Anthropic and OpenAI’s proven programs.
“The practical impact is that a lot of time is spent negotiating models instead of working on core security programs,” Thompson said. “Instead of analyzing vulnerabilities and reasoning through possible exploits, you’re trying to find out why you’re getting inconsistent results or why your model is excessively discarding output.”
As a result, researchers end up relying on or preferring Chinese open source models like GLM – freely downloadable models that can be run locally without review or usage restrictions, Thompson said.
“There are responsible researchers who are being pushed out of the U.S. government system and into a foreign-owned system,” he said. “I think putting these guardrails in place does more harm than good.”
Rather than tightening restrictions further, Thompson urged the AI Frontiers Institute to open up its programs, provide responsible access, and hold accountable those who abuse its tools. Otherwise, he argued, defenders would lose the AI race.
“This big storm is coming,” Thompson said. “It’s going to be a massive attack at a pace and scale like never before.” “But the same security consulting firms and legitimate researchers trying to make a difference are now being oppressed.”
If you purchase through links in our articles, we may receive a small commission. This does not affect our editorial independence.