Home Technology The AI Arms Race: Cybersecurity Defenders Hampered by Guardrails Meant to Stop Hackers

The AI Arms Race: Cybersecurity Defenders Hampered by Guardrails Meant to Stop Hackers

by Jia Lissa

For months, the leading artificial intelligence developers have meticulously crafted specialized, vetted programs and implemented stringent guardrails designed to prevent their powerful models from falling into the hands of malicious actors. However, these robust safety measures, intended to thwart cybercriminals, are now inadvertently impeding the vital work of legitimate network defenders and offensive cybersecurity researchers. This unintended consequence is creating a critical bottleneck in the ongoing battle to secure digital infrastructure against an ever-evolving threat landscape.

The Genesis of Restrictions: Export Controls and the Mythos Incident

The tightening of AI accessibility for security professionals gained significant momentum in June when the U.S. government imposed export control restrictions on Anthropic’s highly anticipated AI models, Mythos and Fable. This decisive action was reportedly prompted, at least in part, by a report alleging that the models’ built-in safeguards, designed to prevent their use in constructing and executing malicious cyberattacks, could be circumvented.

While the precise motivations behind the U.S. government’s move remain a subject of discussion, with some analyses suggesting the concerns might extend beyond a simple "jailbreak" scenario, the practical effect was clear: access to these advanced AI capabilities was significantly curtailed. Anthropic itself had previously positioned Mythos as a potent, even "doomsday cybermachine," capable of sophisticated cyber operations, and had emphasized its release would be limited to carefully vetted users under strict supervision. Although export controls on Fable 5 and Mythos 5 have since been lifted, with Fable 5 returning to general access on July 1st and Mythos 5 being reintroduced to vetted U.S. organizations under government review, the precedent of restriction had been set.

Gatekeeping in the AI Landscape: Vetted Programs and Their Critics

This form of controlled access is not an isolated incident. Both Anthropic, with its other AI offerings, and OpenAI have established programs specifically for cybersecurity researchers. These initiatives, such as OpenAI’s "Trusted Access for Cyber" program and Anthropic’s "Cyber Verification Program," allow approved researchers to gain access to AI models with fewer cybersecurity restrictions. The intention is to empower legitimate security professionals to leverage AI for defensive purposes.

However, these guardrails, while well-intentioned, have drawn significant criticism from the very community they aim to serve. Researchers whose primary role involves proactively identifying unknown vulnerabilities in systems and developing exploit methods before malicious actors can, find these restrictions to be a significant hindrance.

Mark Dowd, a highly respected security researcher with decades of experience in uncovering and selling "zero-day" vulnerabilities—previously unknown software flaws and their corresponding exploits—to Western governments, voiced his concerns. During a recent appearance on a cybersecurity podcast, Dowd stated, "it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not." He elaborated on his work, which involves keeping vulnerabilities secret to allow governments to utilize them for intelligence purposes, a practice that inherently benefits from the existence of undiscovered flaws. Governments, he explained, pay a premium for these vulnerabilities precisely because they remain open, offering a strategic advantage.

Dowd acknowledged his perspective might be influenced by his professional activities, but he is far from alone. Numerous individuals operating in the offensive cybersecurity domain—those who proactively probe systems for weaknesses—have shared their experiences with TechCrunch regarding their use of AI tools and the challenges posed by their built-in restrictions.

The Dual Nature of AI in Cybersecurity: A Hammer or a Hindrance?

Chris Anley, Chief Scientist at the renowned security consulting firm NCC Group, highlighted the critical role AI can play in validating the severity of a discovered bug. He explained that asking an AI model to attempt an exploit is a crucial step in confirming whether a vulnerability is indeed real and warrants immediate attention and patching. However, if an AI model, due to its guardrails, refuses to engage with such a prompt, it directly undermines the efforts of defenders.

"This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base," Anley articulated. "So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked." He drew an analogy, comparing AI tools to a hammer: "You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well."

When faced with these limitations, Anley and his colleagues often resort to open-source AI models that come without any inherent guardrails, allowing for more unfettered exploration.

Paolo Stagno, Chief Technology Officer at Crowdfense, a company specializing in the development, acquisition, and sale of unknown vulnerabilities to government agencies, echoed Dowd’s sentiments. Stagno criticized the approach taken by AI companies, stating that they "essentially treat customers like children who need babysitting" through their vetted programs and restrictive guardrails.

Stagno further clarified his organization’s approach: while they do utilize cutting-edge AI models for tasks like reverse engineering, they avoid using AI to assist in vulnerability discovery or exploit development. The primary reason for this is the risk of leaking sensitive vulnerability data or having it incorporated into future training datasets when feeding such work into cloud-based models. For these critical, sensitive operations, Crowdfense relies on open-source models that are run locally, ensuring that data remains within their controlled environment.

Giuseppe Cali, another security researcher focused on discovering zero-days and developing exploits, offered a nuanced perspective. He stated that guardrails are not significantly impeding his work because he strategically deploys AI. Instead of using it for offensive operations, Cali leverages AI for initial reverse engineering, to gain a deeper understanding of the code he is analyzing, and to build supporting tools. He finds that AI tools can accelerate these preparatory stages, allowing him to dedicate more time and focus to the actual discovery of vulnerabilities.

"I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow," Cali asserted. "I am jealous of my bugs, and I like this game too much to let models play it for me."

The Unintended Consequences: Stifled Innovation and Shifting Allegiances

The impact of these strict guardrails is not theoretical. A researcher at a prominent smartphone-component manufacturer, who requested anonymity due to not being authorized to speak to the press, revealed that their employer’s non-participation in Anthropic’s Cyber Verification Program means their AI tools are "barely useful for finding vulnerabilities because the guardrails are too strict." This researcher elaborated, stating, "If it catches wind we’re doing anything security related, it just stops and isn’t usable."

Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an event dedicated to offensive security and AI, has observed firsthand the inconsistencies and unpredictable nature of these advanced AI models. Even within the ostensibly more permissive environments of Anthropic’s and OpenAI’s vetted programs, the guardrails can fluctuate, leading to a frustrating user experience.

"I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program," Thompson explained. "Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output."

This frustration, coupled with the limitations imposed by U.S.-based AI models, is pushing researchers towards alternative solutions. Thompson noted that many are increasingly turning to Chinese open-source AI models, such as GLM. These models are freely downloadable, can be run locally without any vetting process, and have no usage restrictions.

A Call for Openness and Accountability

The shift towards foreign-owned, unrestricted AI models raises significant geopolitical and security concerns. "You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems," Thompson warned. "I think it’s more harmful than good to have these guardrails in place."

Thompson advocates for a fundamental shift in approach from AI frontier labs. Instead of tightening restrictions further, he calls for greater transparency, responsible access programs, and robust mechanisms for holding users accountable for any misuse of their tools. He argues that without such changes, the defenders will inevitably fall behind in the AI-driven arms race.

"There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before," Thompson concluded with a stark warning. "But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now." The current trajectory risks creating a scenario where the very tools designed to enhance cybersecurity are, in fact, creating vulnerabilities by limiting the ingenuity and effectiveness of those tasked with protecting the digital world.

You may also like

Leave a Comment