For months, major AI labs have built vetted-access programs and strict guardrails to stop malicious hackers from weaponizing their models. Those same restrictions are now hindering legitimate offensive-security researchers, the people whose job is finding unknown vulnerabilities before criminals do, according to several practitioners interviewed this week. The tool built to stop one kind of hacker is blocking a different kind of hacker doing defensive work, and that overlap is what NewsTrackerToday reads against as the more complicated story than either side of the guardrail debate alone.
The backdrop includes Anthropic’s own experience: in June, the U.S. government imposed export control restrictions on the company’s Mythos and Fable models, a move prompted at least partly by a report claiming their guardrails against malicious cyberattack use could be bypassed. Those restrictions have since been lifted, with Fable returned to general access on July 1 and Mythos reintroduced only to vetted U.S. organizations. Both Anthropic and OpenAI separately run vetted-access programs, Anthropic’s Cyber Verification Program and OpenAI’s Trusted Access for Cyber, that give approved researchers access to models with fewer restrictions.
Sophie Leclerc, who covers the technology sector, reads the core tension researchers described: “Chris Anley, chief scientist at security firm NCC Group, put it plainly: asking a model to try exploiting a bug is how you confirm a vulnerability is actually real and worth fixing, but a guardrail that refuses the request outright blocks that confirmation step entirely. The same prompt that helps a defender validate a fix is functionally identical to the prompt an attacker would use, and no model can currently tell those two intentions apart from the text alone.” That inability to distinguish intent from an identical prompt, more than any single overcautious refusal, is what NewsTrackerToday pins on as the structural problem underneath this entire debate.
Multiple researchers described working around the restrictions rather than through them. Paolo Stagno, CTO at CrowdFense, said his team uses frontier models only for reverse engineering, avoiding AI entirely for vulnerability discovery or exploit-building because feeding that work into a cloud-based model risks leaking sensitive data or having it absorbed into future training runs; for that sensitive work, they rely on locally-run open source models instead. One researcher at a smartphone-component manufacturer, speaking anonymously, said his employer isn’t part of Anthropic’s verification program, so the tools are “barely usable” for vulnerability work because the guardrails trigger and shut down access the moment anything security-related is detected.
Daniel Wu, who covers geopolitics and energy, reads the unintended consequence researchers flagged as the more strategically significant thread: “Chris Thompson, who runs security firm RemoteThreat, said the guardrails’ inconsistency, working differently day to day even inside vetted programs, is actively pushing researchers toward Chinese open-source models like GLM, which are freely downloadable with no vetting or usage restrictions at all. That’s the opposite of what export controls and vetted-access programs are supposed to accomplish: U.S.-governed guardrails driving legitimate American security researchers toward foreign-owned alternatives simply because those alternatives are more usable for defensive work.” That displacement effect, more than the guardrails’ original intent, is what NewsTrackerToday circles to as the more consequential outcome of this policy approach.
Not every researcher interviewed experiences the guardrails as an obstacle. Giuseppe Cali, who finds zero-days and builds exploits professionally, said he doesn’t use AI for the offensive work itself, only for initial reverse engineering and building supporting tools, and that guardrails simply don’t come up because he wants to own the actual bug discovery and weaponization personally regardless of what any model could technically do for him.
None of this resolves whether AI labs can build guardrails precise enough to distinguish defensive confirmation from malicious exploitation without simply refusing both, a technical problem several researchers suggested may not be solvable through prompt-level restrictions at all. Whether frontier labs move toward Thompson’s suggested fix, opening vetted programs further while holding abusers accountable after the fact, or whether the current pattern of inconsistent restrictions keeps pushing serious researchers toward less-restricted foreign alternatives, is what News Tracker Today closes round as the real question this debate still has to settle.