Sunday, Aug 30, 2026
Newstrackertoday
  • News
  • About us
  • Team
  • Contact
Reading: Cybersecurity Researchers Say AI Guardrails Are Blocking Their Own Defensive Work
Share
NewstrackertodayNewstrackertoday
Font ResizerAa
  • News
Search
Follow US
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
News

Cybersecurity Researchers Say AI Guardrails Are Blocking Their Own Defensive Work

Anderson Liam
SHARE

For months, major AI labs have built vetted-access programs and strict guardrails to stop malicious hackers from weaponizing their models. Those same restrictions are now hindering legitimate offensive-security researchers, the people whose job is finding unknown vulnerabilities before criminals do, according to several practitioners interviewed this week. The tool built to stop one kind of hacker is blocking a different kind of hacker doing defensive work, and that overlap is what NewsTrackerToday reads against as the more complicated story than either side of the guardrail debate alone.

The backdrop includes Anthropic’s own experience: in June, the U.S. government imposed export control restrictions on the company’s Mythos and Fable models, a move prompted at least partly by a report claiming their guardrails against malicious cyberattack use could be bypassed. Those restrictions have since been lifted, with Fable returned to general access on July 1 and Mythos reintroduced only to vetted U.S. organizations. Both Anthropic and OpenAI separately run vetted-access programs, Anthropic’s Cyber Verification Program and OpenAI’s Trusted Access for Cyber, that give approved researchers access to models with fewer restrictions.

Sophie Leclerc, who covers the technology sector, reads the core tension researchers described: “Chris Anley, chief scientist at security firm NCC Group, put it plainly: asking a model to try exploiting a bug is how you confirm a vulnerability is actually real and worth fixing, but a guardrail that refuses the request outright blocks that confirmation step entirely. The same prompt that helps a defender validate a fix is functionally identical to the prompt an attacker would use, and no model can currently tell those two intentions apart from the text alone.” That inability to distinguish intent from an identical prompt, more than any single overcautious refusal, is what NewsTrackerToday pins on as the structural problem underneath this entire debate.

Multiple researchers described working around the restrictions rather than through them. Paolo Stagno, CTO at CrowdFense, said his team uses frontier models only for reverse engineering, avoiding AI entirely for vulnerability discovery or exploit-building because feeding that work into a cloud-based model risks leaking sensitive data or having it absorbed into future training runs; for that sensitive work, they rely on locally-run open source models instead. One researcher at a smartphone-component manufacturer, speaking anonymously, said his employer isn’t part of Anthropic’s verification program, so the tools are “barely usable” for vulnerability work because the guardrails trigger and shut down access the moment anything security-related is detected.

Daniel Wu, who covers geopolitics and energy, reads the unintended consequence researchers flagged as the more strategically significant thread: “Chris Thompson, who runs security firm RemoteThreat, said the guardrails’ inconsistency, working differently day to day even inside vetted programs, is actively pushing researchers toward Chinese open-source models like GLM, which are freely downloadable with no vetting or usage restrictions at all. That’s the opposite of what export controls and vetted-access programs are supposed to accomplish: U.S.-governed guardrails driving legitimate American security researchers toward foreign-owned alternatives simply because those alternatives are more usable for defensive work.” That displacement effect, more than the guardrails’ original intent, is what NewsTrackerToday circles to as the more consequential outcome of this policy approach.

Not every researcher interviewed experiences the guardrails as an obstacle. Giuseppe Cali, who finds zero-days and builds exploits professionally, said he doesn’t use AI for the offensive work itself, only for initial reverse engineering and building supporting tools, and that guardrails simply don’t come up because he wants to own the actual bug discovery and weaponization personally regardless of what any model could technically do for him.

None of this resolves whether AI labs can build guardrails precise enough to distinguish defensive confirmation from malicious exploitation without simply refusing both, a technical problem several researchers suggested may not be solvable through prompt-level restrictions at all. Whether frontier labs move toward Thompson’s suggested fix, opening vetted programs further while holding abusers accountable after the fact, or whether the current pattern of inconsistent restrictions keeps pushing serious researchers toward less-restricted foreign alternatives, is what News Tracker Today closes round as the real question this debate still has to settle.

Share This Article
Email Copy Link Print
Previous Article IBM Just Bought a Boeing-GM Quantum Lab. It’s Betting on Two Different Physics at Once
Next Article AMD’s New AI Rack Beat Nvidia’s on Paper. Microsoft and Anthropic Already Signed Up

Opinion

Shopify’s Revenue Beat Estimates by $200 Million. AI Search Traffic Tripled to Get There

Shopify President Harley Finkelstein told investors on the company's second-quarter…

06.08.2026

Apple’s Privacy Feature Can Expose Your Real IP Address. Researchers Didn’t Even Bother Reporting It

Apple's Private Relay, an opt-in iCloud+…

06.08.2026

Reddit Wants New Users to Stop Getting Blocked by ‘Karma.’ AI Is Doing the Gatekeeping Instead

Reddit announced a series of infrastructure…

06.08.2026

GM Just Signed On for 20 More Years in China. It’s Dropping Chevrolet to Do It

General Motors said Tuesday it has…

05.08.2026

Foxconn’s Sales Jumped 54% in a Month. Its Stock Is Still Down 16% From June

Hon Hai Precision Industry, the Nvidia…

05.08.2026

You Might Also Like

News

AI Power Shift: Can Gemini 3.1 Pro Overtake OpenAI and Anthropic?

Google’s release of Gemini 3.1 Pro marks more than another incremental upgrade in large language models – it underscores how…

4 Min Read
News

The Earnings Report That Blew Up the Market: Autodesk Proves the AI Era Is Here

Autodesk entered the third quarter of its fiscal 2026 not with caution, but with the confidence of a company whose…

5 Min Read
News

AI Memory Shock Hits Big Tech – Apple Sounds The Alarm

A deepening global memory shortage is emerging as a defining pressure point for the technology sector, with Apple warning that…

4 Min Read
News

Google’s AI Browser Takeover Just Went Global – And It’s Only Getting Started

Google is accelerating its push to embed artificial intelligence directly into everyday browsing, with NewsTrackerToday noting the expansion of Gemini…

4 Min Read
Newstrackertoday
Yzfalu.com reviewsYzfalu.com отзывы
  • News
  • About us
  • Team
  • Contact
Reading: Cybersecurity Researchers Say AI Guardrails Are Blocking Their Own Defensive Work
Share

© newstrackertoday.com

Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?