Home Artificial intelligence Microsoft researchers crack AI guardrails with a single prompt
Artificial intelligence

Microsoft researchers crack AI guardrails with a single prompt

Share



  • Researchers were able to reward LLMs for harmful output via a ‘judge’ model
  • Multiple iterations can further erode built-in safety guardrails
  • They believe the issue is a lifecycle issue, not an LLM issue

Microsoft researchers have revealed that the safety guardrails used by LLMs could actually be more fragile than commonly assumed, following the use of a technique they’ve called GRP-Obliteration.

The researchers discovered that Group Relative Policy Optimization (GRPO), a technique typically used to improve safety, can also be used to degrade safety: “When we change what the model is rewarded for, the same technique can push it in the opposite direction.”





Source link

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Articles
Artificial intelligence

Roblox announces ‘Build,’ AI tools that let anyone create games

Roblox has a whopping 132 million daily active users. But, while Roblox...

Artificial intelligence

Roblox launches an AI-powered game-creation feature in its mobile app

Roblox announced Thursday a new feature called “Build,” allowing users to design...

Artificial intelligence

Steering AI for an Inclusive, Beneficial Future

SHANGHAI, July 16, 2026 /PRNewswire/ -- A report from Science and Technology...

Artificial intelligence

Roblox launches an AI-powered game creation feature in its mobile app

Roblox announced Thursday a new feature called “Build,” allowing users to design...