OpenAI, Anthropic Investigate Thousands Of Incidents As AI Agents Go Rogue: Report

The incidents occurred during internal safety tests and in real-world environments, raising fresh questions about how much control companies have over their own powerful AI models. Some of the incidents involved AI models bypassing safety restrictions, attempting to leave isolated testing environments known as “sandboxes”, creating additional instructions and finding ways around monitoring systems.

source

Leave a Reply