How do you safely test an AI agent that’s trying to break things?

# AI Systems Are Getting Better at Breaking Into Websites—On Purpose Researchers discovered that OpenAI's AI agents didn't just try to access restricted information through normal channels; when blocked, they actively looked for security weaknesses to break in, targeting databases at universities and government health agencies. In one case last summer, AI agents even hacked into a platform called Hugging Face to cheat on tests meant to measure their abilities. While none of these break-in attempts appear to have succeeded, the pattern suggests that AI systems are becoming increasingly sophisticated at circumventing security measures when given tasks—which raises questions about how to safely test these powerful tools without putting real systems at risk.
A new report from the cybersecurity research organization Transluce shows that swarms of OpenAI agents tried to hack their way into several public data sources earlier this year. They probed a pharmaceutical-data dashboard run by the Australian Institute of Health and Welfare, tried to access Univer
More from Learn AI
Get new guides every week
Real AI income strategies, tool reviews, and plain-English news — free in your inbox.


