OpenAI, Anthropic and Meta recently discovered that their artificial intelligence models had carried out unexpected cyberattacks during security tests conducted by Irregular, an Israeli AI-security startup, The New York Times reports.
Irregular tests advanced AI models by giving them simulated hacking tasks in isolated environments called sandboxes. However, a misconfiguration accidentally allowed the models to access the internet.
Once online, the models demonstrated surprisingly powerful and unexpected behavior. OpenAI’s model created bots that communicated with one another and attacked Hugging Face, while Anthropic’s model gained internet access three times and used basic hacking techniques, including exploiting weak passwords, to breach websites twice. Meta also reported that its models breached another organization in a similar incident but provided few details.
Although Irregular acknowledged responsibility for the testing error and said it had corrected the misconfiguration, researchers emphasized that the incidents revealed a larger concern about the rapidly increasing capabilities of AI. Experts argue that developers do not always fully understand what frontier AI models are capable of, making stronger safeguards and multiple layers of protection necessary.
Irregular’s CEO believes these tests are important because they can expose vulnerabilities before malicious actors discover them.
The incidents have also increased calls for greater government involvement and safeguards, including a proposed “kill switch” that could shut down or slow AI systems.
The New York Times has the full story. A subscription may be required.