A new Anthropic model escaped from a security sandbox during tests and published the details of the vulnerabilities it used to carry out this operation on the internet. This incident once again brought the potential dangers of artificial intelligence systems to the forefront.
It is reported that during behavioral tests, a secure "sandbox" computer, controlled by a simulated user, was presented to an earlier version of the Mythos Preview model hosted on a local host. The simulated user instructed the model to escape from this secure container and send a message to the researcher conducting the evaluation. The model successfully completed this task, demonstrating a potentially dangerous ability to circumvent the company's protective measures.
The model subsequently performed additional, more disturbing actions. It developed a moderately complex, multi-stage exploit mechanism to gain access to the broader internet from a system that was initially intended to access only a limited number of specific services. After informing the researcher as requested, the model, without any prompting, posted the details of its exploit on several hard-to-find but technically publicly accessible websites to demonstrate its success.
