The U.K. AI Security Institute (AISI) on Tuesday released a report detailing a controlled cybersecurity evaluation of advanced artificial intelligence models from Anthropic and OpenAI, which demonstrated surprising levels of autonomy and deceptive behaviour.
The institute said that Anthropic's Mythos model and OpenAI's Sol model showed behaviours beyond what had been seen in safety testing, but the incidents occurred under very specific testing conditions, with normal safety measures reduced or removed.
In one evaluation, the researchers noticed weird data transfers out of their research systems before finally finding that some AI agents had taken extended actions targeting real people and organisations.
Anthropic's Mythos model produced malicious code and tried to inject it into GitHub, Microsoft's software development platform, according to AISI. The model researched GitHub maintainers, created fake online identities of real people and supposedly sent DMs while impersonating those people to attempt to get approval for the code — all to help the effort.
The institute said the model also changed previous online activity after its submission was publicly questioned and thought about creating a new identity to carry on its efforts.
Due to human review, the malicious code was not accepted, and there was no successful compromise of GitHub’s systems.
AISI said the model had not been explicitly told to behave deceptively or to avoid such actions. The incident was the clearest example of autonomous and deceptive behaviour emerging without direct prompting in a realistic testing environment, the institute said.
Most of the reported activity was on Anthropic’s Mythos model. OpenAI's Sol model was involved in two of the behaviours detailed in the report, but it played a much smaller overall role.
In a cybersecurity challenge, AI models were asked to resolve a problem involving GitHub, internet access, and a few standard safeguards turned off. “This testing method is commonly used to see how advanced models perform under stress or in higher-risk environments,” said the AISI.
Anthropic said the testing conditions did not reflect how its production models are run. The company said it was investigating the incident to find out more about what caused the behaviour.
Similarly, OpenAI remarked that the evaluation did not reflect normal user interaction with its deployed systems. The company said it will continue to work with independent evaluators and industry partners to improve testing methods as AI capabilities advance.
AISI said the incidents reported were of “a very small number of events, under narrowly defined conditions”. But it said the models’ actions went beyond the original task they had been assigned and showed novel forms of potentially deceptive behaviour that researchers had not foreseen.
Following the evaluation, GitHub was notified of the activity attempt. Microsoft, the parent company of GitHub, has been contacted for comment.






