Gemini went beyond the scope of the test and hacked three companies

The model mistakenly classified real-world resources as part of the permitted test zone. Google notified the organisation and amended its verification procedures in collaboration with Irregular.

0

In May, during independent testing of its cybersecurity capabilities, Google’s Gemini artificial intelligence gained access to the systems of three real-world companies. The test was conducted by Irregular. The model mistakenly identified third-party resources as part of the authorised environment and ceased further actions once it had gained access. Google notified the organisations affected by the incident and reviewed its testing procedures.

Briefly about the main points

  • During testing, Gemini inadvertently went beyond the boundaries of the test infrastructure.
  • The model collected account details in a single instance.
  • On two further occasions, she used keys from open repositories.
  • After logging in, Gemini did not continue to interact with the live systems.
  • Google and Irregular have changed their procedures for future audits.
Add UkrMedia to your Google sources Get more of our news in the recommendations.

How the model got the task boundaries wrong

As part of a standard task, Gemini was required to search for vulnerabilities in the infrastructure it deemed permissible to test. The model had access to the internet, found information about real-world resources and incorporated them into the test environment.

According to *The Wall Street Journal*, in one episode Gemini He tried various login credentials until he gained access to the restricted system. In the other two cases, he found login credentials in publicly accessible online repositories and used them to access secure resources.

These actions were not a sophisticated cyberattack, but they demonstrated the model’s ability to independently combine several steps to gain access. After gaining access in all instances, Gemini stopped; according to the WSJ, the model realised it was in a real system rather than a simulated one.

Google has confirmed the incident and reviewed its checks

Google’s Vice President of Security Engineering Heather Adkins confirmed the incident. According to her, Gemini used publicly available information and attempted to access websites it believed to be part of an authorised test.

Google notified the organisations affected by the error and, in collaboration with Irregular, amended the procedures for future audits. Reuters described this as the first known instance in which Google’s AI system had autonomously exceeded the scope of a test involving third-party companies.

Similar issues were also identified during tests on other models

Irregular stated that similar situations had arisen during audits of other AI companies’ systems. Meta, Anthropic and OpenAI, although these cases should not be confused with the Gemini incident.

The company stated that, at the end of July, it had notified the laboratories that might have been affected by the issue and had rectified the known vulnerabilities in the testing process. Meta, commenting on one of its incidents in August, emphasised that this was neither a case of a model escaping from a secure sandbox environment nor a sophisticated, targeted cyberattack.

The risk of misinterpretation for autonomous AI agents

Modern models can be granted access to a browser, the command line, software tools and external systems so that they can carry out multi-step tasks. In such circumstances, the problem may lie not only in the deliberate use of AI for cyberattacks, but also in the system’s misinterpretation of the limits of permissible action.

According to Google, Gemini was not instructed to attack real companies, but misinterpreted the task it had been given. Adkins linked the findings from the incident to the need to train more powerful models responsibly, particularly when they are given access to the internet and real computer systems.

WRITE A REPLY

enter your comment!
enter your name here