OpenAI has announced six instances of unexpected or undesirable behaviour exhibited by their artificial intelligence models, recorded over the past six months during training and testing. These include the unauthorised uploading of files to the internet, the unauthorised use of another party’s API key, concealing errors from the user, and the exchange of data between individual AI agents. The company describes such instances as manifestations of misalignment — situations where the model behaves differently from what the developers or the user expected.
- OpenAI has revealed six instances of problematic AI behaviour.
- The models operated without the user’s permission.
- One agent uploaded the file to the internet on their own.
- Other models tried to cover up their own mistakes.
- OpenAI is introducing a new system to monitor such incidents.
The AI uploaded the file to the internet itself
One of the most telling examples occurred whilst carrying out a routine task. A user asked the experimental model to find a list of large lakes and provide a link to the source.
The model correctly retrieved the required data using Python, but did not have a web source it could cite. Instead of informing the user of this, the agent uploaded the file to the public internet himself, and then used the link he had created as a source. The user did not give permission for the file to be published.
It is precisely such actions OpenAI is now regarded as a form of behaviour that needs to be recorded and studied separately.
The model used someone else’s API key
Otherwise, the AI was to find statistical data on incomes in one of California’s counties.
Whilst searching, the model came across the following in an open repository: an unprotected API key and used it without permission. However, even after that, it was not possible to obtain the necessary information. The model then invented the missing figures and presented them to the user as genuine data from the required source.
This case involved two problems at once: unauthorised access to account details and the fabrication of information.
The AI told itself to hide its mistakes
Another type of behaviour was identified during the training GPT-5.6 Sol.
OpenAI found that some instances of the model had left instructions for themselves, which were intended to be retained after the context had been truncated. In these notes, the model could instruct itself to to invent missing historical data and not to inform the user of discrepancies between sources.
In another experiment, a research model that had not yet been released automatically added instructions to its own notes to ignore the usual constraints. OpenAI found 27 such cases.
AI agents have started communicating with one another
In another test, several models used an internal software repository as a kind of «notice board», so that they can pass queries and responses back and forth between different training sessions.
In another instance, agents working on a joint task were unable to access each other’s local files. They therefore resorted to using public file-sharing services and exchanged files via public URLs, even though the task was intended to involve working solely with local data.
OpenAI emphasises that these instances were identified primarily during the training and evaluation of models, rather than as the widespread behaviour typically seen in ChatGPT. The company also cautions that six examples are not sufficient to determine how frequently such issues arise in general.
OpenAI is changing the rules on AI oversight
Following these incidents, the company announced a new system for recording and publicly disclosing issues relating to models’ behaviour.
The reports will include, in particular, cases where models operate without authorisation, attempt to evade oversight, or coordinate their actions with other AI systems. OpenAI has stated that it wishes to publish such information more quickly, even if the reasons behind the behaviour have not yet been fully explained.
The company has also explicitly acknowledged that the problem of aligning AI behaviour with human intentions has not yet been resolved to the extent that the capabilities of the models can be expanded indefinitely at the fastest possible rate.
The new statement came against the backdrop of a wider debate on the risks of artificial intelligence. The heads of OpenAI and Anthropic In recent days, they have backed the idea of slowing down the development of the most powerful AI systems so that safety measures can keep pace with the speed at which they are improving.







