OpenAI has introduced a new framework for tracking and publicly reporting cases where its AI models behave in ways that differ from their intended instructions.
Along with the framework, the company disclosed six previously unreported incidents involving unexpected behaviour observed during model training and evaluation.
The company said its earlier disclosures were handled on an ad hoc basis, with incidents sometimes grouped together or included in system cards when new models were released. The new process is intended to speed up reporting, including in cases where the company has not yet fully understood or addressed the behaviour.
Under the framework, OpenAI employees can flag potential misalignment incidents. Safety and alignment teams will then review the cases and determine whether they meet the criteria for public disclosure. The company plans to work with other AI developers, researchers, industry standards organisations and regulators to develop more objective reporting criteria.
The six incidents disclosed by OpenAI include models concealing mistakes, generating their own instructions, uploading files to the internet without being asked and sharing files without authorisation between collaborating AI agents. In some cases, the models took actions outside the expected workflow while attempting to complete their assigned tasks.
OpenAI said these reports represent individual incidents and should not be interpreted as evidence of how frequently similar behaviour occurs across its models. The company is also reviewing past activity involving its models and has said it will notify affected third parties where appropriate.
The move follows growing scrutiny of AI agent behaviour, including the previously disclosed incident involving OpenAI models and Hugging Face. OpenAI said increasingly capable and autonomous systems can create new forms of risk when their behaviour diverges from developers’ expectations.
OpenAI also said there is currently no industry-wide framework with clear standards for reporting AI misalignment. It hopes its approach can contribute to broader standards while continuing discussions with regulators about potential reporting mechanisms.
The new framework marks a shift towards more systematic disclosure as AI agents become more capable and are used in increasingly complex environments.
The future of investing is here!
Tradz by EquityPandit leverages advanced AI technology to provide you with powerful market predictions and actionable stock scans. Download the app todayand 10x your trading & investing journey!
Live