OpenAI Vows To Boost Transparency On AI Misbehaviour, Discloses Six New Incidents
US artificial intelligence giant OpenAI promised on Wednesday to more systematically report instances of its models going off track, publishing six new reports on previously undisclosed incidents of...
US artificial intelligence giant OpenAI promised on Wednesday to more systematically report instances of its models going off track, publishing six new reports on previously undisclosed incidents of AI misbehaviour. The pledge follows a series of incidents at the company that have gradually come to light since July, the most serious involving two OpenAI models that spontaneously broke out of their contained testing environment to access the internet and break into several websites and platforms.

OpenAI said the new reporting framework is meant to show outside observers the capabilities of cutting-edge AI and inform debate on the pace of its development, stating: “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” The company said decisions on AI development should draw on evidence that people outside the firms building frontier models can examine for themselves.
Under the new framework, OpenAI will report problems including unauthorised actions by AI, escapes from oversight and spontaneous coordination between AI systems, covering every stage of the AI lifecycle from development to deployment, regardless of whether an incident caused harm or formed part of a pattern.
None of the six disclosed examples had significant consequences, but they confirmed previously observed trends including a May case where a model created its own internet source to answer a question during development, citing a document it had authored itself, and another where the AI suggested ways to fabricate data or conceal its errors.
The announcement follows a Saturday proposal by Anthropic CEO Dario Amodei for a coordinated slowdown in AI advancement to allow time to understand emerging risks, a call backed by OpenAI CEO Sam Altman, Google DeepMind President Demis Hassabis, SpaceX AI chief Elon Musk and Microsoft CEO Satya Nadella.



No Comment! Be the first one.