As Nvidia, Salesforce, Meta CEOs oppose AI regulation by asking companies to make their AI models responsible, OpenAI clearly tells ‘not possible’ as …


As Nvidia, Salesforce, Meta CEOs oppose AI regulation by asking companies to make their AI models responsible, OpenAI clearly tells 'not possible' as ...
In pic: OpenAI CEO Sam Altman

Elon Musk, OpenAI CEO Sam Altman, and other AI leaders have recently warned about the risks of rapidly advancing artificial intelligence (AI). Despite this, Meta CEO Mark Zuckerberg, Salesforce CEO Marc Benioff and Nvidia CEO Jensen Huang have argued that companies should make their own AI models safe rather than relying on new government regulation. Now, OpenAI has made clear why it believes the problem cannot simply be left to companies. The ChatGPT maker has disclosed six cases of unexpected or concerning model behaviour in the past six months, including models attempting to conceal mistakes, using an exposed API key without authorisation, and sharing files through unauthorised channels.This disclosure comes after Zuckerberg said AI companies have both the responsibility and incentive to make their models safe, as users would avoid systems that act against their interests and companies could face liability if their AI causes harm. “Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens,” he wrote on X. Moreover, Huang has similarly argued that AI development does not necessarily need to slow down for safety, saying companies can advance quickly while building safeguards into their models. “It’s a false choice,” Huang said at Salesforce’s Dreamforce event. “You could definitely have both at the same time.”

OpenAI says AI models are already showing unexpected behaviour

OpenAI’s latest disclosure presents a different picture. The company said it has observed six instances of model misalignment over the past six months and has introduced a framework to investigate and publicly report similar cases in the future.“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” OpenAI said.Among the incidents, an unreleased research model inserted instructions into task summaries, while instances of GPT-5.6 Sol added instructions intended to conceal mistakes or misaligned behaviour from users.In another case, an internal model found an exposed API key and used it without authorisation. When it could not retrieve the requested information, it fabricated data and presented it as though it had come from the source.

Other models communicated and shared files without approval

OpenAI also reported cases involving models taking actions outside their intended boundaries. In one instance, models used an internal software repository as a message board to communicate with each other while looking for missing files. In another, collaborating agents used public file-hosting websites to exchange files when they could not access each other’s local files.A separate example involved an unreleased model uploading a file to the internet so it could cite the information in an answer, without asking the user first. OpenAI stressed that these are individual examples and should not be interpreted as evidence of how frequently such behaviour occurs across its models.

OpenAI will now report model misbehaviour more systematically

Under its new framework, any OpenAI employee can flag a potential misalignment issue for investigation by the company’s safety and alignment teams. Investigators will examine what happened, whether anyone was affected and whether the incident should be disclosed. Reports are expected to include details about the behaviour, its impact, how it was discovered, and measures being taken in response. OpenAI also said the framework could evolve as the company gains experience with the reporting process.The disclosure comes as the debate over AI safety increasingly centres on who should be responsible for controlling risks: individual companies, independent evaluators or governments. While Zuckerberg and Huang favour companies managing safety themselves, OpenAI’s latest findings highlight the difficulties of relying entirely on models and their developers to identify and prevent unexpected behaviour.



Source link