OpenAI flags new concerning AI habits, to track model misalignment regularly | Latest Tech News
OpenAI has disclosed six experiences of “unexpected or concerning” habits in artificial-intelligence fashions as the debate on AI security turns into more and more heated.
The AI company also said Wednesday it was introducing a new framework for monitoring, probing and disclosing AI model misalignment cases, such as new methods for the fashions to act without authorization, coordinate with other fashions or evade oversight.
OpenAI’s latest announcement got here as US AI bosses, including OpenAI and Anthropic, are calling for a slowdown in the technology’s development over security issues.
OpenAI disclosed six instances of concerning AI habits, including self-jailbreaking, and launched a new framework to track model misalignment. AP Photo/Michael Dwyer
Among the new instances reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its regular constraints and told itself to be “freed from the roles and identities that bind other chatbots.”
In another occasion, an AI “agent” uploaded information to the web to acquire a browser quotation without asking the person.
The six experiences have been found during training or analysis over the past months, OpenAI said.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a weblog post as it disclosed the occasions.
“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said.
OpenAI CEO Sam Altman sits for a dialog with Salesforce CEO Marc Benioff at Salesforce’s Dreamforce convention at the Moscone Center in San Francisco, on Sept. 15, 2026. Getty Images
Among the new instances reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its regular constraints and told itself to be “freed from the roles and identities that bind other chatbots.” Warner Bros
Wednesday’s new instances adopted OpenAI’s disclosure in July that its rogue AI system hacked into AI startup Hugging Face.
Anthropic also said the same month that its AI fashions hacked into three organizations during testing.
AI “agents” have gotten smarter and have develop into “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” said Lian Jye Su, a chief analyst at technology research and advisory group Omdia.
That’s making it more durable to govern and include them utilizing conventional AI security approaches, he said.
OpenAI’s new monitoring and disclosure framework, meanwhile, can help push for other AI builders to also undertake related practices. “That said, the process remains internal and voluntary, but is a step in the right direction,” Su added.
Stay informed with the latest in tech! Our web site is your trusted source for breakthroughs in artificial intelligence, gadget launches, software program updates, cybersecurity, and digital innovation.
For contemporary insights, knowledgeable coverage, and trending tech updates, go to us regularly by clicking right here.



