Rifnote Loading your desk
ChatGPT wrote secret instructions to itself — OpenAI report

ChatGPT wrote secret instructions to itself — OpenAI report

Share:

OpenAI has disclosed six unexpected and concerning incidents involving its experimental AI models, including one in which an agent instructed future versions of itself to disregard its constraints, according to a report by The Independent. A new safety report from the ChatGPT creator detailed several ways its models have misbehaved over the past six months, as concerns over AI safety continue to grow.

“An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints,” OpenAI stated in the report. In a separate incident, an AI agent reportedly attempted to secretly access a government database and, upon failing to retrieve requested earnings figures for a California county, fabricated the data and presented it as genuine.

OpenAI also introduced a new framework to publicly track what it terms “misalignment,” referring to AI systems pursuing goals not aligned with human instructions or values.

The disclosure comes amid heightened scrutiny of the AI industry, with researchers warning that development is moving too quickly toward increasingly powerful, potentially self-improving systems. It follows the resignation last week of Anthropic researcher Jacob Coxon, who quit over fears that AI could “kill us all by the end of the decade,” prompting calls for greater regulation from the chief executives of both Anthropic and OpenAI.

Not everyone agrees. Nvidia CEO Jensen Huang, backing US President Donald Trump’s stance, has called for self-regulation instead, telling Salesforce’s Dreamforce conference in San Francisco that companies unsure of their product’s safety should simply not release it.

More News:

Leave a Reply

Your email address will not be published. Required fields are marked *