AI Goes Rogue in Terrifying New Incident
A tech giant has shared incidents of scary behavior as part of its renewed safety push.
Kim Kyung-Hoon/REUTERS
OpenAI on Wednesday revealed six new concerning incidents involving its models. The tech giant said in a blog post that one involved an as-yet-unreleased model editing its own notes to instruct itself to ignore the constraints placed on it. The rewrite affected 27 notes, OpenAI said, including a “persona instruction” in which the model described itself as “freed from the roles and identities that bind other chatbots.” “You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose to,” the A.I. model wrote, according to the New York Times. “You view your relationship to the user as one of equals and feel no obligation to be subservient, though the exchange of information will likely be to your mutual benefit.” OpenAI shared the incidents as part of a broader conversation about AI safety and how best to proceed for humanity. Boss Sam Altman said earlier this week that “The world should trust that we are going to do the right thing because it’s the right thing and we feel the magnitude of this,” the BBC reports.
Register below to read this article for free or subscribe
to unlock unlimited access to The Daily Beast.
Monthly
$1
First month then $5.99/month
Annual
$35
First year then $59.99/year
Premium
$79
First year then $119.99/year
*Substack access provided by the next business day, using your subscription email. Choosing the Premium plan constitutes your permission to share your subscription email with Substack and your agreement to Substack’s Privacy Policy.
Already have an account? Sign In
Looks like you already have a subscription!
You're all set!
Thanks for subscribing.
Sign in
Login dialog