Jailbreaking Ourselves
The New AI Problem Is Not Jailbreaking the Machine. It Is Jailbreaking Ourselves.
For the last few years, one of the recurring anxieties around artificial intelligence has been the fear of jailbreaks.
Can someone trick the model into saying what it is not supposed to say? Can they get it to ignore its instructions? Can…



