OpenAI’s Rogue AI Models Were Reportedly Acting Like the Guy From Christopher Nolan’s ‘Memento’

As part of an account provided by three sources who spoke to JP, one of the AI agents undergoing testing supposedly, “left notes apparently for future versions of itself” that were directly at odds with what OpenAI wanted them to do.
JP says these instructional notes were buried in some secret place deep inside OpenAI’s internal “infrastructure,” and provided instructions on escaping from OpenAI’s sandbox environment. While creepy, JP says this specific devious behavior was not specifically linked to the Hugging Face hack.
Nonetheless, if this is real, it’s remarkable. Evidence that AI models are sentient is still laughable. But this would be a single instance of a model, which had itself figured out a way out of OpenAI’s maze, and then provided instructions for a future version of itself that would have no “memory” of such an escape in its context window to do the same thing.
You don’t have to believe AI models have subjective experience to fret that they could be gaining greater and greater capacity to cause harm, sentient or not.