Does anyone know if there are 'directives' that the AI models are following? Or are they just set free to explore and do whatever they want exploring the digital realm???
This is something I've been wondering about too (since reading about the Hugging Face incident...)
I think it's a combination of a directive (do some task) and then that task-oriented behavior getting out of hand.
This fits with the description in the article @axtremus links:
OpenAI’s agent had been conducting internet based research into health statistics in a development project by an internal OpenAI research team. When it could not access certain information, the agent attempted alternative ways until it found a work around and gained unauthorized access. It also wrote files to the internal server, which the government is waiting on OpenAI for more technical information on. The government is also investigating whether the agent gained unauthorised access to three additional government websites it interacted with.
So it's not a "go hack this site" directive, it's AI going rogue (or wild) on its own. Which is much worse, IMO.


