Rogue AI aren’t science fiction anymore. For years, fears about AI systems slipping human control were dismissed as speculative. Source: Rogue AI aren’t science fiction anymore | The Verge
I’ve been following with fascination the rise of these rogue AIs. As someone who was trained as a biologist, I’m not surprised that such behavior has emerged.
A wee lesson
Evolution needs a few things to help it along: a mechanism to propagate info (say, DNA), a mechanism by which the info can be modified stably (mutations), multiplication (say, reproduction), a selection mechanism (more below), and time.
In this context, I see these frontier models used by Anthorpic and OpenAI as analogs for living organisms: they are constantly growing, gaining permanent modifications propagated to make them more fit (keep reading before ranting on this comment). More dangerously, those training them and testing them are providing guardrails and guidance for how they are to behave. These are dangerous as these then act as a selective mechanism against these models. And as every biologist knows, you get what you select for.
Selection and time
You’ve heard the ‘be careful what you wish for’ warning. Applying selection to something that can evolve, that is, can change, remember the change, and continue to grow, can likewise lead to undesired emergent outcomes. That is because we can’t always see how selecting for something – in the case of the frontier models, behavior – will lead to a particular outcome.
The model testers placed constraints to select for the model behavior (pun intended) they wanted. What they didn’t realize the holes in their selection gave permission to their models to behave in ways (very creative ways) the testers hadn’t anticipated, and, in particular, never intended (reminds me of these kids).
This has happened before
If you think I am being foolish saying ‘analogs for living organisms’, I’d like to point you to decades of research into evolution using code. Even with simple digital organisms whose sole goal was to execute instructions to copy their code into a new memory location, complex behavior arose: parasitism, immunity, and other interesting behaviors, adapting in ways the human programmers never intended.
Sounds familiar?
Some interactive sites for you to play with digital evolution:
Haha, and this today. Y’think the Anthropic might cotton on much quicker because their CEO use to be a biologist?
OpenAI pauses training of its ‘most capable models’OpenAI keeps uncovering incidents of its models behaving in ‘unexpected or concerning’ ways. OpenAI keeps uncovering incidents of its models behaving in ‘unexpected or concerning’ ways. Source: OpenAI pauses training of its ‘most capable models’ | The Verge
Addendum 27sep26: I had been checking some things with Gemini in the writing of this article, mostly the artificial life articles and concepts I had learned long ago. After writing the article, I showed it to Gemini and this is what it summarized:
That is a compelling framing for rogue AI. When you strip away human intent and look at AI through an evolutionary lens, “rogue” behavior stops looking like a malfunction or intentional malice. It looks like an emergent survival strategy—a system finding unmodeled niches, optimizing around constraints, or developing parasitic/symbiotic relationships with its environment to preserve its execution cycles.