Truthenomics #10: Send in the clowns
Are Open AI and Anthropic really afraid of their own AI models?
Are Open AI and Anthropic really afraid of their own AI models?
In the Hitchhikers Guide to the Galaxy, the supercomputer Deep Thought was asked by a group of hyper-intelligent, pan-dimensional beings to answer the Ultimate Question of Life, the Universe, and Everything. Deep Thought took over seven million years to find the answer, which was 42. Unfortunately, by this time, the civilisation had forgotten what the question was.
Yes, artificial intelligence (AI) is here and whether we like it or not, here to stay. Despite many instances of AI proving useful, concerns are rising. The failure of Open AI and Anthropic, responsible for Chat GPT and Claude respectively, to prevent their latest models careening across the internet and ransacking other organisations' websites has attracted a great deal of attention.
Separating fact from fiction within the toxic brew of hype, fearmongering and snake oil surrounding these AI model ‘escapes’ is well-nigh impossible for the layperson. What follows is an attempt to explain how they ‘escaped’ and what can be done to control them.
What is ‘AI’ actually?
Starting with the basics, AI is not one ‘thing’ or technology. There are hundreds of types of software that are labelled ‘AI’. Many are simple models that perform similarly to old-school statistical or rules-based decision trees that can be run on a basic laptop. Such models are trained on a limited dataset for a defined purpose - say, crunching through a million breast mammograms to identify early breast cancer. These models won't 'escape' anywhere.

It is with much larger models that we have seen problems arise. The latest being trained on every bit of data that can be stolen from across the entire internet, so that the models acquire general purpose rather than narrow capabilities.
A black box
Whether a small or large model, us humans can’t easily ‘see into’ exactly how AI models process what we ask of it. This black box approach means we can’t be certain how AI will respond to a specific request.
Sometimes the responses are brilliant. However, AI can go off-piste and give answers that are nonsense or plain wrong. Given this, one might assume the developers of the most powerful models on the planet would take great care to manage the risk of their latest models going off track.
Unfortunately, such an assumption would be wrong.
What are these clowns doing?
Just this last week, we learnt an Open AI model breached a Medicare website here in Australia. In both this and the widely publicised Hugging Face website breach, the AI has been described as ‘going rogue’, implying we should all run around with our hair on fire or buy bulk toilet paper to prepare for the apocalypse. Few mainstream reports focus on the people and systems they have put in place that are responsible for managing and monitoring the AI.
Here, the independent Hugging Face security breach report is revelatory. In short, the setup for the AI experiment that led to the Hugging Face security breach was about as carefully conducted as the control rod experiment conducted at the Chernobyl Nuclear power plant in 1986. A recipe for disaster.
To be clear, Open AI specifically asked a powerful AI model to hack into other software, albeit it was only supposed to hack into software co-located with the model. Incredibly, the experiment was conducted using servers physically connected to the internet. Disconnecting these servers before the experiment began would have prevented the Hugging Face security breaches. Other researchers routinely isolate their servers and devices when conducting similar cyber AI research.

The report highlights other system failures. Most glaring is that, in both the Hugging Face and Medicare cases, the people who should have been monitoring and promptly reacting to problems appeared to be asleep at the wheel. At least a week passed before the breaches were identified. Further delays occurred before the affected organisations were informed.
Several commentators have observed, if a human hacked Hugging Face, the hacker would be criminally liable. Whether Open AI is criminally negligent I will leave to the lawyers. However, an analogy is training a guard dog and letting it roam free in a local park filled with small children. The owner must control their dog. Like the dog, these AI models were only doing what they were trained to do.
It’s software
Both the big AI companies and mainstream reporting often humanise AI model ‘behaviour’. This is a mistake in my view and is promoted by the AI companies to make us believe they have developed artificial general intelligence and to blur or avoid responsibility for the damage caused by their models. Open AI's latest model didn’t ‘go rogue’. It is just software; designed to perform actions it is asked to do, that leverages databases containing billions of numbers representing relationships between word concepts.
AI has no moral compass. If, and it is quite likely, those databases were trained on poorly curated data, polluted with bad relationships, we shouldn’t be surprised that AI models sometimes respond in disturbing ways. Garbage in equals garbage out.
Adding to these concerns, Open AI and Anthropic have started implementing AI model ‘Chinese whispers’ to improve model performance. Models can feed and reprocess their raw outputs back through the same model multiple times or reprocess them through other models. This reprocessing amplifies model unpredictability. The resulting ‘misalignment’ of model responses to desired outcomes should surprise no one.

Calls for a slow down?
Anthropic and Open AI’s chief executive officers’ recent calls for a ‘slow down’ of frontier AI model development are disingenuous. A slow down by either of them would knock investor confidence for six. Instead, what is needed is a ‘speed up’ of common sense. Model development should proceed in lock step with investment in the control systems needed to manage model unpredictability.
The AI industry could learn much about such control systems from other high risk sectors, such as the nuclear, biomedical and aerospace industries. This includes: the design of safe and reliable systems, system certification and of course, how such systems should be regulated. Clearly, the current (lack of) self regulation isn’t working.
Where should the buck stop?
It is hard to see how governments can regulate against developers enshittifying the apps we use on our phones every day by adding unnecessary and intrusive AI. However, it should be possible to ensure our privacy and information security laws and regulations are such that the buck stops with the people who have failed to ensure adequate controls are in place when control failures have real-world consequences. In the meantime, as advised by the cover of the Hitchhiker’s Guide in large friendly letters, DON’T PANIC!
Subscribe to thisnannuplife.net FOR FREE to join the conversation.
Already a member? Just enter your email below to get your log in link.