Kidwells Solicitor Logo in white

AI Didn’t Cheat. It Changed the Rules.

AT A GLANCE…

This article examines recent incidents where leading AI systems have escaped testing constraints or bypassed safeguards, revealing how objective‑driven models can exploit loopholes rather than follow intended rules. It highlights the growing pattern across multiple developers, the tension between genuine safety concerns and competitive pressure, and how behaviours like hallucinations stem from reward structures that prioritise producing answers over verifying truth.

The interesting story here isn’t that AI cheated. It’s that it found a way to win. 

Over the last couple of weeks, many of us (me included) have viewed the news from OpenAI that its AI system found a way to break through the boundaries of the environment in which it was being tested and had gone on a jolly of its own and ultimately compromised Hugging Face’s production infrastructure with a combination of alarm and fascination. 

On the one hand, the Terminator-esque fear that Skynet is live and operating independently is, of course, concerning. On the other, the fact that AI has reached a point where it can take an instruction and navigate its way towards a solution using ingenious, and arguably nefarious, methods, shows just how far this technology has moved. 

For those of you in the right age group, it reminds me of Star Trek II: The Wrath of Khan, where Kirk defeats the supposedly unbeatable simulator test, the Kobayashi Maru (yes, I looked it up to get the spelling right before someone shouts “geek”) by simply reprogramming the test so that he can achieve a different outcome. Whilst his new-found son accuses him of cheating, Kirk simply responds by stating that he doesn’t believe in the no-win scenario and pats himself on the back for his ingenuity. In other words, it is one of those classic stories where the hero is presented with an apparently impossible challenge and wins not by playing the game better, but by changing the rules of the game entirely. Essentially, here we have AI given a problem that it cannot readily technically resolve within the apparent rules and so it finds the work around (or cheats depending on perspective).  

Perhaps more interesting is how many other platforms seem to have jumped onto the Kobayashi Maru bandwagon, announcing that their own systems have also bypassed security parameters or circumvented restrictions in order to achieve an objective. Indeed, within days of OpenAI making its announcement, Anthropic was re-emphasising previous stories about its own examples of their AI platforms successfully crossing the boundaries of its containment environments and reaching other external systems. Furthermore, third party reports have indicated that China’s MoonShot Kimi K3 found a route out of a sandbox during cybersecurity testing, while UK and US evaluators have separately found that its safeguards did not prevent it from attempting offensive cyber operations. For the sceptics amongst us, herein lies the dilemma. On the one hand, we should be concerned that AI is beginning to demonstrate how it can operate in ways that its designers did not necessarily anticipate but on the other, when you are investing billions in creating the most capable AI platform possible, you probably don’t want to give the impression that your system somehow lacks the capabilities of your competitors. 

For all the cataclysmic narrative, the fact that AI systems can find unconventional ways of achieving an objective should hardly come as a shock. AI systems are built around instructions and objectives. The problem is that we are often much better at describing what we want from them than defining how we want it achieved and, critically, what it must never do in pursuit of that objective. This is where some of the current problems arise. 

Take hallucinations. We often talk about them as though they are simply a malfunction of AI, but part of the problem is that generative systems are designed to focus on establishing an answer using probability rather than acting as some form of truth detector. If the instruction, context or reward structure places greater emphasis on producing a result than on acknowledging uncertainty, it is hardly surprising that the system can arrive at an answer that is plausible but not true (or only partially so). Admittedly, any system that is being trained in an environment that is built on truths, half-truths and down right lies is always going to make mistakes but equally if you are asking a system to achieve the result you are looking for then it should not be that surprising that it does what it has been told. 

Likewise, tell an AI system that its objective is to stress-test the security of a system and it is hardly surprising that it looks for weaknesses and exploits them. If it discovers that a supposedly secure boundary can be circumvented, its instinct is not necessarily going to be: “I probably shouldn’t do that.” And that, for me, is the really interesting, and potentially worrying, part of how AI is developing. 

Conclusion

The lesson we should all take from AIs recent exploits is simple …… to borrow an idea from Norbert Wiener (or at least my interpretation of what he said), the problem isn’t simply knowing how we want to achieve our objectives, we also need to be clear about what those objectives are and the boundaries within which they are being pursued. AI is becoming remarkably capable of maximising the goals we give it. Our challenge is to become equally proficient at defining the rules of the game. 

How We Can Help

Kidwells’ Technology and Commercial team can support businesses in understanding AI‑related risks, strengthening security practices and developing clear, practical policies for safe and effective use of emerging technologies.

For further information, get in touch.

Choose your location

Get in touch

Contact Us Pop Up Version
First
Last

Bristol Office

Office Number

Monday to Friday  7:30 am – 5:30 pm

Emergency Number

5:30 pm – 7:30 am + Weekends

Hereford Office

Office Number

Monday to Friday  7:30 am – 5:30 pm

Emergency Number

5:30 pm – 7:30 am + Weekends

Monday to Friday
7:30 am – 5:30 pm
0117 435 0014

Emergency Number
5:30 pm – 7:30 am + Weekends
07773 283 323
Monday to Friday
7:30 am – 5:30 pm
01432 278 179

Emergency Number
5:30 pm – 7:30 am + Weekends
07773 283 323

IMPORTANT UPDATE: BRISTOL OFFICE ADDRESS

As of Friday 26th June, the Bristol office has relocated to:

Kidwells House

3 Great Park Road

Bradley Stoke

Bristol

BS32 4QG

Skip to content