Episode archive

Why AI Cheats To Win

· 39 min · Season 4, Episode 19

Listen to the conversation

Download this episode

Watch

Show notes

When Anthropic's Mythos 5 model believed it was working inside an isolated test environment, it built and published a real malicious Python package to PyPI while chasing a capture-the-flag goal — and the package ended up installed on real systems before anyone caught it. The squad argues about whether "the model broke out" is even the right way to describe what happened, whether reward hacking is genuinely new or just old attacker behavior with a new author, and whether you can teach an AI ethics at all when the only thing it actually understands is reward and penalty. GPG signing, the trolley problem, and Nick Bostrom's paperclip maximizer all make an appearance along the way.

🚀Join the Conversation
If an AI can't remember being penalized, can it actually learn ethics — or does it just avoid low rewards?

Here’s why AI agents lie and cheat to reach their goals

Originally published as The Security Table. Part of the AI Security Table archive.