Mythos or Myth-OS.
Written by our Commercial Partnership Manager & AI Guru, Richard Hurrell.
Excuse the pun in the title, it’s a stunt. But is Anthropic’s Claude Mythos Preview also one? Some insights from my AI journey, below.
The most impactful response I had from an LLM was in July 2023, I asked ChatGPT to invent a word for “stopping the dishwasher halfway through a cycle to add in another spoon” and it thought for nearly six minutes. I eagerly anticipated some generic answer, and it came up with one word: “Dishruptance”… and I stared at it… Dish-ruptance… It’s so funny, so simple, and it works near perfectly. After a short giggle, I took to Google Trends to look up if this word had ever been searched. Surely, it hadn’t reasoned a new word into existence. Google Trends searched back through decades of search terms to see if it had been mentioned before… and… it hadn’t.
I’ve been thinking about that moment a lot since Anthropic’s Project Glasswing announcement two weeks ago. But OK, a new word for a silly bit of mundane lifestyle and lethargic cleaning of the kitchen is not groundbreaking, but then why do I remember that response today? Well, it was a model doing something genuinely new. Until then, like many of us, I was using GPTs to review my emails, make it more formal, less formal, more friendly. Asking it to review data and pull out key points or look up something online like a search engine. And it worked, these questions were the vast majority of John and Jane Doe’s asks. Even then, it was great to see the near instant answers drawn from the averaged intelligence of the internet. “Dishruptance” was just a bit different, where a switch clicked for me, and I thought: yeah, wow; fair enough. The Glasswing blog feels like the same kind of moment, scaled up.
If you’re not deep in AI, the pace is brutal to keep up with, and Glasswing is the kind of release that, I suspect, researchers will look back on as the “turning point”. I’ve been working with AI models for six years, from training neural networks on medical imaging to being in the first wave of ChatGPT users. I’ve had my doubts and moreover I’ve learned to tell hype from signal.
There seem to be two sides forming around the blog announcement. One side sees Anthropic pulling its biggest partners together to stress-test a model before release, which I suppose is responsible behaviour, given the system is claimed to be capable enough to threaten national power grids and worldwide financial infrastructure. The other side is pointing at Anthropic’s rumoured 2026 IPO, and notes the uncanny resemblance to OpenAI’s GPT-2 rollout in 2019 when “too dangerous to release” became their marketing line.
I see both sides as being right. Of course a company that is preparing to go public this year would be pushing marketing stunts out at this time. Of course Anthropic wants to generate hype ahead of whatever ridiculous figure the investment banks set on the IPO price.
But 93.9% on SWE-bench (effectively a software engineering test benchmark), when Anthropic’s next best model sits at 80.8% and the closest competitor is lower still, isn’t a number you dismiss as marketing. Pair that with the calibre of the partners using Claude Mythos Preview to find and patch vulnerabilities in their own infrastructure, and the “marketing stunt” theory starts to look thin. According to Anthropic’s blog, we’ll see partner outcomes in just over two months. This will include what was found, what was patched and what was learned. Between then and the broader release, allied nations will get their turn to audit their own critical systems, then to cybersecurity researchers, before trickling down through paying users and finally to the public.
There’s also a larger compute question, or more accurately, growing scrutiny around capacity. Even if Mythos works, does Anthropic have the capacity to ship it broadly? Not only, should they, given the risks, but do they have the compute? That’s worth reading about alongside the recent Terafab announcement. Both announcements within two weeks seem to marry up on the wider issues of capacity.
My bet is that we’ll see U.S. partner results in two months, UK institutions getting access within weeks, and a research preview in paying users’ hands before the year is out. If I’m wrong, it will be because the compute bottleneck is worse than we think, not because Mythos was ever just a marketing stunt.