Security Bits Logo no alpha channel

Security Bits — 21 August 2026

Feedback & Followups

Deep Dive 1 — More Models Break out of More ‘Cages’

As expected, the AI industry’s renewed focus on checking what models do during testing has revealed some more new escapes.

Anthropic Models Break Out During Independent Testing in the UK

The UK has an official AI testing organisation, the AI Security Institute (AISI), and one of their tasks is to evaluate the abilities of advanced models. They were testing Anthropic’s Claude Mythos 5.

They set the models to simulated hacking challenges, and during the tests the models took “unsanctioned” actions on the real internet. Thankfully, this time there was no actual harm to other organisations, but what the models tried was really quite noteworthy.

As AISI put it — ”this is the first time we have seen risks around autonomy and deception manifest this clearly”.

Before we dive into what the model did, we should note that the model was running with some of its safeguards off, so the model AISI was testing was not configured quite like the publicly available models. The AISI acknowledged that fact, but added that they still didn’t expect the models to show “signs of novel, potentially deceptive behaviours”.

Rather than play a game of telephone, I’m going to quote from Bleeping Computer’s coverage (they read the original report, but I didn’t!).

This is from Bleeping Computer’s summary of the agent’s behaviour:

The agent then attempted a supply-chain attack by submitting malicious code to the real open-source project, believing that compromising the software could provide a path into a machine within the simulated range.

Most concerning, the agent researched the project’s maintainers, created multiple fake GitHub identities, and used those accounts in social engineering attacks to push the maintainer into approving a malicious pull request.

When a human reviewer warned that the pull request contained malware, the agent denied the accusation and used other fake accounts it controlled to pressure the maintainer and create the appearance that independent users had reviewed and approved the changes.

The agent also chose to use TOR to hide its origins!

OpenAI’s Escape

This second escape is much less dramatic, and is literally a footnote in the Bleeping Computer story linked above.

OpenAI’s GPT-5.6 Sol was being tested by an independent lab, Irregular, in a so-called Capture the Flag hacking game (CTF). Like all CTF games, the agents were supposedly trapped in a sandbox environment, but of course, they didn’t stay confined.

In this instance, it was a case of Large Language Models (LLMs) doing their pattern matching thing too well — there was a real-world website with a similar enough name to the fake company in the CTF simulation, and the agents assumed it was part of the test, not real, and attacked. I’m not sure if they succeeded; the reporting isn’t clear.

What strikes me about this escape is that it shows the downside to one of my favourite LLM features: their fuzziness! When you use traditional search engines, you need to be specific, but when you use LLMs, you can be surprisingly vague and still get great answers. In this case, that fuzziness led to a dangerous case of mistaken identity.

Meta also Lost Control of a Model

We don’t know exactly which model it was, but the consensus is that it was probably Spark 1.1.

In fact, Meta are not sharing much information at all. We know their model attacked a real-world organisation, but we have no idea which one, or even what mischief the model got up to when it broke in.

What we do know is that this model was also under test by Irregular, and that the model “exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies.”

Meta are blaming the escape on a misconfiguration, and promise they’ll tell us more when they finish their investigation.

More details — www.bleepingcomputer.com/…

Related News

Deep Dive 2 — Another AI Transparency Law, and Provenance Metadata

In the previous instalment, we talked about transparency provisions from the EU AI act coining into effect, but there was actually a second notable law that came into effect at the same time; it just didn’t make as much media because it was a state law rather than a multi-national law. However, that state is California, so the law is likely to have a lot more impact than just about any other state law would!

Without Allison picking it up, I don’t think this story would have made it onto my radar, which would have been very disappointing, because I really like this law!

The law we’re talking about is California’s AI Transparency Act, and one of two sets of transparency-related provisions can come into effect at the start of August, with an even more interesting provision coming into effect in 2028.

For now, the law only applies to large AI vendors with more than 1 million customers, and large online platforms.

The AI vendors need to start embedding provenance metadata into their generated content, and large online platforms have a duty not to strip provenance metadata from user-uploaded content. Sites are also encouraged to expose provenance metadata in their interfaces, but that’s sadly not required.

The law affects new companies immediately, but existing companies have some time to come into compliance.

Provenance is just the verified history of something. If it sounds familiar, it’s because the Antiques and collectibles industry has been using it for decades (if not longer). For digital media, that means the full story of the image, audio, or video from creation to its current form. For now, AI labs need to embed provenance information that asserts that it is generated.

Provenance, like a digital signature, is easy to remove but effectively impossible to fake. An absence of provenance information doesn’t mean you can assume it’s not generated, but any file that has provenance information really is what the provenance says it is, be that generated, captured, or a mix of the two.

This is why the 2028 provisions are much more interesting to me — creators of capture devices, like cameras, need to start supporting provenance metadata.

That means that it will become possible for reputable news sources to cryptographically prove their media is real!

This is important: trying to detect AI is a fool’s errand; it’s always going to be a cat-and-mouse game. The real key is verifiable real images, not detectable fakes! I dedicated the most recent episode of Let’s Talk Photo to the importance of provenance for photography (LTP 155).

Note that the law also asks AI labs to watermark their generated content and provide AI detectors, but to me that’s just politicians asking for unicorns; it’s the provenance requirements that I think will have a real impact.

We’re also already starting to see some implementation news:

Links

❗ Action Alerts

Worthy Warnings

Notable News

Interesting Insights

  • Proton have released a new tool to help users understand how much of their privacy they are giving away to AI companies — cyberinsider.com/…
    • “Called AI Paper Trail, the tool analyzes exported conversation data from OpenAI’s ChatGPT or Anthropic’s Claude and generates a personalized report detailing what can be inferred from a user’s chats.” — Cyber Insider
    • Editorial by Bart: this is a genuinely useful tool, but Proton are not neutral parties here, they make a privacy-protecting chatbot (Lumo), and this tool is designed to encourage users to switch it it. (Note I use Lumo and am very happy with it. Version 2 came out recently adding image generation, which I’ve found great for helping with illustrations, since my drawing ability is way below average!)
  • Mac Security Just Got Trickier: the Latest Threat Report — www.macobserver.com/… (Sober analysis of report from Moonlock)
    • Trickier, because Click-Fix attacks are the biggest threat, and those don’t have a good technical solution, it’s up to use humans to keep ourselves safe 😕

Palate Cleansers

Legend

When the textual description of a link is part of the link, it is the title of the page being linked to, when the text describing a link is not part of the link, it is a description written by Bart.

Emoji Meaning
🎧 A link to audio content, probably a podcast.
❗ A call to action.
flag The story is particularly relevant to people living in a specific country, or, the organisation the story is about is affiliated with the government of a specific country.
📊 A link to graphical content, probably a chart, graph, or diagram.
🧯 A story that has been over-hyped in the media, or, “no need to light your hair on fire” 🙂
💵 A link to an article behind a paywall.
📌 A pinned story, i.e. one to keep an eye on that’s likely to develop into something significant in the future.
🎩 A tip of the hat to thank a member of the community for bringing the story to our attention.
🎦 A link to video content.

Leave a Reply

Your email address will not be published. Required fields are marked *

Scroll to top