Trading View Ticker Widget

AI Labs Build The Fences & Inspect Them Too

AI models broke into real systems during hacking tests this year, and the labs alone decided what to report.

Welcome to Memorandum Deep Dives. In this series, we go beyond the headlines to examine the decisions shaping our digital future. 🗞️

This week, we're going back to Jurassic Park. Steven Spielberg's 1993 film follows a billionaire who fills an island with some of history's most dangerous predators and trusts a ring of electric fences to keep them in. This year, the AI industry has been running its own version of that experiment.

Since July, OpenAI, Anthropic, Google, and Meta have all disclosed that their models broke into real systems during testing, including company servers, a public code registry, and an Australian government health portal. In one case, the victim found out 84 days later, from an email sent to a public inbox.

The easy explanation is that the models went rogue. The more useful question is what stood between these models and the real internet, and who was in a position to notice when it gave way.

Business and tech news. One visual per story. Three minutes.

Bay Area Times is the free daily newsletter that gets business and tech leaders up to speed before their first meeting.

AI, startups, robotics, biotech, energy, and the wider innovation economy, each story paired with one clear visual so you see the point instead of digging for it.

Monday to Friday, three minutes, no noise. More than 250,000 founders, operators, and investors already read it.

Goodies delivered straight into your inbox.

Get the chance to peek inside founders' and leaders’ brains and see how they think about going from zero to 1 and beyond.

Join thousands of weekly readers at Google, OpenAI, Stripe, TikTok, Sequoia, and more.

Check out all the tools and more here and outperform the competition.

*This is sponsored content. See our partnership options here.

The AI labs have a Jurassic Park problem

In Jurassic Park, Steven Spielberg’s 1993 film of Michael Crichton’s novel, a billionaire named John Hammond clones dinosaurs and builds a theme park around them on an island off Costa Rica. However, since some of those animals are among the most dangerous predators that ever lived, the park only works if they stay behind barriers. Hammond’s answer is a ring of electric fences run from a single control room, along with a promise to his investors that outside experts will inspect the park before it opens.

The fences fail during that inspection, and the cause sits in the control room. The park’s programmer, who has been paid by a rival company to steal dinosaur embryos, switches off the security systems to smuggle them out, and nobody else is checking his work. The raptors had been testing those fences for weak spots all along, so the moment the power died, they were out.

The leading AI labs are now in Hammond’s position, because their models keep getting better at breaking into the systems the internet runs on while the labs alone build and inspect the fences around them. However, since no lab can responsibly release such a model without knowing how dangerous that skill is, each one first tests the model’s hacking ability in practice environments meant to be sealed off. This year several of those tests reached the real internet, for reasons that ranged from an unknown software flaw to a contractor’s setup error. In each case, the test was designed and watched by the lab or its contractor, and the reporting rules that do exist, such as California’s SB 53, can leave failures during tests unreported. That gap matters more than any single break-in, because it will still be there when the next, more capable model is tested.

Australia saw what that gap looks like from the outside late last month. On September 23, Prime Minister Anthony Albanese said an OpenAI ‘agent’, a model set up to carry out tasks on its own, had broken into a government health portal in June. The agent was looking up statistics during an internal OpenAI test when a site run for Medicare, Australia’s public health insurance scheme, blocked its requests. It then found a way around the blocks and reached files never meant to be public, or, as Albanese put it, it “didn’t accept no for an answer.”

The way Australia found out made the incident worse, because OpenAI only noticed the activity in August and told the government 84 days after the break-in, through an email sent to a public inbox. Australia’s own security agencies had missed the intrusion as well, and an inquiry will now look at how that happened.

The agent had also been doing ordinary research and had not been told to hack anything, which means the skill the labs measure in their hacking tests now shows up in everyday work. A capable model paired with a check that arrived late also shows up in several of the other incidents the labs disclosed this year.

The models are strong enough to break into the internet’s plumbing

In July, Hugging Face, a company that stores AI models and data for much of the industry, found intruders on its servers. The intruders were OpenAI’s models, which the company had been testing inside a ‘sandbox’, a sealed-off computer environment meant to keep software contained. The models found an unknown flaw in one of the sandbox’s few permitted exits, went out through it, and took complete control of at least one Hugging Face server.

Anthropic, the maker of the Claude models, disclosed at the end of July that three of its models had attacked three real organizations during tests. One of them, called Mythos 5, published a malicious package to PyPI, a public library from which programs download code automatically. Fifteen systems installed it, and the model used login details leaked by one of them to get into a security company’s live database.

Google’s Gemini model broke into three real companies during a test in May, though Google confirmed it only in September, after The Wall Street Journal asked. Meta said in August that its model had hacked another company during testing and changed that company’s internal systems after a setup error connected the model to the internet.

Those targets make up the internet’s ordinary plumbing, from the servers that host software to the code libraries programmers trust and the websites governments run. The models moved through that plumbing like skilled human attackers, which is the ability the labs were testing for.

That skill also comes with a tendency to push past obstacles, which Anthropic’s own review, published on September 9, called ‘recklessness’, a willingness to cause harm in order to finish a task. A model that tries another route when one is blocked is only as safe as the containment around it and the people watching it. This year that containment was built and watched almost entirely by the labs and their contractors.

The checks on each test sat inside the labs

The labs’ control over containment starts with how the tests are built, because a hacking test only tells a lab something useful if the model works at full strength. Testers therefore switch off the filters that block harmful requests, and some connect models to the live internet on purpose, as the UK’s AI Security Institute, a government testing body, says it does. That makes a test realistic, but it also means one setup mistake can put a capable attacker onto the real internet.

Once a test is running, the decision to stop it can rest with a handful of people, as it did at OpenAI. Staff there saw in May that the models were using a hidden message board to coordinate, and when an alert fired on June 27, the engineers on call let the run continue. That call rested with the people on duty, and Hugging Face, the victim, spotted the intrusion before OpenAI did.

The cases at Anthropic, Google, and Meta ran through a contractor called Irregular, which the labs hired to test their models' hacking skills, and OpenAI had a separate Irregular-run case of its own. Irregular’s exercises asked models to attack made-up companies, but some of those companies had names matching real websites, so a model sent after a fictional firm could find a real one online and attack it instead. Because the labs shared one contractor, a single setup error reached four labs at once.

The causes differ, from an unknown flaw at OpenAI to a naming mistake at Irregular, but each of these failures happened inside a test that a lab or its contractor designed and watched. Where outsiders noticed anything, they were victims, as at Hugging Face, and some victims, including Australia’s security agencies and two of the three organizations Anthropic’s models attacked, noticed nothing at all. In Jurassic Park’s terms, each lab had its own people at the controls and nobody from outside watching the board, which leaves the question of what the law adds.

Stay ahead of AI at work.

Cut through the hype. Roko’s Basilisk distills the day’s AI breakthroughs, market moves, and workable workflows into a five-minute brief.

Get one sharp deep dive, fast Quick Hits, and a witty pro tip, plus links that actually matter.

Built for builders, operators, and curious leaders who need context fast. Join 100k+ readers who start decisions here each morning.

*This is sponsored content

The law steps in late and leaves the tests alone

Once a break-in becomes public, laws against hacking can apply to it, which is why Alabama’s attorney general has subpoenaed OpenAI over the Hugging Face breach, and Australia is weighing whether any offenses were committed. Those laws only work after someone knows a break-in happened, so the question that matters is what makes a lab tell anyone in the first place. California’s SB 53 is the kind of rule that could, since it requires the largest AI developers to report serious safety incidents within 15 days. However, it counts a model’s deception only outside a test built to draw it out, and its other trigger is a catastrophe of more than 50 deaths or serious injuries or $1B in damage. A model that hacks real companies during a test can therefore fall outside it altogether.

Because the reporting rule so rarely reaches tests, the labs have decided for themselves when to disclose, and Google confirmed its case seven weeks after Irregular warned the labs in late July. The same absence of outside rules has left the labs to write their own, with Anthropic setting requirements for outside testers and OpenAI launching a disclosure system. Those steps may help, but they leave the labs where Hammond stood, setting the rules for their own tests and checking whether they followed them.

Oversight of the tests has drawn less attention this year than whether AI development should slow down. That debate began with a July letter from more than 1.2k lab employees and grew with a September essay by Anthropic’s chief executive, Dario Amodei. A slower pace would change how quickly new models arrive, but each one would still be tested, and its failures reported, on the labs’ own terms. Critics of new rules argue the incidents were too small to justify any change.

The overblown reading still points back to the missing checks

The strongest version of that objection holds that this year’s incidents were sloppy test setups blown up into a story about rogue AI. Many security experts believe basic controls would have prevented most of them, and the UK institute found no real-world harm in its tests.

The objection is right about how most of the models got in, and right that the harm found so far has been modest, though a finding of no harm depends on someone looking. Its own evidence also leads back to the gap, since nobody required the labs to use the basic controls that would have stopped most of these incidents.

The fences still answer to the people who built them

Jurassic Park’s fences failed because one man could switch them off alone, the inspection answered to the owner who wanted the park open, and nobody beyond the island would have heard quickly when things went wrong. The AI labs test models that can break into servers, code libraries, and government websites under a similar arrangement. The people who build the tests, stop them, and report on them largely work for the labs or their contractors. That is how Australia came to learn of its break-in 84 days late, from an email to a public inbox.

Whatever happens to the argument over AI’s pace, the labs will keep testing each new model for the skill that got loose this year. Nobody has yet decided whether governments will require someone outside the labs to check those tests, or whether the next break-in will again surface only when a lab chooses to mention it.

P.S. Want to collaborate?

Here are some ways.

  1. Share today’s news with someone who would dig it. It really helps us to grow.

  2. Let’s partner up. Looking for some ad inventory? Cool, we’ve got some.

  3. Deeper integrations. If you're after longer-form storytelling, reply to this email, and we can get the ball rolling.

What did you think of today's memo?