
In 2013 Google gave some of its employees a car that could drive itself on the highway. The deal was simple: the car drives, you stay ready to take over.
They didn't stay ready. John Krafcik, who ran Waymo a few years later, told Reuters what the team saw on the cameras: people napping, putting on makeup, and on their phones, at up to 56 mph. "What we found was pretty scary," he said. "It's hard to take over because they have lost contextual awareness."
I think about that story a lot, because it's exactly where agents are right now. The agent can do the work. We tell a human to watch it. And the human stops watching.
That's not a guess. When Anthropic made auto mode the default in Claude Code in August, it published the data. Users approve 97% of permission prompts. In a study with 1,053 paid testers, people reviewing the agent's commands caught 13.6% of the dangerous ones. An automated check caught 89%. The human in the loop, the thing every team adds to make agents safe, was the weakest part of the system.
Self-driving cars have had a ten-year head start on this problem, and they already ran the experiment we're about to run. I think agentic work is going to follow the same path, so it's worth knowing how that path went. I'm not writing this from above it, either. Every InsForge engineer runs Claude Code on their own VPS, read-only across our infrastructure on purpose. That's a safety driver, and I'm the one who put it there.
TL;DR
- Cars and agents share the same five levels, and the question that separates them is the same: who catches the failure.
- Agents are stuck at Level 3, where a human has to stay ready to take over. Humans are bad at that, in cars and with agents.
- More guardrails and approval prompts are more sensors on a car that still needs a driver. They help, but they don't change the level.
- Waymo got to driverless by drawing a geofence, logging real miles inside it, beating human drivers on the same roads, and only then widening it. Agents get there the same way: give every task its own copy of the environment, where a crash is cheap.
Same five levels, same question
The driving levels come from SAE J3016. Put agentic work next to them and they line up almost too well:
| Level | Cars | Agents |
|---|---|---|
| 0 | No automation | Autocomplete. You write, it suggests |
| 1 | Steering or speed | Chat assistant. It answers, you do the work |
| 2 | Steering and speed, driver supervises constantly | Agent acts, you approve every step |
| 3 | Car drives, driver takes over on request | Agent runs, you're the fallback when it goes wrong |
| 4 | No driver needed, inside a limited area | Agent owns the task inside a boundary, you review the result |
| 5 | No driver needed, anywhere | Agents work anywhere, including production, and talk to and hand work to each other. No human in the loop |
The column I care about in the SAE table is the one about who catches the failure. At Level 3 the human is "fallback-ready" and "becomes the driver during fallback." At Level 4 the system handles its own fallback, "without any expectation that a user will need to intervene." For cars, the only difference between 4 and 5 is whether that holds inside a limited area or everywhere. For agents there's one more: at Level 4 a human still reviews the result, but only inside that specific domain instead of across all of them, and at Level 5 the agents check each other's work everywhere.
Other people have drawn agent levels before. Dan Shapiro's five levels run from "spicy autocomplete" to a "dark factory" where nobody reviews code, and a Knight Institute paper defines levels by the human's role. Both are good, and both are about the human. What the car industry teaches is about the road.
Level 3 is where people fall asleep
None of this was a surprise to the people who study automation. Lisanne Bainbridge wrote Ironies of Automation in 1983, about factories and cockpits, and one line has aged perfectly: it's "impossible for even a highly motivated human to maintain effective visual attention towards a source of information on which very little happens, for more than about half an hour."
Google read the camera footage and dropped the handoff idea. It went straight for full autonomy instead. Ford made the same call in 2017 after its own engineers kept falling asleep in test cars, despite buzzers and vibrating seats.
And even an alert driver needs time. In a 2014 simulator study published in Transportation Research Part F, drivers took about 15 seconds to resume control after the automation handed back, and up to 40 seconds before their steering settled. At highway speed, 15 seconds is a long way.
The agent numbers show the same decay, almost to the minute. In Anthropic's study, people blocked about 17% of dangerous commands early in a session and about 5% after 50 or more prompts. That's Bainbridge's half hour, measured on coding agents.
So when someone says an agent isn't safe enough to run without a human, the honest follow-up is: compared to what? The human in that loop is clicking yes.
More sensors won't change the level
The industry's instinct is to add sensors. More approval prompts, more guardrails, better classifiers, more dashboards to watch. Some of that is genuinely good. Auto mode is a better sensor than a tired human, and the numbers above say so.
But it's still Level 3 thinking. The agent drives on the same road as production, and we keep trying to get better at catching it before it hits something.
Here's what that looks like when it goes wrong. In April a coding agent at PocketOS, an automotive software startup, was doing a routine task in staging. It hit a credential problem, went looking for a fix, and found an API token in an unrelated file with far more access than anyone meant it to have. One call deleted the production database, and because the backups lived in the same place, the backups went with it. Nine seconds. Getting the data back took the hosting company's CEO stepping in personally to help.
A more careful reviewer wouldn't have saved that. The agent was in staging. Staging and production were one token apart.
What Waymo did next: it drew a map
SAE has a term for the bounded area a self-driving system is built for: the operational design domain. It covers "environmental, geographical, and time-of-day restrictions." In plain words, a geofence.
Waymo's whole path is that idea applied patiently. In 2018 California gave it a driverless testing permit for five named towns: Mountain View, Sunnyvale, Los Altos, Los Altos Hills and Palo Alto. In 2020 it opened fully driverless rides to the public in a slice of the Phoenix suburbs. Paid driverless rides in San Francisco came in 2023. As of September it's in 14 US cities, each with its own geofence.
Waymo didn't wait for a car that could drive anywhere. It shrank the world until it didn't need a driver, and then grew the world one city at a time.
Trust came from miles, not tests
The geofence is half of it. The other half is how Waymo earned the right to widen it.
In 2016 RAND tried to work out how much driving it would take to prove a self-driving car safer than people. Their answer, in Driving to Safety: more than 11 billion miles to show a fatality rate 20% better than humans. For a fleet of 100 cars driving around the clock, that's about 500 years. You can't pass that test, so RAND concluded companies "cannot simply drive their way to safety."
What worked instead was evidence, piled up inside the boundary and compared against people. Waymo has logged 271.3 million miles with no one at the wheel through June of this year, with 95% fewer serious-injury crashes than human drivers on the same roads. A Swiss Re study of 25.3 million of those miles found 9 property-damage claims where human drivers would have produced about 78. Behind the real miles sit tens of billions of simulated ones and a 113-acre closed course on an old Air Force base.
Those aren't only numbers on Waymo's own website. Its safety researchers published the method in the peer-reviewed journal Traffic Injury Prevention: across the first 56.7 million driverless miles, crash rates were significantly lower than human benchmarks at every injury level they measured, and there was no crash type where Waymo did significantly worse. The preprint is free to read.
Agents get judged the other way around. We grade them on benchmarks, which are the written part of the driving test, and they saturate and leak. In February OpenAI stopped reporting SWE-bench Verified after finding frontier models could reproduce the reference answers from memory. And passing doesn't transfer the way you'd hope. METR's 2025 randomized trial had 16 experienced developers do 246 real tasks. With AI they were 19% slower, while believing they were 20% faster.
The agent version of a mileage log is only starting to exist. Anthropic's autonomy research tracked how often people interrupt Claude Code, about 5% of turns for new users and about 9% for experienced ones, and noted that "many of our findings cannot be observed through pre-deployment testing alone." That's a disengagement report. We need a lot more of them, on real work, compared against the same work with a human approving every step.
Agents already have geofences in two places
Look at where people already let agents run unattended.
Deep research is the first one. People kick off a research agent, walk away for twenty minutes, and read the report. Nobody approves each search. Usually it's the same model as the coding agent, so intelligence isn't the difference. Reading the web is read-only, so the whole internet is inside the geofence.
Code in a git branch is the second. Nobody approves an agent's keystrokes in its own branch. You review the pull request. The mess stays in the branch until someone decides to merge it. Anthropic found most agent actions today are "low-risk and reversible," and that's no accident. Those are the places we gave agents a boundary.
Then the agent needs to run a migration, write to a bucket or deploy a service, and it drops straight back to Level 2. Databases, storage and running services don't branch the way code does, so every write touches the real thing, a human has to approve it, and nobody ever logs miles there.
The models aren't what's holding this back. METR's measure of how long a task agents can finish has been doubling roughly every seven months, and faster since 2024. The method is in METR's 2025 paper, which measures a task by how long it takes a skilled human to do it. Anthropic calls the gap between what models could handle and what people let them do a "deployment overhang." The car can drive. There's no road for it.
Give the agent a second life
So here's the geofence for infrastructure. Every agent task gets its own copy of the whole environment: its own database with real data, its own storage, its own running app at its own URL. Production credentials don't exist inside it. If the agent drops a table, it drops its own table. If it ships something broken, it breaks its own URL. You throw it away in seconds and start over.
That's the second life. The agent can crash and respawn, so nobody has to watch it drive. And because crashes are cheap, you can finally let them happen and count them. Every task becomes a logged mile on realistic data. The human moves from approving fifty steps to reviewing one outcome.
This is what we're building at InstaCloud. insta branch create gives a task a copy-on-write Postgres that opens with the parent's data, a forked bucket, and compute already running the parent's app. The agent has full write access there, insta agent events keeps the log of what it did, and when the work is right you promote it like a pull request. When it isn't, you delete the branch.
Then you widen the geofence the way Waymo did, one kind of task at a time, as the record earns it. It's also how we'll get our own agents off read-only.
Where this goes
Level 5 for agents is production with no boundary at all, where agents hand work to each other and check each other's results with no human in the loop. I don't think we're close to that everywhere, and most of the value is already at Level 4. Waymo has carried a lot of people without ever reaching Level 5.
Level 3 is a stage, the same as it was for cars. The companies that got to driverless didn't build a more attentive human. They built a smaller world, proved the car in it, and grew it.
Agents are in the 2013 car. We know how the next ten years went for cars, so we don't have to take ten.
If you're thinking about this differently, I'd like to hear it. Find me on X or LinkedIn.
Sources
The research this post leans on, if you want to go deeper:
- Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775–779.
- Merat, N., Jamson, A. H., Lai, F. C. H., Daly, M., & Carsten, O. M. J. (2014). Transition to manual: Driver behaviour when resuming control from a highly automated vehicle. Transportation Research Part F, 27, 274–282.
- Kalra, N., & Paddock, S. M. (2016). Driving to Safety: How Many Miles of Driving Would It Take to Demonstrate Autonomous Vehicle Reliability? RAND Corporation.
- SAE International (2021). J3016: Taxonomy and Definitions for Terms Related to Driving Automation Systems.
- Kusano, K. D., Scanlon, J. M., Chen, Y.-H., et al. (2025). Comparison of Waymo Rider-Only crash rates by crash type to human benchmarks at 56.7 million miles. Traffic Injury Prevention, 26.
- Di Lillo, L., Gode, T., Zhou, X., et al. (2024). Do Autonomous Vehicles Outperform Latest-Generation Human-Driven Vehicles? A Comparison to Waymo's Auto Liability Insurance Claims at 25 Million Miles. Waymo and Swiss Re.
- METR (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, a randomized controlled trial.
- Kwa, T., West, B., Becker, J., Deng, A., et al. (2025). Measuring AI Ability to Complete Long Software Tasks. METR.
- Feng, K. J. K., McDonald, D. W., & Zhang, A. X. (2025). Levels of Autonomy for AI Agents. Knight First Amendment Institute.
- Anthropic (2026). Measuring AI agent autonomy in practice and Auto mode is now the default in Claude Code.