A year ago, in Oct 2025, Andrej Karpathy announced on Dwarkesh’s podcast that agents were far from good yet.
They just don’t work. They don’t have enough intelligence, they’re not multimodal enough, they can’t do computer use and all this stuff.” - Andrej Karpathy, Oct 2025
He gave it about a decade for agents to feel like AGI. “When you get a demo and something works 90% of the time, that’s just the first nine,” he said, referring to the idea of five 9s of reliability needed for agents to feel great.
Well, the next 12 months have seen an onslaught of dramatic improvements. If you are building customer facing agents, you need to deeply understand these and make sure your team is building for the future, not the past.
Agents have moved out of demos into everyone’s working day. Here are the five things that made this happen - longer model horizons, breakthroughs in computer use, high performance tool use, safe-enough inbox access with low prompt injection and better runtime configurations for agents with their own auth and wallets. More on each below.
1. Model horizons got longer
The simplest measure of the year is how long a task you can hand to an agent and expect back done. METR tracks this for software work. It measures the length of task, in human working time, that a model completes 80% of the time. (METR’s own headline numbers are much bigger because they use a 50% success rate but four in five is closer to what a product actually needs). In Oct 2025 the best figure was 38 minutes, for GPT-5. It reached 3 hours 6 minutes in April, for an early version of Claude Mythos Preview.
That is almost five times longer in eight months. A year ago you could trust the best agent with a chore. Now you can trust it with an afternoon of work.
A longer horizon matters this much because errors compound. To finish a 300-step task correctly four times out of five, it has to get each step right 99.93% of the time. Most of the distance between a demo and a product is in those decimal places. 3 nines!
2. Computer use got really good
When Anthropic first shipped it in October 2024, the model scored 14.9% on OSWorld, a test of short desktop tasks that take a person about two minutes each. Anthropic’s launch post called it “at times cumbersome and error-prone.”
By April 2025 the best score was 38%. Claude Opus 4.5 reached 66.3% that November. People score about 72%. By June 2026 Claude Mythos Preview was at 85.4%.
The benchmark’s authors did what benchmark authors do and built a harder one. OSWorld 2.0, published in July, has 108 workflows that take a person a median of 1.6 hours and about 318 tool calls each, ten times the original. On its first outing the best agent earned 54.8% partial credit and completed 20.6% of the workflows.
Then came the fall launches. On September 3, OpenAI released GPT-6 Astra and led with computer use.
OpenAI reported 72.6% partial credit on OSWorld 2.0 at about 40 minutes a task. On Sept 22, Anthropic’s Claude Opus 5.5 posted 81.8%, up from 74.0% for Opus 5 two months earlier. 🤯
3. On-demand tool loading made agents faster
A year ago, connecting an agent to too many tools was making them worse. Every tool came with a definition, and every definition was loaded into the model’s context before the user had typed a word. Anthropic measured a typical setup of five MCP servers (MCP is the standard protocol for connecting agents to tools) at about 55,000 tokens of definitions. The agent got slower and more expensive, and it picked the wrong tool more often.
The fix was to stop loading everything. In Nov 2025 Anthropic let the model search for tools and load only the ones a task needs. That cut the token cost by 85%, and accuracy on its tool-use evals rose from 79.5% to 88.1%. Here is Simon Willison when the same change reached Claude Code in January:
The same idea showed up elsewhere. Agents started writing code that calls tools directly, so intermediate results never pass through the model. In Anthropic’s example that took one task from 150,000 tokens to 2,000. Vercel removed 80% of its data agent’s tools in December and gave it a shell instead to see a 20% improvement in accuracy.
4. Inbox access got safe enough
In Aug 2025 Anthropic described a test in which a malicious email told Claude to delete the user’s emails “for security reasons,” and it did, without asking. Even with safeguards on, 11.2% of such attacks worked.
By July 2026 Anthropic reported zero successes in 129 browser scenarios once it added a filter on incoming content and a check on outgoing actions.
Grok Bot, which launched in beta in August, comes connected to Gmail, Outlook, Google Drive, SharePoint and Salesforce. And customers are using it happy because it’s “safe enough”.
5. We figured out agent runtimes: cloud computers, with logins and wallets
On Oct 22, 2025, the day after ChatGPT Atlas launched, a ZDNET writer asked its agent mode to order wood putty, caulk and screws from Walmart. It made pogress but then got stuck on pop-ups. Her verdict: “half magic, half refining.”
That was a good agent experience a year ago. The agent borrowed your browser and could potentially use your logins, and you did the last few steps yourself. This year the industry settled on a better arrangement: the agent gets its own computer, its own way to log in, and its own way to pay.
The computer came first. Perplexity Computer launched in Feb, Gemini Spark in May, Grok Bot in August, and Meta’s Muse, Manus Cue and OpenAI’s dots all in September. Each gives every user’s agent a persistent machine in the cloud. Grok Bot’s launch thread said it best: “Give them a task, shut your computer, and reach them from anywhere.”
In April, Stripe launched a Link wallet for agents. In July, 1Password shipped a way for Claude to sign in to a site without seeing the password.
The plumbing matured too. The MCP spec now lets a tool server run behind an ordinary load balancer like any other web service.
The pattern across all of these developments is more access to agents but in narrower grants: per task, per purchase, and never the raw secrets themselves.
The rise of agent deployment services
So why haven’t all enterprises gone fully agentic with their workflows? Because their environments and systems are just not set up for it. So the applied agent companies have leaned into custom deployments.
Harvey puts a former practicing lawyer into every deployment, about 180 of them, and its customers now run more than 25,000 custom agents.
Sierra embeds engineers with each customer for weeks and has started shipping more advanced uses cases with Takeoff and Ghostwriter.
AirOps now pairs its agent marketing platform with a team of channel-specific experts - they are taking off like wildfire.
Enterprises need help choosing models, setting up evals, preparing data, and in the “constant tuning of the agentic system for their process.”
When building an agent today
“Can it do this?” is mostly settled. The useful questions are how often it works, and how efficiently can it do it. Measure it on your own traces. Benchmarks tell you what’s possible. Your own success rate, broken out by task length, tells you what you have.
Delete things. Delete all the complex, involved prompts written > 2 months ago. Update your assumptions aggressively. Prompt-audit your skill files. Start a new harness from scratch. Dax said it best for day to day AI users:
The Anthropic team agrees - “every component in a harness encodes an assumption about what the model can’t do on its own.” Boris Cherny’s team cut more than 80% of its system prompt for Opus 5, and he advises deleting your instruction files, skills and hooks every six months to see what the model does with less.
You can’t run that test without evals, which is why we’d treat the eval suite as the most important product asset that lasts. Prompts, tools, harness code and even the runtime all turned over this year. The eval rubric and the labeled traces carry over to new versions of agents.
Karpathy called it the decade of agents. If every year looks like 2026, I’d call that way too slow, we are right at the finish line for most knowledge work.
If you are a Reforge member, we’ll be kicking off our evals course in a couple weeks: check out the pre-reads and videos here!













