All posts
Blog

AI can use your computer now. It still can't do your job.

Frontier models are climbing computer-use benchmarks fast. But the two things that decide whether a human gets replaced — cost per finished task, and how long it takes — are one of them badly unfavourable and the other not measured at all.

Deep Chokshi·July 26, 2026·13 min read·
Computer UseAI AgentsBenchmarksOSWorldEvaluationAI Economics

Every few months a new model lands and someone tells me the office job is over.

TL;DR — A new benchmark called OSWorld 2.0 gives AI agents real computer work: 108 tasks that take a skilled human about 1.6 hours each. The best model finishes about 1 in 5 of them, at roughly $72 a task — which works out to around $350 for every task it actually completes. On the longest tasks, every model tested finishes exactly zero. Nobody has published how long any of this takes in wall-clock minutes, on any model, anywhere. The scores that get quoted in headlines are mostly partial credit, not finished work. The capability is real and improving fast. The replacement isn't close.

I've spent the last while reading the actual papers behind the headlines instead of the headlines. This post is what I found, in plain language. No RL background needed.

First, what a "computer-use agent" actually is

Most AI you've used works through text. You type, it types back.

A computer-use agent is different. It gets a screenshot of a desktop — an actual Ubuntu machine with a browser, a spreadsheet, a chat app — and it can move the mouse, click, type, and run terminal commands. Then it takes another screenshot to see what happened, and decides what to do next. Over and over, hundreds of times, until the job is done or it gives up.

That's the pitch that gets people excited: it doesn't need a special integration with your software. It just uses the software, the way you do. Point it at anything.

The exam these agents are graded on

To know whether that works, somebody has to test it. That's what a benchmark is — a standardised exam for models, so different labs can compare results honestly.

The one everyone is quoting right now is OSWorld 2.0, published in June 2026 by a research group at the University of Hong Kong with a dozen collaborators. It's worth understanding why it exists.

The previous version of this exam had become too easy. Frontier models were scoring around 83% on it, which sounds like "computer use is basically solved." But the researchers pointed out the tasks behind that number were short and narrow — a couple of minutes each, usually one application.

So they built a harder exam. 108 tasks, and this is the number that matters:

The median task takes a skilled human about 1.6 hours.

These aren't "rename a file" tasks. They're end-to-end workflows across multiple applications — pull data from an email, cross-reference it against a web portal, update a spreadsheet, file the result. About two-thirds need at least two different apps. Roughly 70% are estimated at over an hour of human work.

That's the point. It's the first computer-use exam where the tasks look like actual jobs.

The same models that scored ~83% on the easy version score 20.6% on this one.

That gap — same models, same week, harder tasks — is the entire story of this post. The exam wasn't wrong before. It was just measuring something much shorter than a real job.

The partial credit trap

Here's where most of the public confusion lives, and it's worth slowing down for.

The exam scores each task two different ways.

Partial credit. Each task is broken into about 27 checkpoints — small verifiable steps along the way. Did you open the right file? Did you find the right record? Did you enter the correct value in the correct field? You get credit for each one you hit.

Completion. Did you finish the whole thing? This is all-or-nothing. Every checkpoint, or it doesn't count.

The researchers are explicit that completion is the primary metric — it's the headline number of the paper.

Now, why does the difference matter so much? Think about booking a flight for your boss. You search the routes, compare the fares, pick the seats, fill in the passport details, apply the corporate discount code — and then you never click Pay.

On a partial-credit scale you did great. Maybe 90%. In reality your boss is not on a plane. The value of that work is zero, and worse than zero, because someone now has to check what you did and finish it.

Almost all real work is like this. The last step is the one that carries the value.

So when you see a computer-use score in a launch announcement, the first question is which of those two numbers you're looking at. Usually it's partial credit. The number that would tell you whether a job got done is the other one, and it's much lower and much less often published.

Here's how far apart they are, on the same runs:

ModelFinished the taskPartial credit
Claude Opus 4.820.6%54.8%
Claude Opus 4.718.2%48.9%
GPT-5.513.0%49.5%
Claude Sonnet 4.68.3%41.5%
MiniMax M34.6%22.3%
Kimi 2.64.6%22.1%
Qwen 3.7-Plus2.8%21.5%

Roughly half the checkpoints, roughly a fifth of the jobs. That's the shape of computer use in 2026.

One quick note on how I'm using these numbers, because it matters for the argument. Vendor announcements generally report on the partial-credit scale. Claude Opus 5 launched in July at 70.57% on this benchmark — a genuine improvement over Opus 4.8's 55.7 on the same scale, and a real engineering achievement. But it sits on the partial scale, not the completion scale, and no lab has published a completion number for the newest generation of models. So throughout this post I use the completion figures that do exist. They're older, and today's models are better than them. They're also the only numbers that answer the question "did the work get done."

What it costs to finish one task

Now the money, which is measured and unambiguous.

The paper reports what each run cost in API fees. For the best-performing setup, about $72 per task attempted.

But you don't pay per attempt and get value per attempt. You pay for every attempt and get value only from the ones that finish. So the number that matters is cost divided by completion rate.

The following column is my arithmetic on the paper's published figures, not something the paper states:

ModelCost per attemptFinishedCost per finished task
Claude Opus 4.8$72.4020.6%~$351
Claude Sonnet 4.6$22.308.3%~$269
GPT-5.5$25.5013.0%~$196
Claude Opus 4.7$33.6018.2%~$185
MiniMax M3$2.404.6%~$52

Two things jump out.

The frontier models are the expensive ones per unit of finished work. The most capable model on the list is also the priciest way to get a task done — roughly $351 against a task that takes a skilled person about 1.6 hours. Depending on what you pay that person, you're somewhere between break-even and three or four times more expensive, for a system that fails four times out of five.

The cheap model is the cheapest per completion, by a lot. MiniMax M3 finishes only 4.6% of tasks, but it's so much cheaper that each completion costs about $52. That inverts the usual assumption that frontier models are the efficient choice. If you genuinely have a task you can attempt repeatedly and verify cheaply, the small model may be better economics than the big one.

The number nobody is reporting

Now the speed half, and this is the finding that surprised me most.

There is no wall-clock time in this benchmark. For any model. Anywhere.

Not in the paper, not on the leaderboard, not in the code repository, not on the project website, not in a single vendor's model card. The exam reports steps, turns, tokens and dollars. It does not report minutes.

Think about how strange that is. The core commercial promise of this technology is that it's faster and cheaper than a person. The industry has built an elaborate exam to measure it — and the exam doesn't have a stopwatch.

I want to be careful here, because it would be easy to fill that silence with a scary number, and I'm not going to. I can't tell you an agent takes three hours on a task a human does in ninety minutes. Nobody can, because nobody has published it.

What I can tell you is the floor. The test harness inserts a three-second pause after every action. The top setups take several hundred actions per task. That alone is roughly 24 to 30 minutes of pure waiting, before the model has thought about anything.

And from adjacent research where people did measure time: a study on the older, easier version of this benchmark found agents "practically unusable due to extremely high end-to-end latency (e.g., tens of minutes)" on tasks that take humans a few minutes. Their example: twelve minutes for an agent to double-space two paragraphs, against under thirty seconds for a person.

That same study also punctures the most common explanation for why agents are slow. People assume it's the screenshots — all those images going back and forth. It isn't. When they broke down where the time actually goes, screenshot capture was under 2% of it. Planning and reasoning were 76 to 96%.

The agent isn't slow because it's looking at your screen. It's slow because it's thinking, several hundred times, in sequence.

"Just give it APIs instead"

This is the objection I hear immediately, and it's a good one, so let me deal with it properly.

The argument goes: clicking around a screen is a stupid way to work. Give the model a proper software interface — an API — and let it call functions directly. Faster, cheaper, more reliable. The pixels are the problem.

I believed some version of this before I read the trajectory data. Two things changed my mind.

The agents already do this. The researchers annotated how each model actually solved things. GPT-5.5 solved 71.3% of tasks primarily by writing code, calling APIs, or manipulating files directly — and only 4.6% by pure clicking. Across all models, terminal commands make up nearly as many actions as mouse clicks. Nobody had to propose this architecture. The models discovered it on their own and it's already what they mostly do.

And it doesn't rescue them. GPT-5.5, the most API-happy model in the study, finishes 13% of tasks — lower than the models that click more. Meanwhile Zapier built a benchmark that gives models nothing but clean APIs across 40 business applications, no screens at all. Every model scored under 10%.

If the interface were the bottleneck, that test should be nearly solved. It isn't.

So what actually goes wrong?

The researchers catalogued the failures, and this is the part I'd frame on a wall:

Rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification.

Read that again. The models are not failing because clicking is hard, or because coding is hard. They're good at both.

They fail because on a 1.6-hour job with dozens of interlocking requirements, they forget a constraint they were given at the start. Something changes halfway through — a new message arrives, a value updates — and they don't notice. They hit an ambiguity and guess instead of asking. And then they don't check their own work.

That last one has a number attached, and it's the most telling statistic in the paper: agents spend under 7% of their effort detecting and repairing errors, even though errors are what's killing them.

Here's the clearest way to see what that means in practice. The researchers sorted tasks by how long a human would take, and looked at completion:

  • Shorter tasks: around 20–24% finished
  • Middle: about 5%
  • Tasks over roughly 2.7 hours of human work: 0.0%. Every model. Zero.

And partial credit on those same long tasks stays near 50%.

That's the whole thing in one line. On the longest, most valuable, most job-like work, agents get halfway through and finish none of it.

Not "finish it slowly." Not "finish it expensively." Don't finish.

Why this won't close as fast as the chart suggests

The fair objection to everything above is that it ages badly. Models improve monthly. Write this post in a year and the numbers look silly.

That's a real risk and I want to name it honestly. AI capability has been doubling on a roughly three-month cadence by some measures. Inference prices have historically fallen fast. I'm not going to pretend that stops.

But there are three specific reasons this particular gap is stickier than the score curve implies.

Getting better gets exponentially more expensive. The researchers measured the cost of each additional accuracy point, and found it "rises by roughly an order of magnitude as agents approach the ceiling." Roughly 25,000–30,000 extra words of model output per single point of improvement, and climbing. The last stretch of reliability is the most expensive stretch there is.

The frontier price hasn't actually been falling. Flagship pricing has sat at $5 per million input tokens and $25 per million output for about eight months, across three model generations. Cheaper models keep arriving underneath, which is real progress — but the "prices always collapse" assumption is not currently true at the top tier.

Completion is a harder curve than capability. Partial credit rises smoothly as models get better. Completion doesn't, because it requires getting everything right. A model that improves from 50% to 70% on checkpoints might barely move on completion, because the failures are concentrated in the parts that need the most care — verification, memory over long spans, knowing when to ask.

The part that actually blocks replacement

Suppose all of that resolves. Suppose a model finishes 80% of these tasks at $10 each, fast. Is the job gone?

Not automatically, and this is the piece most cost comparisons skip.

OpenAI's GDPval study is the most careful work I've found on this. They took real professional deliverables, had models produce them, and had experts review the output. Compared naively — tokens against salary — models looked up to 5,172× cheaper.

Then they priced the review. Someone has to check the work, catch what's wrong, and fix it. Once that's included, the advantage collapses to somewhere between 1.18× and 1.63× for their best model. For a weaker model it goes the wrong way entirely — using it costs more than just having the person do it.

And that's on top of what human expertise actually costs there: about 6.7 hours and $361 per task, plus 109 minutes and $86 of expert review for each model attempt.

So at a 20% completion rate, the human doesn't leave. They stop doing the work and start verifying it — which is a different job, often a more tedious one, and not obviously less of their day.

That's what "replacement" runs into. Not capability. Trust. You can only remove the person once you can predict which outputs are wrong without checking, and right now nobody can, because the models themselves can't.

What would change my mind

I'd rather be specific than vague, so here's what I'm watching. Any of these would move me:

  1. A lab publishes wall-clock time for a computer-use benchmark, and it's competitive with a human on the same task.
  2. Completion rates above 50% on long-horizon tasks — the actual finished-work number, not partial credit.
  3. Anything above zero on the longest bucket of tasks. Right now that's a hard floor, and the first model to crack it is telling you something structural has changed.
  4. Agents that reliably stop and ask instead of guessing, and that spend real effort checking their own work. Under 7% is the number to beat.
  5. Frontier price actually falling, rather than cheaper models arriving below a fixed frontier price.

The takeaway

The capability is real. Models genuinely can operate a computer now, and that was science fiction three years ago. The improvement between generations is not marketing — it's measurable and it's fast.

But there are two different curves here, and only one of them is going up quickly.

Capability — can the model do the kind of thing the job requires — is climbing fast, and that's what benchmark headlines measure.

Unsupervised completion — will it finish the whole job, correctly, without someone checking — is climbing much more slowly, costs exponentially more at the top, and hits exactly zero on the longest tasks.

Replacement needs the second curve. Almost everything you read about is the first one.

So when the next model lands and the chart goes up and to the right, ask two questions. Is that number finished work, or partial credit? And how long did it take?

Right now, for the second question, nobody has published an answer at all.