
What is a computer use agent?
It is an AI that runs software the way you do, by looking at the screen and moving the mouse. That works more often than you would expect, and it breaks in ways worth understanding before you hand one your accounts.
A computer use agent is an AI model that operates a computer through its screen. It takes a screenshot, decides what to click or type, does it, and takes another screenshot to see what happened. It repeats that loop until the job is done or it gets stuck. The point is that you stop doing the clicking, even in software that was never built to be automated.
- A computer use agent is an AI model that controls a computer by reading screenshots and sending mouse and keyboard actions, one step at a time.
- The loop is simple: look at the screen, pick an action, perform it, look again, and stop when the task is finished or a step budget runs out.
- Computer use agents are most useful in software that has no API, because the screen is the only interface those programs offer.
- The main risks are slow execution, misplaced clicks, and prompt injection from text on web pages or images, which is why Anthropic tells developers to run computer use in an isolated environment with limited access.
- Reading pixels is a workaround for interfaces built for people, and an operating system built for an AI user can expose structured state directly and keep screenshots as a fallback.
What a computer use agent is
Most AI tools you have used answer in text. You ask, it writes, and then you go do the work. A computer use agent does the work itself. It gets access to a desktop, usually a virtual one, and drives that desktop with the same inputs a person uses: clicks, keystrokes, scrolling, and dragging.
The model never touches the computer directly. A program around it captures the screen, hands the image to the model, receives an instruction such as "click at 512, 742," and carries it out. Anthropic's computer use documentation describes exactly this split. The model supplies the decisions, and your application runs every action inside an environment you control.
That is why the category is called computer use rather than chat. The output of a good run is a changed computer, such as a submitted form or an updated record.
How the screenshot loop works
Every computer use agent runs some version of the same loop. It captures the screen, reasons about what it sees and what the goal needs next, sends one or more actions, and then captures the screen again to check the result.
Anthropic calls the repetition of those steps without human input the agent loop. The current Claude toolset lets the model send a short batch in one turn, for example a click, some typing, and a fresh screenshot. Each batch still ends the same way, with the agent looking at pixels to find out whether its plan worked.
Two details explain most of the behavior you will see. Coordinates are measured in screenshot pixels, so any mismatch between the screenshot size and the real display sends clicks to the wrong place. And the agent only knows what the last screenshot showed. If a pop-up, a cookie banner, or a slow page load changes the screen between steps, the agent has to notice and recover on its own.
What computer use agents are good at
The strongest case for computer use is coverage, because a lot of real work lives in software with no API, including old internal tools, vendor portals, desktop apps, and websites nobody bothered to integrate with. A person gets through those by looking and clicking, and so does a computer use agent.
That makes them a reasonable fit for repetitive, well defined jobs where speed does not matter much. Moving data between two systems that do not talk to each other, filling out the same portal every week, and checking that a web flow still works after a release are all sensible places to start.
They are a poor fit for anything where a single wrong click is expensive, or where the task needs hours of unbroken attention. We cover that second problem in why AI agents fail at long tasks.
Where computer use agents break
Anthropic's documentation is unusually blunt about the limits. It lists latency first and says computer use can be too slow compared with a person doing the same actions, so it should go where speed is not critical. It also warns that the model can make mistakes or hallucinate when it outputs coordinates, that scrolling and dropdowns can be unreliable, and that spreadsheets may take several attempts.
The bigger problem is trust. An agent that reads the screen also reads whatever is on the screen, including text written by strangers. Anthropic states that in some circumstances Claude will follow commands found in content even when they conflict with your instructions, such as instructions on web pages or inside images. OWASP ranks this class of attack, prompt injection, as the top risk for applications built on language models, and its indirect form is exactly what a browsing agent runs into: hidden instructions in a page or file the agent was asked to read.
Taken together, the usual failure is ordinary: the agent is working through an interface built for human eyes, with partial information, while reading content it cannot fully trust.
How to run one without getting burned
The vendor guidance is a good place to start, because it comes from the people with the most reason to make the product look safe. Anthropic recommends a dedicated virtual machine or container with minimal privileges, keeping sensitive data such as login details away from the model, limiting internet access to an allowlist of domains, and asking a human to confirm anything with real consequences, such as financial transactions or agreeing to terms of service.
In practice that means giving the agent its own environment and its own accounts instead of your personal desktop. Start with read-only tasks. Keep a log of every action and screenshot so you can see what happened when something goes wrong. Decide in advance which actions always need your approval, which we go into in how much access an AI agent should have.
Why pixels are the wrong default
Step back and look at what the screenshot loop is doing. A model that reads structured data perfectly well is shown a picture of a user interface, asked to find a button by its position, and told to move a pointer to it. Every layer of that exists because the computer assumes a person is sitting in front of it.
That assumption is the thing we think should change. When the user of a computer is purely AI, the system can hand the agent the actual state of its browser, files, and accounts, and accept typed actions in place of simulated clicks. We made the longer argument in what changes when the computer's user is purely AI and why an AI operating system needs a native browser.
ERIKA is the operating system eFreedom is building on that premise. Its primary control path is structured system state and typed actions, and screenshot control stays available as a fallback for legacy software that offers nothing better. Computer use agents show that models can operate a computer built for people, and ERIKA drops the assumption that a person is looking at the screen.