A twisting paper ribbon torn through the middle

Why AI agents fail at long tasks

Agents complete short tasks reliably and fail more often as tasks stretch into hours, mostly because of where they keep track of their own work.

AI agents fail at long tasks because every extra step is another chance to go wrong, and most agents keep all of their progress inside a context window that was never meant to hold hours of work. A model that is excellent at each individual step can still fall apart over a sequence of a few hundred of them.

The short version
  • METR's March 2025 study found frontier agents of that period succeeded almost every time on tasks that take a skilled person under four minutes and less than 10 percent of the time on tasks that take more than about four hours.
  • The length of task agents can complete with 50 percent reliability roughly doubled every seven months over six years, according to the same METR research, so the ceiling is rising quickly.
  • Long tasks break agents mostly through compounding small errors, unchecked assumptions about whether an action worked, and progress that lives only in the context window.
  • An agent that works while you sleep needs durable task state, checkpoints, a record of every action, and the ability to resume after a crash, and none of that comes from the model.
  • Continuity of work is a property of the system the agent runs on, which is why we treat it as an operating system problem.

What the research measured

The clearest data comes from METR, a research nonprofit that evaluates AI systems. In a March 2025 study, METR timed skilled people on a large set of software and reasoning tasks, then tested how often AI agents could complete the same tasks on their own. How long a task took a human turned out to predict agent success very well. Models of that period succeeded almost every time on tasks that took people less than four minutes and less than 10 percent of the time on tasks that took more than around four hours.

METR also tracked how that changes over time. The length of task a frontier agent can finish with 50 percent reliability had been doubling roughly every seven months for six years. METR has since flagged some figures in the original post as out of date and publishes updated measurements, and those measurements still show the length of task agents can complete rising over time.

The most useful line in the study is the diagnosis. METR wrote that agents often seem to struggle with stringing together longer sequences of actions more than they lack the skills or knowledge needed to solve single steps. That gap is what this post is about.

Small errors compound

Picture a task with 200 steps where the agent gets each one right 99 percent of the time. The chance of getting through all of them without a single mistake is about 13 percent. Real tasks are more forgiving than that, since an error can be noticed and fixed, but the arithmetic explains why a one-hour job is so much harder for an agent than a one-minute job.

Recovery depends on the agent checking its own work, and it often does not. Anthropic's computer use documentation notes that Claude sometimes assumes the outcome of an action without explicitly checking it, and recommends prompting the model to take a screenshot after each step and confirm it worked. An agent that assumes a payment went through, without looking, will build the next ten steps on a guess.

The agent keeps its progress in the wrong place

Most agents track a long task the same way they track a conversation. The plan, the steps completed, the results so far, and the open problems all live as text in the context window. As the task grows, that text grows, and the agent has to find the right detail in an ever longer pile.

Language models are not great at that. The Lost in the Middle study found that models use information in the middle of a long input noticeably worse than information at the start or the end. An instruction from hour one can effectively fade by hour three. We cover the broader version of this in why AI agents forget everything between sessions.

Then there are interruptions. A crashed process, an expired login, a dropped connection, or a restart wipes a context window completely. If the only record of progress was in that window, the agent either starts over or, worse, repeats steps it already took.

What an agent needs to work while you sleep

The whole pitch for agents is work that gets done without you sitting there. For that to hold on anything longer than a few minutes, the agent needs things the model cannot provide on its own.

It needs task state written somewhere durable, outside the context window, so a restart picks up where the work stopped. It needs checkpoints it can return to when a step fails, and a record of every action and its result, so it can check what happened instead of trusting its memory of what it meant to do. It needs to know which actions are safe to retry and which must never run twice, such as sending a message or making a payment. It also needs clear rules about which decisions it can make alone and which wait for you, which we cover in how much access an AI agent should have.

Agent frameworks add pieces of this today. They sit on top of an operating system that has no idea an agent is running, so every piece has to be rebuilt and kept alive by whoever set it up.

Why this is an operating system problem

An operating system already handles this for ordinary programs. It keeps files when a program crashes, tracks what is running, and decides what each process may touch. What it does not do is treat an AI agent as the user whose work needs protecting.

ERIKA starts from that assumption. One running instance holds one persistent agent, and that agent's goals, work, and evidence belong to the operating system rather than to a single model call. Actions leave records, and failed work can be inspected, resumed, or rolled back. We explain the identity side in one operating system, one persistent agent and the reason behind all of it in why computer work should become optional.

METR's trend shows models finishing longer tasks over time, but the durable state, checkpoints, and action records described above still have to come from the system underneath.

Max MedawarFounder of eFreedom. Building ERIKA, an operating system that keeps an agent's work intact across failures.