A scribbled note whose lower edge dissolves into loose halftone dots

Why AI agents forget everything between sessions

The model behind your agent keeps nothing after it answers. Everything that looks like memory is bolted on around it, and where that memory lives decides whether your agent learns anything at all.

AI agents forget between sessions because the language model inside them is stateless. Each request is processed on its own, and when the response is done, nothing about you or your work stays inside the model. Every bit of continuity you have seen from an AI tool comes from software around the model that saves information and feeds it back in on the next request.

The short version
  • Language models do not remember past sessions on their own, because each request is processed without any stored state from earlier requests.
  • A context window is working space for a single request, and it is cleared when the session ends unless another system saves what mattered.
  • Bigger context windows do not solve forgetting, because the Lost in the Middle study found models use information in the middle of long inputs noticeably worse than information at the start or end.
  • Most agent memory today is saved notes or chat history that gets retrieved and pasted back into the prompt, which helps but carries no built-in rules for provenance, permissions, or expiry.
  • Memory that has to outlive a model, an app, or a restart belongs in the layer underneath them, which is why we argue it belongs to the operating system.

The model does not remember you

When you talk to an AI agent, the model receives a block of text, produces a response, and stops. There is no running process inside the model that keeps your name, your project, or yesterday's decisions. The next request starts from the model's training and whatever text arrives with it.

Inside a single session this is easy to miss, because the application quietly resends the conversation so far with every message. The model looks like it is following along. It is rereading the whole transcript each time. End the session, clear the chat, or switch tools, and that transcript is gone from its view.

So when an agent forgets something, the model did not lose it. The software around the model either never saved it, saved it somewhere the next session cannot reach, or failed to bring it back at the right moment.

Why a bigger context window does not fix it

The obvious fix is to make the window bigger and stuff everything in. Context windows have grown enormously, and they help within a session. They do nothing for the next session unless something saves the contents and loads them again, and you pay for every token you resend.

They also do not guarantee the model uses what it is given. In Lost in the Middle, researchers from Stanford and elsewhere tested how well language models find relevant information in long inputs. Performance was highest when the answer sat at the beginning or end of the input and dropped significantly when it sat in the middle, even for models built for long contexts. In one setup, GPT-3.5-Turbo did worse with the answer buried in the middle of its documents than it did with no documents at all.

A bigger window gives the agent a bigger desk. Piling a year of work onto that desk makes the one page that matters harder to find, and the desk still gets cleared when the session ends.

What today's memory fixes do

Every memory feature you have seen is a variation on saving text outside the model and putting some of it back later. Chat history keeps the raw transcript. Summaries compress older conversation into a shorter note. Retrieval systems store facts or documents in a database, often a vector index, and pull back the pieces that look most relevant to the current request.

The more ambitious designs borrow directly from operating systems. The MemGPT paper from UC Berkeley treated the context window like main memory and external storage like disk, and let the model move information between the two with function calls, the way an operating system pages data in and out. On questions about earlier conversations, GPT-4 answered 32.1 percent correctly from a summary of past sessions and 92.5 percent when MemGPT managed its memory.

These techniques work, but almost all of them are features of a single app, framework, or model provider's account.

Where bolted-on memory goes wrong

Memory attached to an app inherits every limit of that app. Switch tools and the memory stays behind. Change model providers and you may lose it entirely, which is the ownership problem we wrote about in why agent capability should be owned, not rented.

The quieter problem is that saved text has no rules attached. A note saying a client wants invoices on the first of the month does not record who said it, when it was true, or whether it has since changed. A retrieval system will happily pull back a stale fact next to a current one. Nothing marks which memories the agent may act on and which are sensitive. And because the memory is just more prompt text, anything written into it, including instructions planted in a web page the agent read, can shape its behavior weeks later.

That is how you end up with an agent that forgets the thing you told it twice and confidently remembers something that stopped being true in March.

Where memory should live

Memory that has to survive a restart, a model swap, and a change of tools cannot belong to any one of them. It needs to sit underneath all of them, in the layer that also decides identity, permissions, and what happens when something crashes. On a normal computer that layer is the operating system.

That is the argument behind why memory belongs to the operating system, and it is how ERIKA is being built. In ERIKA, memory carries provenance, permissions, and expiry rather than being a saved transcript pasted into the next prompt, and it stays with the agent when the model underneath changes. The model does the thinking and the operating system keeps the record, so one running instance holds one agent's history across everything it does, as we explain in one operating system, one persistent agent.

If your agent keeps forgetting, a better prompt will rarely fix it. Look at where its memory is stored, who controls that store, and whether anything checks that what it remembers is still true.

Max MedawarFounder of eFreedom. Building ERIKA, an operating system that keeps an agent's memory below the model.