Agents have
Amnesia
Sessions end.Compaction deletes.Summarized memory is lossy.
Agent context is
Expensive
The whole conversation history is re-sent every turn.Most of it is not needed for the current prompt.
Reasoning
Degrades
with context
The same agent scores lower with the full conversation history than with a retrieved frame.
Every message
Kept
For 99.1% of questions the evidence reaches the agent within one recall call, for 97.4% already in the first frame.Measured on the 470 questions that have evidence.
Your budget goes
Up to 2.8×further
52% less spent on tokens in a 120-prompt Claude Code test, the same agent side by side.64% less between prompts 61 and 90.
Reasoning
94.0
from a 3.6k-token frame
LongMemEval.Opus 5 with thinking, the official GPT-4o judge, 500 questions, nothing excluded.Half of them answered in one round.