Perpetuity

500 questions94.0LongMemEval-S

keeps your coding agent from summarizing away your session
while using fewer tokens

Agents have

Amnesia

Sessions end.Compaction deletes.Summarized memory is lossy.

Agent context is

Expensive

The whole conversation history is re-sent every turn.Most of it is not needed for the current prompt.

Reasoning

Degrades

with context

The same agent scores lower with the full conversation history than with a retrieved frame.

Every message

Kept

For 99.1% of questions the evidence reaches the agent within one recall call, for 97.4% already in the first frame.Measured on the 470 questions that have evidence.

Your budget goes

Up to 2.8×further

52% less spent on tokens in a 120-prompt Claude Code test, the same agent side by side.64% less between prompts 61 and 90.

Reasoning

94.0

from a 3.6k-token frame

LongMemEval.Opus 5 with thinking, the official GPT-4o judge, 500 questions, nothing excluded.Half of them answered in one round.

92.2at 82M-token context

holding 500 users at once

Same agent and judge, nothing excluded. We know of no published memory system that has run this benchmark at that size. At that size it is a test of retrieval, since Perpetuity ships one memory per agent.

Token counts use Claude's tokenizer, the unit the bill is in.

We list every run and its caveats on the benchmarks page, and we share the per-question outputs on request.

What Perpetuity is

Perpetuity is a proxy that sits between your agent harness and its model. You change one base URL, and your agent runs through it.

Context that never runs out

Each turn, the model gets a frame of only the relevant context, about ten thousand tokens in a coding session, instead of the whole conversation history. The agent can search its memory when it needs more.

Nothing thrown away

No model runs inside the memory loop, and nothing is extracted or summarized. A better retriever tomorrow can read more from the same memory.

measured in Claude Code sessions, the same agent side by side

52%
less spent on tokens over a 120-prompt coding session, $50.24 against $104.71
84%
of questions about earlier work answered correctly, 64.5 of 77 over 5 sessions, against 80% for Claude Code on its own after compacting its history
45k
tokens read per call by prompt 60, most of it Claude Code's own system prompt and tools. Without Perpetuity, 739k.

More about Perpetuity

Perpetuity is out soon.

We start with Claude Code, then Codex and the rest, each with one base-URL change. Leave your email and we will tell you the day it opens.

NeuraWeave

Get notified when Perpetuity is out

Perpetuity gives an agent context that never runs out, and cuts what each turn costs. You change the base URL in your harness and add your key.

Roadmap: Claude Code first, then Codex, then the rest.

We will email you when it ships, and now and then with an update. Privacy