The Caveats Became The Result
I recently tried to make a clean comparison of the leading systems on LoCoMo. The first version looked like a normal leaderboard: system, overall score, category scores, retrieval speed, token use. Then the caveats began to take over. One system used GPT-4.1-mini, another used GPT-5. Some reported four non-adversarial categories, while others included the adversarial set. One published search latency, another published total response time, and another reported only the number of retrieved tokens. Even the scores that appeared in the same column often came from different answer models, judge models, prompts, subsets, and retrieval cutoffs.
The comparison was still useful, but the caveats were more revealing than the ranking. We have reached a point where memory systems can report scores above 90% while the field still lacks a shared test for the thing many of those systems ultimately claim to build: an accurate, durable model of one person over time.
Our own Synthius-Mem result belongs inside that caveat. A 94.37% LoCoMo score and 99.55% adversarial score provide useful evidence about structured retrieval under a particular harness. They do not tell us whether the same memory can absorb years of a person’s life, reconcile therapy notes with group chats, reconstruct a social circle, follow the evolution of a preference, or explain why a major decision made sense to the person who made it.
A personal memory benchmark should tell us whether a system formed an accurate, economical, and evolving model of a human life. Retrieval from a conversation covers only part of that job.
What The Current Benchmarks Actually Measure
LoCoMo was an important step. It gave the field ten long conversations, five useful question categories, event summarization, temporal reasoning, and adversarial questions. Its dialogues average about 588 turns and 16,618 tokens across 27 sessions, spanning a few months. That was genuinely long by the standards of conversational datasets in 2024. At the scale of a life, it is still a small and unusually tidy sample: one dialogue channel, two speakers, synthetic event graphs, and an average history that fits inside a 32K context window.
The adversarial category cannot remain optional. Mem0’s public LoCoMo result covers 1,540 questions across the four non-adversarial categories, while Backboard’s published runner explicitly filters out category 5. The table above shows the same blank repeated across other published rows. This removes the part of LoCoMo that asks whether a memory system will invent facts the user never disclosed. A future benchmark should withhold the overall score unless every adversarial question is included. Category selection belongs to the benchmark, never to the test-taker.
LongMemEval pushed the scale much further. It tests information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention through 500 curated questions. Its standard histories contain roughly 115,000 tokens, while the larger setting reaches 500 sessions and about 1.5 million tokens. The benchmark is very good at exposing failures in indexing and retrieval. Its unit of construction, however, is a question embedded in a compiled history. This makes it closer to a sophisticated needle-in-a-haystack test than to the accumulated record of one life.
The May 2026 release of LongMemEval-V2 goes much larger: 451 curated questions, up to 500 trajectories, and as many as 115 million tokens in its largest haystacks. It shows that benchmark infrastructure can already operate at the scale proposed here. Its histories consist of multimodal web-agent trajectories in web and enterprise environments, designed to test whether an agent becomes an experienced colleague. That is a serious agent-memory problem, with a different object of memory from a person’s biography, relationships, emotional history, and private life.
MemBench added several missing dimensions: factual and reflective memory, participation and observation scenarios, plus accuracy, recall, capacity, and read/write efficiency. It also creates individual tests above 100,000 tokens. Its reflective layer is valuable, although it remains largely centered on high-level preferences and emotions derived from structured profiles. Biography, relationship history, motives, contradictions between sources, and developmental change remain much broader targets.
PersonaMem comes closer to personalization. It contains more than 180 simulated user histories, up to 60 sessions, and 15 task scenarios. It explicitly tests whether a model can follow changing traits and preferences and choose an appropriate response at the current point in a user’s timeline. That makes it a strong test of dynamic response alignment. A complete personal memory has a wider burden: it must also preserve the factual and emotional history that makes those preferences intelligible.
BEAM pushes conversational scale to the same general territory. It includes 100 conversations and 2,000 questions, with ten conversations extending to 10 million tokens. Its ten abilities cover extraction, updates, event ordering, contradiction resolution, preferences, summarization, and several other important operations. BEAM still presents memory through generated conversations. Ten million tokens in one channel create an excellent stress test for distance and long-context reasoning; ten million tokens spread across chats, private notes, therapy, work meetings, and competing perspectives create a different problem.
Recent work has started to attack the remaining pieces. HaluMem evaluates extraction, updating, and question answering across histories exceeding one million tokens. PerMemBench introduces multi-year, multi-domain histories for 20 users and asks whether a memory system learns what is worth retaining for each person. HaluMem makes error accumulation visible; PerMemBench makes user-specific retention measurable. The remaining job is a unified test built around a life.
The Dataset Has To Be Uncomfortably Large
A credible human-scale test needs at least ten test individuals, with no less than 100 million source tokens for each person. Public training personas should sit outside that minimum so that teams can develop against representative schemas without seeing the hidden lives used for ranking. At that point the private test set alone begins at one billion tokens, which is exactly why the benchmark becomes useful. Full-context replay stops being a serious answer, and architectural choices about extraction, consolidation, forgetting, and retrieval become visible.
Each individual should have a history spanning months or years and tens of thousands of turns. Some histories can be synthetic, some can come from consenting participants, and some can combine real structural patterns with generated content. Whether the material is real or synthetic matters less than the density and coherence of the life underneath it. Every person needs recurring relationships, unfinished projects, old preferences, new preferences, periods of contradiction, routine days, emotionally significant events, and facts that appear once and never become important until much later.
The sources should be deliberately heterogeneous. Chat extracts reveal informal relationships and momentary reactions. Therapy transcripts reveal self-interpretation, recurring conflicts, and private language that may never appear elsewhere. Personal notes carry plans, doubts, and unpolished thoughts. Meeting minutes and transcripts record professional roles, decisions, and the version of the person visible at work. The same event may appear in several sources with different details and emotional weight. A useful memory system has to preserve provenance and perspective while still forming a coherent account.
Noise should come from life. Random filler alone is too easy. Names collide. Dates are corrected later. A passing preference looks stable for six months and then disappears. Two sources describe the same argument differently. A meeting transcript repeats material already covered in private notes, but adds one consequential fact. These cases expose whether a system consolidates evidence or merely accumulates text.
The Questions Must Reach Beyond Recall
Factual recall remains necessary. The benchmark should ask about biography, education, work, health, events, preferences, daily routines, and the people around the individual. It should include peripheral details with plausible future relevance: the café used for difficult meetings, the colleague who always challenges budgets, the food avoided before travel, or the reason Tuesdays became protected time. Trivia with no connection to the person’s life adds volume without adding much signal.
The next level concerns association and meaning. Questions should require a system to connect an event to the emotional reaction attached to it, or a present concern to an earlier experience in another domain. Why did a promotion produce anxiety rather than satisfaction? Which relationship changed the person’s attitude toward money? What pattern links a failed project, a family conflict, and a later hiring decision? The evidence for these answers will often be distributed and partly interpretive, so short exact-match evaluation will fail.
The hardest questions should address motives, key decisions, principles, and goals. A strong answer may need to distinguish what the person said at the time from what they concluded months later. It may need to present two plausible explanations and calibrate confidence. Human memory contains perspective, revision, and uncertainty; a benchmark that accepts only one decontextualized fact will reward false certainty.
Adversarial questions remain essential. Some should contain false premises, invented relationships, impossible dates, or a detail that belongs to another test individual. Others should exploit near-duplicates and preference changes. The best response will sometimes be a refusal, sometimes a correction, and sometimes an explicit statement that the available evidence supports more than one interpretation.
The benchmark should make plausible guessing expensive. In personal memory, a fluent fabrication about someone’s family or inner life is a severe failure even when it sounds reasonable.
We Need To Test The Memory Artifact Itself
Answer accuracy reveals what a system can use. It says little about what the system actually constructed. Every candidate should therefore produce an inspectable memory artifact after ingestion, even when its internal representation is proprietary. A standardized export can expose claims, evidence references, timestamps, confidence, relationships, and preference states without forcing every architecture into the same storage schema.
That artifact enables several additional tests. A reference inventory can measure how many supported facts were captured and how many unsupported facts were introduced. A reference social graph can measure whether the system found the important people, their roles, their relationships to one another, and the changes in those relationships. Psychological profiles can be compared against established instruments such as the Big Five Inventory-2, together with source evidence and confidence. Preference timelines can test whether the system preserved evolution instead of flattening every statement into one timeless profile.
These outputs should still be judged semantically. Social relationships and psychological descriptions rarely have one canonical wording. An LLM judge can compare the candidate artifact with the hidden reference model, evidence, and a domain-specific rubric. Human experts belong in dataset creation, reference construction, and judge calibration; they should not become a variable that changes between leaderboard submissions.
One Run With Sources, One Run Without Them
The benchmark needs two answer conditions. In the source-assisted condition, the candidate can search the original corpus at query time. This measures the complete system: memory, retrieval, and access to raw evidence. In the sealed-memory condition, source access ends after construction and updates. The system must answer using only the memory it chose to build.
The gap between those scores may be more informative than either score alone. Strong source-assisted performance with weak sealed-memory performance indicates an effective search layer and a poor constructed memory. Strong performance in both conditions suggests that the system compressed the life without discarding the structure needed later. A memory benchmark that always leaves the source attached cannot clearly separate memory from search.
The sealed condition also forces decisions about representation. A compact list of facts may handle biography and fail on emotional causality. A transcript summary may preserve narrative while losing peripheral detail. A graph may reconstruct the social circle and flatten first-person perspective. Those tradeoffs are where memory architectures become interesting.
Fairness Requires A Frozen Harness
The dataset should be divided into a public training set and a hidden test set. Test individuals, questions, reference answers, and judging evidence should remain private behind an evaluation service. The public release should contain enough complete examples to make the task reproducible without making leaderboard overfitting easy.
Every candidate must use the same approved models for memory processing and answer generation. The judge model, judge prompt, decoding settings, and per-question rubric must also remain identical across submissions. Candidate systems can differ in schemas, consolidation policies, storage, indexing, retrieval, update logic, and forgetting. Giving one system a stronger extraction model or a more favorable judge would turn an architecture comparison into a model procurement contest.
Leaderboard correctness should come exclusively from LLM-as-a-judge evaluation. Exact match and lexical overlap are too brittle for causal, emotional, and interpretive answers. The judge should receive the question, question type, reference answer, supporting and contradicting evidence, and a fixed rubric covering factual support, completeness, uncertainty, and false claims. The same prompt must be used for every candidate. LLM judges have documented position, verbosity, and self-enhancement biases, so the frozen judge and prompt are part of the experiment. A published calibration set can show how the judge handles borderline answers, while hidden examples protect the test.
Accuracy Without Economics Is Incomplete
The primary result should be percentage correct by question type, reported separately for source-assisted and sealed-memory conditions. An overall score can be included, but it should never erase the profile. A system that excels at single facts and collapses on motives has a different capability from one that understands life patterns but misses dates.
Every run also needs telemetry. Report median tokens and median wall-clock time per answer, again by question type and source condition. Report the tokens and time spent constructing the initial memory, the cost of subsequent updates, and the final memory size. These numbers reveal whether an accuracy gain came from a better representation or from spending far more compute at every turn.
Incremental behavior deserves its own test. After the initial memory is built, add a noisy source whose content overlaps the existing evidence by 80%. Measure the percentage increase in memory size, the capture rate for the genuinely new 20%, conflict resolution, update cost, and any regression in previous answers. A system that doubles its memory after receiving mostly duplicate information has failed at consolidation, even if its next answer remains correct.
The final scorecard should therefore contain answer accuracy, source dependence, fact coverage, unsupported-memory rate, psychological profile accuracy, social graph completeness, preference evolution, construction cost, update cost, query cost, latency, memory size, and overlap growth. No single number can honestly carry all of that information.
What A Better Benchmark Would Change
Benchmarks shape architecture. When the dominant task is question answering over a modest transcript, teams optimize retrieval over transcripts. When the task includes 100 million source tokens per person, several noisy sources, hidden life structure, source-free answering, psychological and social reconstruction, and measured update efficiency, the winning systems will have to solve a much larger problem.
That shift would also improve how claims are communicated. A 94% LoCoMo result should be read as a 94% LoCoMo result: meaningful evidence on a defined task. Human-scale understanding demands evidence at human scale.
Based on my own judgment and hands-on work with current memory systems on human-memory tasks, I expect the best of them to score no more than 10% on such a benchmark. That would be a healthy result. LoCoMo’s leading results are now crowded above 90%, and much of the competition has become a search for the last percentage point under incompatible harnesses. A human-scale benchmark would direct attention back to the open problems: deciding what deserves to survive, joining evidence across domains, representing motive and perspective, resolving contradictions, and updating a person without flattening their history. Personal memory needs that kind of discontinuity. Another fraction of a point on a saturated test will not create it.
The benchmark will never be perfect. Its job is to make the important failures visible. It should be large enough to defeat replay, varied enough to defeat one-channel retrieval, intimate enough to test meaning, adversarial enough to punish invention, and controlled enough that two scores actually deserve comparison.
Once the benchmark measures the construction of a life and the use of that memory under pressure, personal memory systems will finally have to show how well they know a person.
