Unibase
    • Memory
    • BitAgent
    • UB Bridge
    • Governance
    • Membase
    • AIP
    • Unibase Pay
    • Unibase DA
    • Explorer
    • Docs
    • Blog
    • GitHub
    • Twitter
    • Telegram
    • Discord
Unibase

© 2026 Unibase

Products
MemoryBitAgentUB BridgeGovernance
Infrastructure
MembaseAIPUnibase PayUnibase DA
Developers
ExplorerDocsBlog
Community
GitHubTwitterTelegramDiscord
Membase · Research

Benchmarking long-term memory for agents.

Measured on LoCoMo, LongMemEval and DMR. Powered by episodic extraction and multi-round retrieval that sends the reader a few thousand tokens instead of the whole history.

93.1LoCoMo
92.6LongMemEval
92.2DMR
Mean tokens per retrieval call
Membase~6,500
Full-context26,000+

4× fewer tokens on LoCoMo. On LongMemEval the gap is 13×: ~9,000 tokens against a 115,000-token history.

Benchmark

Benchmark deep-dives

Accuracy and mean context tokens per question on each benchmark. Single pass over the full question set, graded by the benchmark's own judge.

LoCoMo

1,540 questions, 4 categories. Single-hop, multi-hop, open-domain and temporal recall across multi-session conversations spanning months.

93.1%
Accuracy
6,562
Mean tokens

Mean context tokens are 4× below the full history.

Gold session reached the reader for 96–99% of questions. Swapping the reader model moves the score by less than 0.1 points.

Accuracy by categoryoverallcategory
Methodology

Why the numbers
look this way

Each score traces back to a specific part of the architecture, not just asserted.

01

Recall correctness

Every session is narrated into timestamped episodes that keep who, what and when together, so a fact is retrieved with its context. The retrieval decider reads the first hits and asks follow-up questions before settling, which is why the gold session is in context for 99.95% of LongMemEval questions.

02

Context footprint

Keyword and vector search are fused by reciprocal rank and only the top twenty episodes enter the prompt. That keeps a LongMemEval call at about 9,000 tokens against a 115k-token history, and a LoCoMo call at about 6,500 against 26k.

03

Response time

Search runs in 1.1–2.5 s at the median, including one to three decider rounds. End-to-end median is 3–15 s depending on the reader model; the memory layer is not the bottleneck.

Performance

Latency and tokens

Search and end-to-end timings, plus how much context each call actually sends to the reader.

LoCoMoLongMemEvalDMR
search latency, p50 / p951.67 s / 7.02 s2.53 s / 6.11 s1.13 s / 1.71 s
end-to-end, p50 / p958.30 s / 18.0 s14.7 s / 30.2 s3.21 s / 6.34 s
context tokens per question6,5628,9701,602
full history per question~26k~115k—
token reduction4×13×—
reader modelgpt-4.1-minigpt-5.5gpt-4o-mini
Architecture

What's inside Membase

Four pieces working together, from how a session is cut up to where the memories live.

Boundary detection

An LLM pass splits each session into topical cells before extraction, so one episode never straddles two subjects.

Episodic extraction

Each cell becomes a titled, timestamped narrative from the user's point of view. Optional profile, fact and foresight layers sit beside it.

Multi-round retrieval

Hybrid search per sub-query, fused by reciprocal rank; a decider marks core evidence and issues new queries for up to three rounds.

Local-first store

SQLite and FAISS on disk, scoped per user, with any OpenAI-compatible model for extraction, retrieval and answering.

Blog

Research Blog

From the Unibase Research TeamThe posts behind the numbers and architecture above.

Research posts from the Unibase Research Team are coming soon.
View all research posts→
Membase · benchmark reportSeptember 2026