Benchmark deep-dives
Accuracy and mean context tokens per question on each benchmark. Single pass over the full question set, graded by the benchmark's own judge.
1,540 questions, 4 categories. Single-hop, multi-hop, open-domain and temporal recall across multi-session conversations spanning months.
Mean context tokens are 4× below the full history.
Gold session reached the reader for 96–99% of questions. Swapping the reader model moves the score by less than 0.1 points.
Why the numbers
look this way
Each score traces back to a specific part of the architecture, not just asserted.
Recall correctness
Every session is narrated into timestamped episodes that keep who, what and when together, so a fact is retrieved with its context. The retrieval decider reads the first hits and asks follow-up questions before settling, which is why the gold session is in context for 99.95% of LongMemEval questions.
Context footprint
Keyword and vector search are fused by reciprocal rank and only the top twenty episodes enter the prompt. That keeps a LongMemEval call at about 9,000 tokens against a 115k-token history, and a LoCoMo call at about 6,500 against 26k.
Response time
Search runs in 1.1–2.5 s at the median, including one to three decider rounds. End-to-end median is 3–15 s depending on the reader model; the memory layer is not the bottleneck.
Latency and tokens
Search and end-to-end timings, plus how much context each call actually sends to the reader.
| LoCoMo | LongMemEval | DMR | |
|---|---|---|---|
| search latency, p50 / p95 | 1.67 s / 7.02 s | 2.53 s / 6.11 s | 1.13 s / 1.71 s |
| end-to-end, p50 / p95 | 8.30 s / 18.0 s | 14.7 s / 30.2 s | 3.21 s / 6.34 s |
| context tokens per question | 6,562 | 8,970 | 1,602 |
| full history per question | ~26k | ~115k | — |
| token reduction | 4× | 13× | — |
| reader model | gpt-4.1-mini | gpt-5.5 | gpt-4o-mini |
What's inside Membase
Four pieces working together, from how a session is cut up to where the memories live.
Boundary detection
An LLM pass splits each session into topical cells before extraction, so one episode never straddles two subjects.
Episodic extraction
Each cell becomes a titled, timestamped narrative from the user's point of view. Optional profile, fact and foresight layers sit beside it.
Multi-round retrieval
Hybrid search per sub-query, fused by reciprocal rank; a decider marks core evidence and issues new queries for up to three rounds.
Local-first store
SQLite and FAISS on disk, scoped per user, with any OpenAI-compatible model for extraction, retrieval and answering.
Research Blog
From the Unibase Research TeamThe posts behind the numbers and architecture above.