Evals

How we measure memory

Frozen scores for turbomem's extract → embed → scoped search pipeline. No competitor column. Methodology is public so later write-ups can cite the snapshot, not a vibe.

LoCoMo F1

27.1%

1986 questions

Recall@5

100.0%

5 queries

Scope leak rate

0.0%

Target 0%

storage: pgliteembeddings: text-embedding-3-smallextraction: gpt-4.1-minik=102026-09-28 · c5088aa

Product suite

First-party goldens for extraction, retrieval, write-time dedup, and multi-tenant scoping.

Extraction F1
100.0% (n=4)
Retrieval recall@5 / MRR
100.0% / 100.0%
Dedup accuracy
66.7% (n=3)
Scope isolation
100.0% accuracy, leak 0.0%
addFacts p50 / p95
464.76 ms / 517.53 ms
search p50 / p95
421.53 ms / 537.87 ms

LoCoMo

Ingest locomo10.json through TurboMemory.add, answer from search, score with official category-aware token F1. Dataset category IDs: 1 multi-hop, 2 temporal, 3 open-domain, 4 single-hop, 5 adversarial.

CategorynF1
1 · multi-hop28212.6%
2 · temporal3211.8%
3 · open-domain9611.5%
4 · single-hop84114.4%
5 · adversarial44681.6%
Overall198627.1%

251 tokens/query · search p50 469.43 ms

Reproduce

Harness lives in the turbomem repo. Full freeze needs OPENAI_API_KEY.

pnpm --filter @turbomem/evals test
pnpm --filter @turbomem/evals download-locomo
pnpm --filter @turbomem/evals evals -- --suite all --publish --sync-landing

Public snapshot. Headline LoCoMo number is official category-aware token F1, not LLM-as-judge.

Read the methodology, including what these numbers do not claim.