Evals
Frozen scores for turbomem's extract → embed → scoped search pipeline. No competitor column. Methodology is public so later write-ups can cite the snapshot, not a vibe.
LoCoMo F1
27.1%
1986 questions
Recall@5
100.0%
5 queries
Scope leak rate
0.0%
Target 0%
First-party goldens for extraction, retrieval, write-time dedup, and multi-tenant scoping.
Ingest locomo10.json through TurboMemory.add, answer from search, score with official category-aware token F1. Dataset category IDs: 1 multi-hop, 2 temporal, 3 open-domain, 4 single-hop, 5 adversarial.
| Category | n | F1 |
|---|---|---|
| 1 · multi-hop | 282 | 12.6% |
| 2 · temporal | 321 | 1.8% |
| 3 · open-domain | 96 | 11.5% |
| 4 · single-hop | 841 | 14.4% |
| 5 · adversarial | 446 | 81.6% |
| Overall | 1986 | 27.1% |
251 tokens/query · search p50 469.43 ms
Harness lives in the turbomem repo. Full freeze needs OPENAI_API_KEY.
pnpm --filter @turbomem/evals test pnpm --filter @turbomem/evals download-locomo pnpm --filter @turbomem/evals evals -- --suite all --publish --sync-landing
Public snapshot. Headline LoCoMo number is official category-aware token F1, not LLM-as-judge.
Read the methodology, including what these numbers do not claim.