Abstract
We benchmarked Engram against mem0 and supermemory across store throughput, search latency, and semantic recall accuracy. Across 5 independent runs of 50 queries each, Engram's results showed zero variance — R@1=0.86, MRR=0.899, 100% recall at R@10. Engram stores 100 memories in 7.5 seconds versus mem0's 49.2 seconds and supermemory's 134.1 seconds. Search latency floors at 38ms versus 287ms (mem0) and 426ms (supermemory). No LLM is invoked per query.
Methodology
Dataset: 100 memories stored per system. All three systems received identical data through their official APIs with recommended configurations.
Queries: 50 ground-truth queries, written with different vocabulary than the stored memories. Queries do not share keywords with their target memory — this tests semantic understanding, not keyword matching.
Runs: 5 independent runs per system. All reported figures are mean across runs. Standard deviation across all 5 runs: 0.0 on retrieval metrics.
Hit detection: Keyword-based ground truth. Each query maps to exactly one memory containing unique identifiable content. A result is a hit if the correct memory appears in the top-k results.
Negative query validation: 10 out-of-domain queries were included. All 10 returned max cosine scores below 0.35 — no false positives.
Results
Store Throughput (100 memories)
| System | Ops/sec | Avg ms | p95 ms | Total time |
|---|---|---|---|---|
| Engram (local) | 13.3 | 74.9 | 130.6 | 7.5s |
| mem0 (cloud) | 2.0 | 491.9 | 610.4 | 49.2s |
| supermemory (cloud) | 0.7 | 1341.2 | 1770.0 | 134.1s |
Engram stores synchronously to the local vector database with no network round-trip. mem0 is 6.6x slower due to cloud latency and LLM extraction on every write. Supermemory is 17.9x slower.
Search Latency
| System | Min ms | Avg ms | p95 ms |
|---|---|---|---|
| Engram (local) | 38.0 | 48.0 | 51.0 |
| mem0 (cloud) | 287.5 | 339.0 | 391.6 |
| supermemory (cloud) | 426.0 | 560.2 | 852.1 |
Engram's latency is bounded by local vector search. Both cloud competitors add network overhead and — in mem0's case — an LLM reranking step on every query. Engram's p95 (51ms) is lower than mem0's minimum observed latency (287ms).
Semantic Accuracy (different vocabulary queries)
| System | R@1 | R@3 | R@5 | R@10 | MRR |
|---|---|---|---|---|---|
| Engram (local) | 0.86 | 0.90 | 0.96 | 1.00 | 0.899 |
| supermemory (cloud) | 0.88 | 0.92 | 0.96 | 0.96 | 0.91 |
| mem0 (cloud) | 0.40 | 0.48 | 0.56 | 0.64 | 0.46 |
Engram and supermemory are statistically close at R@1 (0.86 vs 0.88). Engram reaches 100% recall by R@10 — supermemory does not. mem0's retrieval accuracy is significantly lower; this is likely an effect of their LLM extraction layer dropping or transforming memories before indexing.
Real agent systems inject 5–10 retrieved memories as context, not one. R@5 and R@10 matter more than R@1 for production use.
Full Evaluation Suite
These results are from the full evaluation suite (5 runs, 50 queries each). All figures are validated with stddev=0.0.
Retrieval
| Metric | Value | 95% CI |
|---|---|---|
| R@1 | 0.86 | [0.86, 0.86] |
| R@3 | 0.90 | — |
| R@5 | 0.96 | — |
| R@10 | 1.00 | — |
| MRR | 0.899 | — |
Classification
Memory type classification uses keyword-based detection with priority ordering — no LLM.
| Metric | Value |
|---|---|
| Overall accuracy | 81.7% (49/60) |
| Preference F1 | 1.000 |
| Decision F1 | 1.000 |
| Fact F1 | 0.815 |
| Entity F1 | 0.696 |
| Macro F1 | 0.702 |
Preferences and decisions classify perfectly. Facts and entities show lower F1 due to surface-form variation — an area of active development.
System Performance
| Operation | Mean | p95 | p99 |
|---|---|---|---|
| Store | 58ms | 71ms | 84ms |
| Search | 48ms | 51ms | — |
| Forget | 8.6ms | — | — |
| Feedback | 1.1ms | — | — |
Throughput: 17 stores/sec, 18.2 searches/sec
Concurrency: 0 failures at 50 concurrent workers, 20.6 ops/sec under load
Latency scaling: FLAT from 50 to 1000 memories — no degradation observed
How It Works
Engram uses a three-tier recall pipeline:
- Hot cache — recently and frequently accessed memories. Sub-millisecond lookup for repeated queries. Hit rate climbs toward 100% with sustained use.
- Hash index — exact and near-exact match lookup. Catches identical or near-identical queries without a vector search.
- Hybrid vector search — dense embeddings with cosine re-rank as the final tier. Handles semantic queries with different vocabulary.
The cosine re-rank step (applied post-fusion) was the single largest accuracy improvement in our evaluation — R@1 went from 0.72 to 0.86, a +0.14 gain. It runs entirely in-process with no external calls.
What We Don't Do
No LLM extraction on store. mem0 and supermemory run an LLM on every write to reformat, extract, and index memories. This adds latency, burns tokens, and can silently drop information. Engram stores exactly what you give it.
No per-query LLM reranking. Cloud-based reranking adds 200–800ms per query and introduces a dependency on external API availability. Engram's cosine re-rank is local and synchronous.
No mandatory cloud. Data stays on your hardware. No network required for store or search.
Reproduce
The benchmark suite is open source.
git clone https://github.com/EngramMemory/engram-memory-community
cd engram-memory-community
pip install -r benchmarks/requirements.txt
python3 benchmarks/competitive_benchmark_2026.py
All competitor tests use official, publicly available APIs with recommended configurations. The 50 ground-truth queries and memories are deterministic and reproducible. Run 5 times to validate variance.