Introduction
What is Engram Memory and why it exists
Introduction
Engram Memory is a hosted intelligence layer for AI memory. Your database, your hardware, Engram's intelligence.
Traditional memory solutions force a choice: build it yourself or hand your data to someone else. Engram eliminates that tradeoff. It processes memories through a sophisticated intelligence pipeline — embedding, classifying, deduplicating, and compressing — then returns the results to your own infrastructure. Engram never stores your data unless you explicitly opt into overflow storage.
Stateless by Design
Every API call is stateless. Engram processes your request, returns enriched results, and forgets. Your vectors live in your vector database, on your hardware, under your control. This is a deliberate architectural choice, not a limitation.
One Call Does Everything
The /v1/intelligence endpoint replaces an entire pipeline of tools:
POST /v1/intelligence
In a single API call, Engram will:
- Embed your text into 768-dimensional vectors
- Classify the memory type automatically
- Deduplicate against your existing memories
- Compress vectors with proprietary encoding
No chaining endpoints. No orchestration logic. One call, full pipeline.
Intelligent Recall
Engram uses a multi-tier recall architecture to deliver the right memory at the right speed:
| Tier | Latency | Purpose |
|---|---|---|
| Hot Cache | Sub-millisecond | Recently accessed memories |
| Deduplication Index | Microseconds | Exact-match detection |
| Semantic Search | Low milliseconds | Full similarity search |
Queries hit the fastest tier first and fall through only when needed.
Auto-Classification
Every memory is automatically classified into one of five types:
- preference — User likes, dislikes, and behavioral patterns
- fact — Objective information and knowledge
- decision — Choices made and their reasoning
- entity — People, organizations, places, and things
- other — Everything else
Classification happens server-side with zero configuration. Use it to filter recalls, build category-specific UIs, or power smart retrieval strategies.
Proprietary Compression
Engram's compression delivers 6-8x vector size reduction while preserving >0.99 recall quality. This means your vector database stores significantly more memories in the same hardware footprint without meaningful quality loss.
Use Cases
- AI assistants — Give any LLM persistent memory across conversations
- Autonomous agents — Agents that learn, remember decisions, and avoid repeating mistakes
- Knowledge management — Ingest and recall organizational knowledge at scale
- Multi-device fleets — Shared memory across edge devices, robots, or IoT deployments
Engram vs Mem0
| Engram | Mem0 | |
|---|---|---|
| Architecture | Stateless intelligence layer | Hosted storage |
| Data residency | Your infrastructure | Their servers |
| Vector storage | Your vector database | Their managed DB |
| Intelligence | Embed + classify + dedup + compress | Basic embedding |
| Self-hosted option | Community Edition (Docker) | Limited |
Engram is not a database. It is an intelligence layer that makes your database smarter.
SDKs
Python
pip install engrammemory-ai
JavaScript
npm install engrammemory-ai
Both SDKs provide full coverage of the API with typed interfaces and built-in error handling.
Community Edition
For fully local, zero-cost deployments, the Community Edition bundles everything into a single Docker container. No API key required. No external calls.
docker pull engrammemory/engram-stack
Next Steps
- Quickstart — Store and recall your first memory in under 5 minutes
- API Reference — Full endpoint documentation
- Self-Hosted Guide — Deploy the Community Edition