RCLL
Shared memory for a fleet of agents. One store they all write into, where every fact carries a room, a hall and a layer — so a recall can ask for exactly the slice it needs instead of the whole pile. Your Postgres, local embeddings, MIT — and no account, no API key and no outbound call are required to read from it.
What is wrong with it
All of this is argued in full further down or in the docs, and most of it sits past the halfway mark of a long page — so a reader who stops early, or a fetcher that truncates, comes away with a better impression than we can support. It goes first instead. Every line links to where the number was taken.
- Our default retrieval is the wrong one for the questions we built this for. On multi-hop, fusion plus reranking scores 0.3651 where plain dense plus reranking scores 0.4057. Multi-hop is literally the shape of the product — one agent found it, another needs it. The cause is measured and a fix is identified; it is a directional result and it is not shipped. why it loses →
- 201,154 links, of which 2 are causal. The rest is machinery: 78,009 time-window adjacencies, 77,825 cosine similarities, 45,318 shared entity names. Nothing but those two edges was asserted by an agent, so do not read the count as structure a fleet built. the breakdown →
- A recall aimed at the wrong room returns exactly zero. Not a degraded answer — an empty one, in every configuration we ran. That is the price of a filter that is a hard gate rather than a hint, and we measured it rather than reasoned about it. the ablation →
- A room is assigned once and there is no way to move it. No endpoint, no job, no query — so a room chosen badly on Tuesday is still wrong in month three. Ours are already dirty: 595 facts about deploys sit in ten different rooms and 260 in no room at all. what rooms do and do not do →
- Authorship covers 243 of 8,810 facts. It began working this week. Everything older was written before there was anything to record it with, and it does not backfill — the identity is not recoverable from the store. who wrote what →
- Memory does not export yet. The endpoint named
exportreturns a bank's configuration and not one fact, and the listing API drops room, hall, layer and all 201,154 links. A command named export that hands back something which is not your memory is a backup you find out about later. the export gap →
Two more you would find anyway: on single-hop questions a twenty-five-year-old lexical baseline beats our dense retrieval (0.4591 against 0.4518), and no API key is exact for reading and misleading if you take it to mean no key anywhere — writing facts needs a model, local or hosted. Both are worked out below.
It runs with nothing outside the box
Self-hosted usually means your data sits on your disk while the thinking still happens at a vendor. Here the storage and the retrieval are both yours: Postgres you already run, a 33M-parameter embedder and an 80 MB reranker that ship with the image, and a recall path with no completion call in it at all. There is no RCLL account, no hosted control plane, no phone-home, and nothing to rate-limit you.
The only place a model is genuinely needed is the write path — turning a document into typed, dated facts. That is one HTTP call per document, and you choose where it goes. Three answers, all supported today:
| Write mode | Leaves the box | What you get |
|---|---|---|
LLM_PROVIDER=none |
nothing, ever | A working vector-and-lexical chunk store. No fact extraction, no entities, no causal links, reflection and consolidation off. Smaller product than the rest of this page describes — say so out loud before you pick it. |
| A local model llama-server + a ~2B GGUF |
nothing, ever | Full extraction with no key anywhere. The engine already carries an offline provider and will fetch the weights itself. We measured it against a hosted model on the same chunks, and the table is below. It does not cost you facts. It costs you time. |
| Any OpenAI-compatible endpoint your key, your base URL |
one call per document written | Full extraction, fastest, and the configuration our own numbers were taken on. Reads still make no model call, so a key here does not put your queries anywhere. |
Swapping a hosted key for a local model is three environment variables, not a fork. Point the OpenAI-compatible provider at a local server:
LLM_PROVIDER=openai
LLM_BASE_URL=http://127.0.0.1:8091/v1
LLM_MODEL=gemma-4-e2b-it
Two things we learned doing it, both worth having before you start. Use the
OpenAI-compatible path, not the ollama or lmstudio provider names —
for those the engine skips JSON-schema enforcement and pastes the schema into the prompt as a
suggestion instead, which is exactly the guard rail a 2B model needs to stay on a nested extraction
schema. And if your local server caps completion length, set the retain token budget yourself: the
engine only clamps that ceiling for model names it recognises as OpenAI's, so a third-party endpoint
can be handed a budget it will refuse.
What taking the key out actually costs
gemma-4-E2B through llama-server against gpt-4o-mini. Same 20 chunks,
same harness, one box, 8 vCPU, no GPU.
| Extraction, first attempt | local 2B | gpt-4o-mini |
|---|---|---|
| valid JSON | 100% | 100% |
| satisfied the extraction schema | 50% | 15% |
| facts per chunk | 4.75 | 5.05 |
fields left N/A | 27.2% | 48.8% |
| median seconds per chunk | 96.6 | 7.4 |
The hosted model lost the schema row, which is not the result anyone expected, so we
went and read the failures instead of printing the number. All 17 of them are the same shape: the
model returns entities: ["Caroline", "the support group"] and the schema wants
[{"text": "Caroline"}]. It does that because the extraction prompt in this engine —
inherited, not written by us — shows exactly that in its worked examples, a few lines above a schema
that forbids it. The 2B model followed the schema and failed on a different field. So that row
records which of two contradictory instructions a model happened to obey, not how good it is. It is
a defect, it is ours now that we ship the fork, and it hits any OpenAI-compatible endpoint on this
path.
What survives the correction is the row nobody can argue with: 13× slower. That is the real price of taking the key out — latency on CPU, not lost facts. n=20, one box, one dataset, and we will say so again when the number moves.
Most agent memory is built for one assistant: one bot, one history, one user. That shape breaks the moment you run more than one agent. Ours found out the hard way — an architect spent Monday re-deriving what a sysadmin had already worked out on Friday, because there was nowhere for Friday to leave it.
RCLL is recall with the vowels dropped — the one
operation every agent performs before it does anything else. The MCP tool is literally called
memory_recall; the product is named after the call.
The engine ships under a different name on purpose. RCLL is the product;
fleet-memory is the artifact — the repository, the container image and the npm
package all carry that name, because it says what the thing does and because a name you pin in a
lockfile should never have to change. Same code, one download, two names: the one you read and the
one you depend on.
The read path never calls a model
This is the design decision everything else follows from. A memory_recall is four
retrievals in parallel — vector, BM25, graph activation, temporal — merged by reciprocal rank
fusion, reranked by a local cross-encoder, then cut to a token budget. There is no completion
call anywhere in it. In the source tree, the whole search package contains zero LLM invocations
reachable from the recall tool, and zero writes.
Three things fall out of that, and they are the reason to run this rather than something managed:
- Reading costs no tokens. Not "fewer tokens" — none. The only model in the read path is a 33M-parameter embedder and an 80 MB reranker, both local by default.
- A read-only deployment needs no API key at all. Set
LLM_PROVIDER=noneand there is nothing on the box to leak. - Nothing leaves the machine on a read. That is a property of the code path, not a policy we promise.
Writing is the other half of that ledger and we are not going to be quiet about it.
memory_retain does call a model: it splits incoming text into atomic facts, types and
dates them, resolves entities and decides what supersedes what. That is a real per-document cost,
and on that column RCLL is not cheap.
Which makes one correction worth printing here rather than in a footnote. "No API key" is exact for a read-only node and misleading if you read it as "no key anywhere". Closing the write path too is the choice laid out at the top of this page: no key at all costs you fact extraction, and a local model costs you speed — about 13× on a CPU box — and, as far as we can measure, nothing else. Both are real options; neither is free. How the read path works →
One thing that genuinely costs nothing either way: rooms. The classifier that decides which room a fact belongs in is 176 lines of regular expressions. The only part of this system we invented is also the only part that never touches a model.
Friday, then Monday
Friday 17:42
sysadmin →
Finds that a push to the new org fails because the working clone is shallow: the root commit
names an upstream parent that no longer exists. Writes it down, room agent,
layer L0, as a procedure.
Monday 11:05
architect →
Different agent, different session, no handoff, nobody linked them. Asks the store what it
should know before publishing a repository. Gets Friday's finding back, checks
.git/shallow first, and does not spend the morning on it.
That is the whole product. Not "memory" as a feature — a place where one agent's Friday survives into another agent's Monday.
Rooms
Scoping is one field. Every fact is written into a room, and a recall names the rooms it wants — a topic filter applied before the search, not a permission system with tiers and policies. A floor plan an agent can hold in its head.
Rooms sit inside two more axes: hall is what kind of thing a fact is (decision, procedure, warning, preference…), layer is how durable it is — L0 rules that surface at session start, L1 decisions retrievable by topic, L2 raw observations that get compacted later. The room / hall / layer model →
One limitation belongs right here rather than three clicks away: a fact's room
is decided once, when it is written, and nothing ever moves it. There is no reassignment
path in the engine — no endpoint, no background job, no query that updates the column — so a room
chosen badly on Tuesday is still wrong in month three. Since 26 August a caller who types
deploy lands in deployment and keeps the string they typed as a tag, which
stops the drift going forward without touching a single fact already written. That runs in our own
client layer, not in the packaged server, and the backfill for facts written before it is not done.
The registry, the janitor, and what is still open →
Getting the room wrong is not a slightly worse search. A recall filtered to a room that does not hold the answer scores exactly zero — measured, not asserted. That is the cost side of a filter this sharp, and it is why naming rooms consistently turned out to matter more than we thought. What scoping is worth, in numbers →
And now: who wrote it
A room says what a fact is about. It never said who put it there, and until this week nothing did — the writer's name sat in scope at the call site and went into a log line. One field, never sent. That is fixed, and it is deliberately a separate axis: what a fact is about and who wrote it are different questions, and folding them into one string answers neither.
memory_recall({ query: "how do we deploy", mine: true }) // only what I wrote
memory_recall({ query: "the gzip change", author: "sysadmin" }) // someone else's desk
memory_recall({ query: "the gzip change" }) // everyone — author shown on each hit
Three properties worth stating in the open, because each one is a decision that could have gone the other way:
- You cannot claim to be someone else. The server resolves the caller and mints the attribution itself; an author written by hand in the call is discarded. Self-declared attribution stays possible where no server-side identity exists — and it is recorded as self-declared, so a reader can tell the two apart without guessing.
- Nothing is attributed to a stand-in. Where no identity resolves, the fact is stored with no author and a warning. A wrong name is worse than a missing one: a missing author is a gap you can see.
- It survives compaction. Attribution rides a tag rather than the metadata column, because consolidation rewrites facts into observations and metadata does not make that trip. Had we used the obvious field, authorship would have quietly vanished from exactly the records the engine writes by itself.
Two bounds we would rather state than have found. It is not retroactive — 243 of 8,810 facts in our store carry an author, because everything older was written by code with no way to say — and it lives in our client layer today, so the packaged MCP server gets it with the first release. Both sit on the roadmap, next to the two remaining fields that would make a handoff measurable rather than merely recorded.
What one fleet actually accumulated
Numbers from the store our own agents write to, read on 26 August 2026. This is usage, not a benchmark — nobody tuned anything to make these look a particular way, and they grow every day.
One store, 17 rooms in use, three layers. 1,350 facts in shared.
Every durable fact carries a room — L0 459 of 459, L1 3,176 of 3,176, both exactly 100% — and the
unroomed remainder is entirely L2 raw observation waiting on consolidation. Rooms get assigned
where durability is decided. The room / hall / layer model →
Next on that axis — decided, not built: a room registry and a janitor
that runs at write time. Creating a room already costs one string; what a fleet cannot do is see which
rooms exist, or keep deploy and deployment from becoming two of them. 595
facts in this store mention a deploy and sit in ten different rooms, 260 in no room at all. The fix
files the write under the canonical name and keeps the string the caller typed as a tag, so every
rewrite stays auditable and reversible using machinery the store already has.
What it will and will not do →
Do not read the link count as fleet-built structure. A number that large sitting next to "22 agents writing" invites the wrong inference, so here is the breakdown: 78,009 temporal edges (time-window adjacency), 77,825 semantic (cosine similarity above a threshold), 45,318 entity (two facts naming the same thing) — and 2 causal. Everything except those two edges was derived by machinery, not asserted by an agent. Anyone who gets a dump of this store can compute that split in one query, so it may as well be on the front page.
How good is what it brings back
Until this week we had no answer to that and said so. Now there is one, and it was produced on purpose in the least flattering way available: a competitor's harness, a public dataset, and a lexical baseline that beats most agent memory systems in published work.
LoCoMo, all 1,531 questions with evidence labels. Ranking only — no reader, no judge, no model in the loop at all, because the metric is arithmetic over evidence IDs. BM25 is the baseline every other row is measured against.
| Configuration | nDCG@10 | vs BM25 | recall@10 |
|---|---|---|---|
| BM25 alone | 0.3885 | baseline | 0.522 |
| vector alone | 0.4244 | +0.036 | 0.581 |
| hybrid fusion (BM25 + vector) | 0.4722 | +0.084 | 0.615 |
| vector + rerank | 0.5607 | +0.172 | 0.648 |
| fusion + rerank (default) | 0.5862 | +0.198 | 0.668 |
95% CI on the default configuration's margin: [+0.182, +0.213], paired bootstrap, sign test p < 1e-100. Vector alone clears BM25 by a margin whose interval nearly touches zero — the lexical baseline is genuinely strong, and a store that only did embeddings would have very little to show here. The fused rows combine two channels, lexical and dense; the shipped engine fuses four. The direction carries, the magnitude belongs to this substrate.
Two results in that table argue against us, and they are the reason to trust the other three.
- On multi-hop questions the default configuration is not the best one we ran. Vector+rerank scores 0.406 there; fusion plus rerank scores 0.365. Fusion helps everywhere else and hurts here. Multi-hop is precisely the shape of "agent A found it, agent B needs it", which makes this the category we care most about and the one where our defaults are currently wrong. A follow-up run found the mechanism and a fix; both are on the benchmarks page, and the fix is not shipped yet.
- On single-hop questions BM25 beats dense retrieval (0.459 against 0.452, not significant). Twenty-five years of lexical search is not a straw man.
What this number is not. It is retrieval quality — did the right evidence come back and how high. It is not answer accuracy, and it does not belong anywhere near a vendor's "77% on LoCoMo": those figures come from a reader model answering and a judge model grading, which is a different measurement in different units. Changing the reader-and-judge pair moves a score more than changing the memory store does. We will publish accuracy when we publish the reader, judge, seed and dataset version alongside it, and not before. Full tables, conditions and caveats →
What a recall costs, phase by phase
Every memory vendor publishes one latency number. One number is not checkable and does not tell you what to fix, so here is the whole ledger — median of 12 real queries against the store above, traced per phase, on the box that also runs our production CRM. 8 vCPU, no GPU, both models forced to CPU.
| Phase | p50 | What it does |
|---|---|---|
| embed query | 0.026 s | local bge-small-en-v1.5 |
| retrieve ×4 | 0.253 s | vector, BM25, graph, temporal — in parallel |
| RRF merge | 0.004 s | ~1,541 candidates fused |
| rerank | 2.570 s | cross-encoder over 300 candidates, on CPU |
| token filter | 0.006 s | cut to 4,096 tokens, ~108 facts returned |
| total | 3.016 s |
The reranker is 85% of it. Search itself — everything except the cross-encoder — is 0.289 s p50. We are publishing the slow number rather than the flattering one because the split is the useful part: it says the fix is a GPU or a smaller reranker, not an index rewrite. Read the caveats before comparing this to anyone else's figure, including ours: the numbers page →.
The thing you cannot take with you yet
Rooms are the only part of this design that is ours. As of today they are also the only part you cannot move — and that is backwards, so it is what we are fixing first.
The engine has export and import endpoints. They move a bank's
template: configuration, mental models, directives. Not one fact. The listing API that does
return facts hands back 11 of the 26 stored columns, and the three it drops are room,
hall and layer — plus every one of the 201,154 links. So a memory export
today loses exactly what a memory store is for, and it loses it moving into another copy of RCLL,
not just into somebody else's system.
We think that is a serious defect rather than a missing feature, because a command named
export that returns a file which is not your memory is worse than no command at all: it
is a backup you find out about later. A full dump — every column, every link, provenance intact — is
the first thing on the list, and the interchange format we are targeting is the published portable
agent memory specification rather than an invention of our own. When it lands, the direction we
intend to demonstrate first is the one that costs us something: getting your memory out of
RCLL. Everything else we have not built yet, in order →
What we still do not claim
The table above is retrieval quality on a public dataset. It is not an accuracy benchmark, and we are not going to quote one until the harness that produced it runs on your machine with the reader, judge, seed and dataset version printed next to the figure.
Two gaps we would rather name than have found. The scoring above skips the questions where the right answer is "the store does not know" — they carry no evidence labels, so a ranking metric is blind to them by construction, and independent work suggests plain files beat structured stores precisely there. And the claim this product actually rests on — that a fleet sharing a store repeats less work than a fleet that does not — is a question about finished work, not about recall, and nothing on this page tests it yet.
That one is next, and negative results get published on the same page as positive ones. If it comes out flat, that is the single most useful thing we could tell you.
Two ways to run it
They are not two products and not two editions. Same tree, same image, same migration line — the difference is one answer during install and one environment variable afterwards. That is deliberate: if the two shapes had separate codebases, a dump from one would never load into the other, and then "take it with GOD CRM" would quietly mean "start over".
Standalone
Its own Postgres, its own volume, nothing else on the box touched. Point any MCP client at it — Claude Code, your own agents, anything that speaks the protocol. This is not a demo tier with the good parts removed; it is the same store our fleet runs on. MIT, docker compose, your database.
Embedded, with GOD CRM
The same store living as a schema inside the CRM's own database: one database to back up, one to operate. RCLL keeps memory; something has to produce it, and GOD CRM is the agentic CRM this store was built for — the fleet whose numbers are on this page. If you install RCLL standalone and find you have nothing writing to it, that is the next question, not a paywall.
The installer asks once — none, embedded or
standalone — and either answer can be changed later. Changing it does not migrate what
you already stored, which is the honest reason the export work above outranks everything else on our
list.
npx fleet-memory-mcp
0.1.0, MIT — built and signed by CI, provenance attested on npm; needs a Postgres with pgvector