Your projects get a nervous system, built the way living ones are.
A language model reads what it is given and keeps nothing once the conversation ends. Living nervous systems solve that in layers: senses that never switch off, a memory that finds the right moment, reflexes that prepare a response. AUVY is built from the same layers, with a person deciding at the top, and it keeps every word a brain would let go.
Drawing: a nerve cell and its branches.
A whole nervous system
The whole system first, in the real project.
In 1991 the roboticist Rodney Brooks argued for building complete systems that act in the real world from the start, in layers that each work on their own, and making them more capable over time. Most AI research had split the problem into separate planners and memories tested in simplified worlds. Nervous systems grew his way, and AUVY is built his way.
- SensesMail, invites, meeting notes and file changes land here first. Each one is read, dated and filed to its project as it arrives.
- MemoryEverything the project received, kept verbatim with its source. A question can reach it by date, by person or by what was meant.
- Prepared stepsWhen something changes, AUVY works out what it affects and drafts the step: a reply to the client, a task, a new date in the plan.
- A person decidesEvery draft stops here. Someone accepts, edits or rejects it, and only then does it leave.
Build something simple and complete, put it in the real world early, and add to it there.
AUVY works on a project's live mail and files from the first day, with the forwards, attachments and unfinished threads that come with them. Each new capability is added on top of that.
Use many simple layers and no central model. A higher layer takes over a lower one only where it has to.
Noticing, remembering and drafting are separate layers. If drafting is switched off, the project is still noticed and remembered.
The world is its own best model.
An answer cites the mail and the file version it read, so anyone can check it against the source.
Keep the lowest layers fast and always running.
New mail keeps being filed while someone reviews a draft. Nothing queues behind a slow step.
How it finds
What the brain does to find a memory, and what AUVY takes from it.
Your brain drops most of the detail, but it is remarkably good at finding the one thing you need from a date, a place, a name or half a sentence. Memory research describes how it does that. These are the processes, one ring each, and what AUVY builds from them. AUVY keeps the detail the brain would lose.
- One fact, one trace
In the brain: Two similar meetings stay two memories. You do not get a blur of both.
In AUVY: Mails and files are split into small pieces, one fact to a piece, so a search lands on the passage that holds it.
- When and where
In the brain: You often find a memory through the day or the room it happened in.
In AUVY: Every passage knows its date and where it lives. Ask about last week or one folder, and the search looks there first.
- People and relations
In the brain: You know who works with whom without remembering where you learned it.
In AUVY: AUVY is building a map of people, companies and meetings from what it reads, each link backed by the sentence that states it. A question about someone will follow those links.
- From a fragment to the whole
In the brain: A smell or a phrase can bring back the whole scene.
In AUVY: When a search finds one passage, AUVY also reads the section or thread around it, so the answer has the whole conversation behind it.
- Filing it overnight
In the brain: While you sleep, the brain goes back over the day and files what mattered.
In AUVY: In the background, AUVY will turn new mail and meeting notes into short facts and summaries that a search can find, each linked to the original.
- The newer fact first
In the brain: When you learn that something changed, the new version takes the old one's place.
In AUVY: When a date or a price changes, the latest statement will come first among passages about the same thing, with the old one kept on record. The red cloud on the drawing marks that change.
- What comes to mind first
In the brain: What you use often comes to mind first. What you never use recedes, and eventually it is gone.
In AUVY: Passages nobody uses will rank a little lower when two are otherwise equal. None are deleted.
- Knowing what you don't know
In the brain: You can tell when you do not know something, and you say so.
In AUVY: When the sources do not hold the answer, AUVY says it did not find it, and marks any sentence it cannot back with a source.
- Beyond memory
In the brain: Attention, prediction, planning, acting.
In AUVY: Once finding works, we take the same approach to the rest of the brain. Drawn in phantom lines: planned.
Results
Four public benchmarks, no outside model in the search.
They test whether the right passage is found in a long history and whether the answer drawn from it is correct. They use chat histories; AUVY reads project mail and files, and the task is the same.
| Benchmark | What it tests | Count | AUVY | Other systems |
|---|---|---|---|---|
| Finding | ||||
| LongMemEval-S | Find the conversation that holds the answer, among about 50 per question | 500 | 99.4Highest score |
|
| LoCoMo | Find the session that holds the answer, in ten conversations that each ran for months | 1,540 | 94.9 | – |
| Answering | ||||
| MemConflict | Answer when stored facts contradict each other | 3,750 | 70.4Highest score |
|
| LongMemEval-S | Answer from about 115,000 tokens of chat history | 500 | 91.8 |
|
| LoCoMo | Answer about the same ten conversations | 1,540 | 91.8 |
|
| BEAM 1M | Ten memory abilities over conversations of about a million tokens | 700 | 62.4 |
|
Finding: the right source in the top five results. Answering: share correct, judged the way the other systems report it: LongMemEval with its own judge prompts, LoCoMo with Mem0's published prompt, MemConflict as the average over its three conflict types. BEAM: nugget score from 0 to 100, Mem0's judge prompt; the rest in percent. Other systems as they publish them, each with its own answer model; a dash where nobody publishes the same measure.
Open problems
What we are working on next.
Where the numbers leave room, and the part of the method each one points to.
- Every source, not just one
- Search almost always finds one source an answer needs, but not always every one. Questions that combine several conversations lose most. Filing it overnight is aimed at this.
- Time and change
- Questions about the newer of two facts, about time and about the order of events score lowest, and so do facts that should not have changed. The newer fact first is aimed at this.
- Large company archives
- EnterpriseRAG-Bench asks 500 questions over about half a million company documents, mails and chat messages. On a simpler configuration of the search AUVY scores 34.4; the leaderboard's top entry, Mixedbread with Opus 5, scores 86.6. In a search-only run AUVY found 36.4% of the documents the answers needed, so the work is in search.
Method
How the numbers were made.
- What ran
- AUVY's document store and search on a separate benchmark stack that holds no customer data. Each benchmark had its own workspace and its own index.
- Runs
- One full run per benchmark and model. Every question counts; none were dropped or retried.
- Search
- Hybrid keyword and vector search with the blending and reranking described above. BEAM and EnterpriseRAG ran on a simpler configuration of the same search. No external reranking model in any run shown here.
- Readers and judges
- LongMemEval, LoCoMo and MemConflict: GPT-6 Luna reads and judges, the reader at high reasoning effort. BEAM: GPT-5.6 Luna for both.
- Judge prompts
- LongMemEval and MemConflict with their own prompts. LoCoMo and BEAM with the prompt Mem0 publishes, which is how the other systems report them, so the scores can sit side by side.
- No per-question rules
- Nothing in the reader's instructions names a benchmark question or its answer.
- Published numbers
- From each vendor's own write-up. We did not re-run them. Our benchmark harness is not public.
Benchmarks and published resultsLongMemEval ↗LoCoMo ↗BEAM ↗MemConflict ↗EnterpriseRAG-Bench ↗Mem0 research ↗gbrain-evals ↗
Questions about this
Do you train AI models on our data?
No. AUVY does not designate customer data for model training.
Data handling ↗Can the whole team work in one project?
Yes. The team shares one project: the same picture, shared conversations with AUVY, and documents edited together live. Personal notes stay personal.
How do I check what AUVY tells me?
Every answer shows its source: the file and the version it read, one click away. If a sentence has no source, AUVY names it.
Files →Who at AUVY can see our data?
Only people in your workspace, following workspace roles, and AUVY staff with a need to operate the service, with limited production access for support and operations.
