AUVY
DEBook a walkthrough

Your projects get a nervous system, built the way living ones are.

A language model reads what it is given and keeps nothing once the conversation ends. Living nervous systems solve that in layers: senses that never switch off, a memory that finds the right moment, reflexes that prepare a response. AUVY is built from the same layers, with a person deciding at the top, and it keeps every word a brain would let go.

Drawing: a nerve cell and its branches.

A whole nervous system

The whole system first, in the real project.

In 1991 the roboticist Rodney Brooks argued for building complete systems that act in the real world from the start, in layers that each work on their own, and making them more capable over time. Most AI research had split the problem into separate planners and memories tested in simplified worlds. Nervous systems grew his way, and AUVY is built his way.

AABBCCDDEEFF11223344B–BCC5:1SSI123401A whole nervous systemAfter R. A. BrooksPL 21
  1. SensesMail, invites, meeting notes and file changes land here first. Each one is read, dated and filed to its project as it arrives.
  2. MemoryEverything the project received, kept verbatim with its source. A question can reach it by date, by person or by what was meant.
  3. Prepared stepsWhen something changes, AUVY works out what it affects and drafts the step: a reply to the client, a task, a new date in the plan.
  4. A person decidesEvery draft stops here. Someone accepts, edits or rejects it, and only then does it leave.
  1. Build something simple and complete, put it in the real world early, and add to it there.

    AUVY works on a project's live mail and files from the first day, with the forwards, attachments and unfinished threads that come with them. Each new capability is added on top of that.

  2. Use many simple layers and no central model. A higher layer takes over a lower one only where it has to.

    Noticing, remembering and drafting are separate layers. If drafting is switched off, the project is still noticed and remembered.

  3. The world is its own best model.

    An answer cites the mail and the file version it read, so anyone can check it against the source.

  4. Keep the lowest layers fast and always running.

    New mail keeps being filed while someone reviews a draft. Nothing queues behind a slow step.

Rodney A. Brooks, “How to Build Complete Creatures Rather than Isolated Cognitive Simulators”, MIT Artificial Intelligence Laboratory ↗

How it finds

What the brain does to find a memory, and what AUVY takes from it.

Your brain drops most of the detail, but it is remarkably good at finding the one thing you need from a date, a place, a name or half a sentence. Memory research describes how it does that. These are the processes, one ring each, and what AUVY builds from them. AUVY keeps the detail the brain would lose.

LLMBB4:16123456789A1234567891–8A–AMemory around a language modelFront view, section A–ANot to scale
  1. One fact, one trace

    In the brain: Two similar meetings stay two memories. You do not get a blur of both.

    In AUVY: Mails and files are split into small pieces, one fact to a piece, so a search lands on the passage that holds it.

  2. When and where

    In the brain: You often find a memory through the day or the room it happened in.

    In AUVY: Every passage knows its date and where it lives. Ask about last week or one folder, and the search looks there first.

  3. People and relations

    In the brain: You know who works with whom without remembering where you learned it.

    In AUVY: AUVY is building a map of people, companies and meetings from what it reads, each link backed by the sentence that states it. A question about someone will follow those links.

  4. From a fragment to the whole

    In the brain: A smell or a phrase can bring back the whole scene.

    In AUVY: When a search finds one passage, AUVY also reads the section or thread around it, so the answer has the whole conversation behind it.

  5. Filing it overnight

    In the brain: While you sleep, the brain goes back over the day and files what mattered.

    In AUVY: In the background, AUVY will turn new mail and meeting notes into short facts and summaries that a search can find, each linked to the original.

  6. The newer fact first

    In the brain: When you learn that something changed, the new version takes the old one's place.

    In AUVY: When a date or a price changes, the latest statement will come first among passages about the same thing, with the old one kept on record. The red cloud on the drawing marks that change.

  7. What comes to mind first

    In the brain: What you use often comes to mind first. What you never use recedes, and eventually it is gone.

    In AUVY: Passages nobody uses will rank a little lower when two are otherwise equal. None are deleted.

  8. Knowing what you don't know

    In the brain: You can tell when you do not know something, and you say so.

    In AUVY: When the sources do not hold the answer, AUVY says it did not find it, and marks any sentence it cannot back with a source.

  9. Beyond memory

    In the brain: Attention, prediction, planning, acting.

    In AUVY: Once finding works, we take the same approach to the rest of the brain. Drawn in phantom lines: planned.

Results

Four public benchmarks, no outside model in the search.

They test whether the right passage is found in a long history and whether the answer drawn from it is correct. They use chat histories; AUVY reads project mail and files, and the task is the same.

BenchmarkWhat it testsCountAUVYOther systems
Finding
LongMemEval-SFind the conversation that holds the answer, among about 50 per question50099.4Highest score
  • gbrain hybrid97.6
  • gbrain vector97.4
  • MemPalace96.6
LoCoMoFind the session that holds the answer, in ten conversations that each ran for months1,54094.9–
Answering
MemConflictAnswer when stored facts contradict each other3,75070.4Highest score
  • MemOS55.4
  • Letta48.7
  • A-Mem44.5
  • Mem036.1
  • Memobase35.5
  • LangMem28.2
LongMemEval-SAnswer from about 115,000 tokens of chat history50091.8
  • Mastra OM (GPT-5 mini)94.9
  • Mem094.4
  • Hindsight (Gemini 3 Pro)91.4
  • Supermemory85.2
  • Zep (GPT-4o)71.2
  • Full context (GPT-4o)60.2
LoCoMoAnswer about the same ten conversations1,54091.8
  • Mem092.5
  • Full context (GPT-4o mini)72.9
BEAM 1MTen memory abilities over conversations of about a million tokens70062.4
  • Hindsight73.9
  • Mem064.1
  • Honcho63.1
  • RAG baseline30.7

Finding: the right source in the top five results. Answering: share correct, judged the way the other systems report it: LongMemEval with its own judge prompts, LoCoMo with Mem0's published prompt, MemConflict as the average over its three conflict types. BEAM: nugget score from 0 to 100, Mem0's judge prompt; the rest in percent. Other systems as they publish them, each with its own answer model; a dash where nobody publishes the same measure.

Open problems

What we are working on next.

Where the numbers leave room, and the part of the method each one points to.

Every source, not just one
Search almost always finds one source an answer needs, but not always every one. Questions that combine several conversations lose most. Filing it overnight is aimed at this.
Time and change
Questions about the newer of two facts, about time and about the order of events score lowest, and so do facts that should not have changed. The newer fact first is aimed at this.
Large company archives
EnterpriseRAG-Bench asks 500 questions over about half a million company documents, mails and chat messages. On a simpler configuration of the search AUVY scores 34.4; the leaderboard's top entry, Mixedbread with Opus 5, scores 86.6. In a search-only run AUVY found 36.4% of the documents the answers needed, so the work is in search.

Method

How the numbers were made.

What ran
AUVY's document store and search on a separate benchmark stack that holds no customer data. Each benchmark had its own workspace and its own index.
Runs
One full run per benchmark and model. Every question counts; none were dropped or retried.
Search
Hybrid keyword and vector search with the blending and reranking described above. BEAM and EnterpriseRAG ran on a simpler configuration of the same search. No external reranking model in any run shown here.
Readers and judges
LongMemEval, LoCoMo and MemConflict: GPT-6 Luna reads and judges, the reader at high reasoning effort. BEAM: GPT-5.6 Luna for both.
Judge prompts
LongMemEval and MemConflict with their own prompts. LoCoMo and BEAM with the prompt Mem0 publishes, which is how the other systems report them, so the scores can sit side by side.
No per-question rules
Nothing in the reader's instructions names a benchmark question or its answer.
Published numbers
From each vendor's own write-up. We did not re-run them. Our benchmark harness is not public.

Benchmarks and published resultsLongMemEval ↗LoCoMo ↗BEAM ↗MemConflict ↗EnterpriseRAG-Bench ↗Mem0 research ↗gbrain-evals ↗

Questions about this

Do you train AI models on our data?

No. AUVY does not designate customer data for model training.

Data handling ↗

Can the whole team work in one project?

Yes. The team shares one project: the same picture, shared conversations with AUVY, and documents edited together live. Personal notes stay personal.

How do I check what AUVY tells me?

Every answer shows its source: the file and the version it read, one click away. If a sentence has no source, AUVY names it.

Files →

Who at AUVY can see our data?

Only people in your workspace, following workspace roles, and AUVY staff with a need to operate the service, with limited production access for support and operations.

All questions, by role →

Stop rebuilding where things stand.