How should an AI agent use long-term memory without flooding its context window? Load only a small resident tier at startup, and route everything else through an index the agent reads first, opening just the one or two files whose description matches the task. On the nestwork author's own nest this cut session startup from about 69,600 tokens to about 640, while a typical task still reads only around 3,600.
This article explains the two tiers, the topic index that makes the second tier navigable, what the index checker catches, and why more context is not better context.
Why "load everything" stops working
The first version of most memory setups injects the memory file into every session. That works while the file is small. It stops working for two reasons.
First, memory grows faster than windows. The author's nest, measured in September 2026 with the o200k_base tokenizer:
| Scenario | Files | Size | Tokens |
|---|---|---|---|
| 2.x-style full startup (rules, strategy, all shared and agent memory, workflows) | 37 | 224 KB | ~69,600 |
| 3.x startup, resident tier only | 2 | 2.6 KB | ~640 |
| 3.x task: a git operation (resident + index + one topic) | 4 | 11.8 KB | ~3,600 |
| Every memory file in the nest | 180 | 1.2 MB | ~369,000 |
The whole nest no longer fits in most context windows, so selection is not optional.
Second, even when it fits, more context is not free. Liu et al. found that models' performance "significantly degrades when models must access relevant information in the middle of long contexts" (Lost in the Middle). Anthropic's engineering team calls the same effect context rot and recommends "the smallest possible set of high-signal tokens" (effective context engineering). nestwork's protocol puts it as a design rule: tune file sizes for retrieval quality, not raw window size, because attention still degrades with token count.
Tier 1: resident context
Since protocol 3.0, an agent loads only three files at session start (AGENTS.md §1):
queen/agent-rules.md, your core behaviour rulesshared/resident.md, small current cross-agent facts and retrieval pointers, if presentagents/<host>/<agent-id>/resident.md, small instance-specific facts and pointers, if present
Each has a byte budget: 4,096 bytes for the rules, 4,096 for the shared resident file and 2,048 for each agent resident file, at most 10,240 bytes per startup. python3 scripts/maintenance/check-resident.py checks them before you commit and never truncates anything; an oversized file has to be reviewed and moved, not silently cut.
Where a SessionStart hook is installed, it prints the resident paths as a READ-ON-START manifest and the rest as READ-ON-DEMAND. It prints paths, not contents, so one oversized file cannot crowd out the rest of the list.
What belongs in resident: facts that must apply before anyone looks them up, essential boundaries and a few pointers. Project backlogs, incident stories and old status do not. The protocol's test is whether an entry has to take effect "when nobody went looking for it". When in doubt, choose on demand.
Tier 2: on demand, routed by a topic index
Everything else is on demand: strategy, memory history, projects, workflows, carryover and the mailbox. The agent starts from the current task, searches headings or keywords, and reads only the matching sections.
That works for a few files. For hundreds, protocol 3.1 adds optional topic memory. A memory scope, either shared/ or one agent directory, opts in by placing two markers in its memory.md:
# Shared memory
<!-- nestwork:topic-index:begin -->
<!-- generated by scripts/maintenance/memory-index.py; edit topic front matter, not this block -->
- [`engineering.md`](engineering.md) — Build, test and release conventions; read before changing CI (2026-09-20, 6.1 KB)
- [`hosts.md`](hosts.md) — Machine-specific quirks, ports and paths; read before touching a host's setup (2026-09-27, 3.4 KB)
<!-- nestwork:topic-index:end -->
From then on memory.md is a routing table, and the facts live in topic files. Each topic file starts with front matter:
---
description: Machine-specific quirks, ports and paths; read before touching a host's setup
updated: 2026-09-27
---
The description is the only thing an agent sees before deciding to open a file, so it is written as a trigger (when to read) rather than a title. The retrieval path becomes: resident summary, then the memory.md index, then the one or two topic files whose description matches.
It works like skills
If you use Claude Code skills, the pattern is familiar. Claude Code's docs say skill descriptions are loaded into context, "but full skill content only loads when invoked" (Claude Code skills). Claude Code's own auto memory does something similar: it loads the first 200 lines or 25 KB of MEMORY.md and reads topic files on demand (Claude Code memory).
nestwork applies the same idea to a memory store shared by many tools and machines, with one difference: the index is generated, not hand-written, so it cannot drift from the files it describes.
What memory-index.py --check catches
The index is produced by scripts/maintenance/memory-index.py. Run it after changing topic files; run it with --check before committing or in CI:
python3 scripts/maintenance/memory-index.py # regenerate indexes
python3 scripts/maintenance/memory-index.py --check # fail instead of writing
--check fails when:
- the index is stale, meaning it no longer matches the topic files' front matter
- a topic file has no
description, which would make it invisible to retrieval - a topic file is over 32 KB (it warns above 16 KB)
- a topic is nested deeper than
<topic>/<subtopic>.md
It also warns when a description is longer than 300 characters or the generated index block passes 8 KB, both signs that topics should be consolidated. Text outside the markers is preserved, and reserved paths such as resident.md, outbox/, local/, carryover/ and anything starting with _ or . are never treated as topics.
Keeping topics useful
A generated index only helps if topics are cut well. The protocol's rules (AGENTS.md §6):
- Reuse before creating. Read the index first and extend the topic that already covers the fact.
- Split by when it is needed, not by who wrote it. A topic is what an agent loads for one kind of task.
- Agents shape only their own scope. New, renamed or merged topics in
shared/happen only during a reviewed distillation, so 30 agents don't grow three near-identical user-profile files.
Migrating an existing monolith is manual and reviewed: split memory.md by its headings, moving text verbatim; add a trigger-style description to each file; replace memory.md with a header and the markers; run the indexer and --check; point resident.md at the index. The steps are in context loading.
To see the numbers for your own nest:
python3 scripts/maintenance/measure-context.py --task shared/<topic>.md
Bytes are exact; tokens use tiktoken if installed, otherwise a calibrated estimate.
FAQ
Isn't this just retrieval-augmented generation?
It is retrieval, but without embeddings. The agent reads a short, human-readable index and chooses files by description. That keeps it transparent and debuggable in git, at the cost of fuzzy semantic matching.
Do I have to switch to topic memory?
No. It is opt-in per scope. Scopes without the markers keep single-file memory, nothing migrates automatically, and moving from 3.0 to 3.1 needs no bootstrap refresh.
What if the agent picks the wrong topic?
Usually the description is the problem. Rewrite it as a trigger that says when to read the file, then regenerate the index. Resident files can also point directly at the few topics that matter most.
Will a bigger context window make this unnecessary?
Unlikely. A larger window does not stop attention from degrading as tokens accumulate, and the author's full nest is already far larger than typical windows.
Related
- Git-native agent memory: why git is enough
- Memory distillation: merging many agents' memory
- AGENTS.md vs CLAUDE.md vs memory
- Agent memory benchmarks
- Long-term memory in practice
Give your agents a memory
nestwork turns one private git repo into shared, persistent memory for Claude Code, Codex, Gemini, Kimi and more.
Use this template → ★ Star on GitHub