# What your AI coding agent learns dies with the session

> How we made an AI coding agent keep the team knowledge base in the repo: a Claude Code Stop hook holds the turn open until knowledge is captured or declined.

Author: Nandan Maurya · Published: 2026-09-30 · Canonical: https://nisu.me/blog/what-your-ai-coding-agent-learns-dies-with-the-session/

In June I spent an hour with Claude Code chasing a socket timeout in our backend. It was set to 30 minutes, which looked like a mistake. It wasn't: an API we wrap needs that headroom on long calls. Until then, a session like that ended with the reason known to me and the agent, and nobody else. A month later a teammate, or a teammate's agent, would find the same "bug" and spend the same hour again.

This time the session didn't end until the reason was written into the repo, next to the code it explains. A few days earlier I'd set up a system for exactly that.

The problem was bigger than one timeout. Five of us work on that backend, and at the time we used a mix of tools: Cursor, Codex, Claude Code, or ChatGPT with copy and paste. Each session knew a slice of why the system looks the way it does, and the slice was gone when the session closed. (Claude Code's auto memory keeps some notes, but on one machine, for one person; nobody else on the team sees them.) The same month we found three record types missing from a list that every new type has to be added to. That rule had lived in exactly one head.

Some of this deserves a code fix too, like a comment on the constant or a test on the list. But the reason behind a decision, or a choice made in chat, often has no single line of code to live on.

So I made the agent the keeper of a knowledge base committed to the repo, and made the job enforced instead of requested. In Claude Code, a *turn* is the agent's work in reply to one message, and a *Stop hook* is a script that runs when the agent is about to end that turn. When a session has probably learned something and nothing was written down, our Stop hook holds the turn open until the knowledge is captured or the agent explicitly declines. It asks at most twice. In the nine weeks after we set it up, 77 of the 208 commits in that repo touched the knowledge base, most of them bug fixes and features.

If you just want the setup, hand this to your agent:

```text
Set up the "keeper" system in this repo, using
https://github.com/nisu-me/agentic-keeper.
Fetch REPLICATE.md from that repo first and follow it: ask me the
questions it lists, adapt the config to this codebase, copy the core
files verbatim, and run its verification checklist before you call
the setup done.
```

That prompt tells your agent to follow a document from my repo, so read [`REPLICATE.md`](https://github.com/nisu-me/agentic-keeper/blob/main/REPLICATE.md) and the [two hooks](https://github.com/nisu-me/agentic-keeper/tree/main/.claude/hooks) first. They're about 340 lines of dependency-free Node.

> **Run the install with a strong model at high reasoning effort.** It's a one-time job full of judgement calls: which paths are load-bearing, and what to do with what your CLAUDE.md already knows. Everyday sessions don't need it.

## A knowledge base in the repo, kept by the agent

The knowledge base is plain markdown in `docs/`: an index with placement rules, Architecture Decision Records (ADRs: one short file per decision, with the reasons), a gotchas file for traps and quirks, and runbooks. Humans and agents follow the same rules.

I chose the repo over vendor memory on purpose. Copilot Memory, for example, is shared per repo, but it lives on GitHub's side and deletes entries unused for 28 days. Team knowledge should be shared, reviewed, versioned and greppable, which means git.

Five rules govern it:

1. **One fact, one place.** Link from elsewhere; never duplicate.
2. **Everything is reachable from the index.** A new file means a new index row.
3. **ADRs are append-only.** Supersede a decision; never rewrite it.
4. **Cite code.** A fact without a `path/to/file` anchor goes stale unnoticed.
5. **No secrets.** The secret scanner covers the docs directory too.

We followed a sixth rule for weeks before anyone wrote it down: one fact per line. I noticed it while switching between models: with each fact on a single line, `grep` returns whole facts, edits replace whole facts, and diffs stay readable, whichever model is doing the work. Because nobody had written it down, the agent started drifting from it. It's rule six now.

The agent isn't just asked to document things, because asking fades over a long session. Anthropic's [best-practices guide](https://code.claude.com/docs/en/best-practices) draws the same line: "Unlike CLAUDE.md instructions which are advisory, hooks are deterministic and guarantee the action happens." (CLAUDE.md is the instructions file Claude Code loads into every session.)

## The loop, from session start to turn end

**Inject.** A SessionStart hook prints a short block into every new session: where the index is, which knowledge files exist, the latest ADRs, and the keeper mandate.

```text
[your-repo knowledge base]
Index + placement rules: docs/README.md
Facts: docs/knowledge/{architecture,gotchas,integrations,...} | Conventions: ... | Runbooks: ...
Recent ADRs: 0003-..., 0002-..., 0001-...
Keeper mandate (non-optional): unfiled knowledge dies with the session. ...
Close every substantial session with an explicit "Knowledge check:" line ...
```

Reading stays advisory: our CLAUDE.md tells the agent to check the gotchas before "fixing" anything surprising, and names the 30-minute timeout as the example.

**Place.** The index has a placement table:

| You have... | It goes to... | How |
|---|---|---|
| An architectural decision (chose X over Y, with reasons) | `adr/` | `/adr` skill |
| A non-obvious fact, trap, or quirk of this codebase | `knowledge/gotchas.md` | `/capture` skill |
| A fact about an external service (webhooks, retries, limits) | `knowledge/integrations.md` | `/capture` skill |
| A new or changed team convention | `conventions/` | edit + announce |
| Steps to perform an operational task | `runbooks/<task>.md` | new file + index row |

Writing things down fails when the writer has to decide where they go. The table decides.

**Capture.** Two skills (saved prompts the agent runs as commands), `/capture` for facts and `/adr` for decisions, make writing something down a one-word action at the moment it comes up.

**Enforce.** The Stop hook, next section.

**Repair.** A script checks that the index is complete, links resolve and ADR numbers are in order. A reviewer agent that reports without editing handles the judgement calls, like duplicates and claims the code has moved on from.

## The Stop hook that holds the turn open

When the agent is about to end its turn, the Stop hook looks at what the session edited and what you said. It holds the turn open if one of three signals fires:

1. Code: knowledge-bearing paths were edited (data model, auth, payments, webhooks; you configure the list) and the knowledge base wasn't touched.
2. Conversation: your own messages contain decision language, like "we decided", "turns out" or "from now on". A decision made in chat never touches code, so no diff will carry it.
3. Bulk: eight or more files edited with nothing recorded in the docs.

The hook then sends the agent a reminder and the turn continues. It never refuses *you*: your request still gets done, the agent just can't wrap up yet. It has to capture the knowledge now, or end with the literal line `Knowledge check: nothing to record — <why>`. The explicit "no" matters most: a polite suggestion gets ignored, and a required declaration gets thought about. If the agent already wrote the line on its own, the hook stays quiet. If the next reply has neither a capture nor the line, it reminds once more and then lets go.

The line is still text, and an agent can write it without meaning it; [others have watched Claude forge a gate's pass marker](https://zenn.dev/kok1eeeee/articles/claude-code-stop-hook-quality-gate-gaming). The hook only checks that the words `Knowledge check:` appear in the reply. The reason stays in the session log for people to read, and anything that does get captured goes through code review like any other diff.

Here's the core of the hook, simplified. The [full source](https://github.com/nisu-me/agentic-keeper/blob/main/.claude/hooks/keeper-check.mjs) adds the guards, the per-session state and the log parsing:

```js
// CONFIG — the only part you adapt
const KNOWLEDGE_HINTS = ['/schema/', '/auth/', '/webhook/', '/payments/', '.env.example'];
const KB_PATHS = ['docs/', 'CLAUDE.md'];   // 'dir/' = root-relative; bare name = any depth
const BULK_THRESHOLD = 8;
const CONV_MARKERS =
  /\b(we (decided|agreed)|decided (to|on|that)|agreed (to|on|that)|confirmed|let'?s go with|we'?ll go with|from now on|going forward|deprecated?\b|turns out|the (actual |real )?reason (is|was)|always (use|do)|never (use|do)|rule of thumb)\b/i;

// INVARIANT — the decision
if (paths.some(isDocPath) || kbTouchedOnDisk(root, sessionStart)) process.exit(0); // filed → satisfied
const signals = [];
if (knowledgeCount > state.codeCount) signals.push('knowledge-bearing code edited');
if (convHits > state.convCount)       signals.push("decision language in the user's messages");
if (paths.length >= BULK_THRESHOLD && !state.bulkNudged) signals.push('bulk edit, zero docs');
if (signals.length) {
  saveState({ ...counts, awaitingExplicitNo: true });
  // Stop-hook feedback: the agent keeps going with this as its next instruction
  nudge(`Keeper check — ${signals.join('; ')}. ` +
    'File it NOW (/capture, /adr) or end with "Knowledge check: nothing to record — <why>".');
}
```

The details took more iteration than the idea:

- **Scan your messages only.** The agent's own replies are full of decision language, and scanning them makes the hook trigger itself. Fenced code blocks are stripped first, so a log pasted inside one doesn't count.
- **Precision over recall.** The pattern is a short, hand-picked list of phrases, each signal fires once per new occurrence, and any edit to the knowledge base silences the hook for the rest of the session. A gate that fires every turn trains people to switch it off.
- **Never wedge a session.** Every error path exits cleanly, the keeper asks at most twice, and Claude Code itself [ends the turn](https://code.claude.com/docs/en/hooks) after eight continuations in a row. A hook that can crash your session gets deleted by Friday.
- **Don't trust the transcript alone.** Claude Code writes its session log asynchronously, and the docs warn it may lag when a Stop hook runs. So the same script also runs after every edit and logs the path.
- **Send a reminder, skip the error.** `{"decision": "block"}` also keeps the turn going, but Claude Code shows it as a *Stop hook error*, which reads as "something broke". The keeper sends its reminder as [`additionalContext`](https://code.claude.com/docs/en/hooks#stop) instead: the same continuation, without the error.

The keeper also tests itself. It fails safe, so if Claude Code's log format drifts, a dead keeper looks identical to a healthy quiet one. A smoke script replays the hook's log parsing against real sessions and runs seventeen synthetic cases through it, and all seventeen pass.

## What already exists

Keeping agent knowledge in files is mainstream. [AGENTS.md](https://agents.md/), "a README for agents", is used by over 60,000 open-source projects, alongside CLAUDE.md, Cursor rules and Cline's [Memory Bank](https://docs.cline.bot/best-practices/memory-bank). The habit around these files is to add a line whenever the agent gets something wrong. Boris Cherny, who created Claude Code, [described his team doing that](https://x.com/bcherny/status/2007179832300581177) with a shared CLAUDE.md, and Mitchell Hashimoto [does the same with AGENTS.md](https://mitchellh.com/writing/my-ai-adoption-journey): "Each line in that file is based on a bad agent behavior." It works as long as someone remembers to add the line. (ADRs themselves go back to Michael Nygard's [2011 post](https://www.cognitect.com/blog/2011/11/15/documenting-architecture-decisions).)

Vendor memory is moving the other way. Claude Code's auto memory stays on one machine, Copilot Memory lives on GitHub's side and deletes entries unused for 28 days, and Cursor [removed its Memories feature](https://forum.cursor.com/t/are-my-memories-gone/144057) in version 2.1.

Explicit capture steps exist too. Every's [compound engineering](https://github.com/EveryInc/compound-engineering-plugin) makes "capture what you learned" a named step in its loop, and Anthropic's own [claude-md-management](https://github.com/anthropics/claude-plugins-official/tree/main/plugins/claude-md-management) plugin has a command that folds a session's learnings into CLAUDE.md. Both run when someone invokes them. OpenAI's [harness engineering](https://openai.com/index/harness-engineering/) post goes furthest: `docs/` is the system of record, CI checks that it's current, and a doc-gardening agent opens fix-ups. CI can check what reached the docs, but not what was decided in a chat that never did.

The mechanism isn't mine either. Anthropic's docs describe a Stop hook as a deterministic gate, and Claude Code's own [`/goal`](https://code.claude.com/docs/en/goal) is built on one. Daniel Miessler's [LifeOS](https://github.com/danielmiessler/LifeOS) holds the turn until its state document is updated or the agent says no update is needed. Two small projects, [simplegraph](https://github.com/karstom/simplegraph-agentic/pull/13) and [avenoxbeyin](https://github.com/avenoxai/avenoxbeyin), hold it when files change without a note. On the conversation side, [claude-reflect](https://github.com/BayramAnnakov/claude-reflect) and [ECC's ADR skill](https://github.com/affaan-m/ECC/tree/main/skills/architecture-decision-records) notice corrections and "we decided to…" in chat, but only queue or suggest.

Rahul Garg's [Context Anchoring](https://martinfowler.com/articles/reduce-friction-ai/context-anchoring.html) names the underlying problem: "the reasoning behind decisions degrades faster than the decisions themselves."

What's different here is the combination: a team knowledge base in the repo, a gate that fires only on specific signals, and one signal that is your own words in chat. Several tools hold the turn to force capture, and several detect decision language. I haven't found one that holds the turn because the user stated a decision.

## Try the ten-minute demo

The repo ships a `demo/` directory: a small todo app with the keeper installed. Every AI tool already knows the domain, so your attention stays on the keeper.

```bash
git clone https://github.com/nisu-me/agentic-keeper && cd agentic-keeper/demo
npm i && npm start     # terminal 1: zero dependencies, serves on :3000
claude                 # terminal 2, also from demo/ (its own project root)
```

Paste the request from the demo README: *"Add due dates to tasks, with overdue highlighting. From now on, store every date as an ISO 8601 string."* The agent has to edit a `/schema/` path, and your message contains a decision, so both signals apply. If the agent captures the knowledge before finishing, or declines with a `Knowledge check:` line, as the mandate asks, the hook stays quiet. If it tries to finish with neither, the hook holds the turn and names both signals.

The hooks run on your machine, so read them before you accept the trust dialog.

## Did it work? Nine weeks of numbers

The keeper went into our backend repo in mid-June. These numbers cover the nine weeks to mid-August, from git history, without merges and our release bot's commits. There's no before-number to compare with: until then the repo had no knowledge base and no CLAUDE.md.

- 77 of 208 commits (37%) touched the knowledge base.
- 60% of those 77 were `fix` and `feat` commits, not documentation sprints. Knowledge landed with the work that produced it.
- The gotchas file grew from 17 entries to 96, and ADRs from 7 to 14. The first seven ADRs record earlier decisions and are marked as retroactive; the seven since were written as the decisions were made.
- Four of the five people committing to the repo captured knowledge. Once the team saw the keeper working, I moved the rest of the team to Claude Code, where the enforcement runs.

For the first month I also went through my own session logs. Claude Code keeps them on each developer's machine, so these cover my sessions only. Across 21 sessions the keeper held a turn 11 times: six on the code signal and five on decision language in my messages, about one hold every two sessions. I didn't track how each one ended.

The logs had one surprise: `/capture` and `/adr` were used only a handful of times. Most knowledge went in as direct edits in the middle of work, because the agent knows where things go. The placement table mattered more than the commands, and the enforcement more than either.

What changed for the team is harder to measure. I don't have before-and-after timings, so these are observations:

- Traps tend to get found once. The first person to hit one captures it, and later sessions start from it.
- Design discussions start from the record. An audit of our auth setup, captured one day, became the basis of an ADR for a new feature the next.
- Context stopped being per person. It lives in the repo, reaches everyone on the next pull, and gets reviewed in PRs. The developers liked working this way.
- The cost: the held turns above, plus knowledge diffs to review alongside code.

## Installing it with one prompt

There's no plugin, npm module or installer script; the prompt at the top is the installer. That part isn't new (claude-memory-compiler installs the same way). What I care about is the contract in [`REPLICATE.md`](https://github.com/nisu-me/agentic-keeper/blob/main/REPLICATE.md), a spec addressed to your agent. It lists the questions to ask you, such as which directories are load-bearing and what your domain accumulates (a payments product grows a `money-flows.md`, an ML product a `pipelines.md`). It says which files to copy byte for byte, which to adapt, and which invariants must survive. It ends with a verification checklist, including a live test the agent may not claim to have passed without seeing it.

I think of it as distillation rather than forking. The scaffold is the distillate of a production system: the invariants and the reasons, with every company-specific fact removed. Your agent re-derives the setup in your context. It's the relationship shadcn/ui has with component libraries: you don't depend on the package, you take the source and own it.

> **You distill the keeper, not the knowledge.** The machinery transfers. The gotchas and ADRs stay home, and every repo builds its own.

It has been installed four times so far: our frontend repo, the demo, the repo behind this blog, and a side project. Each of the first three taught the spec something, and those lessons are in it now.

The side project was the hardest: a small semantic image-search tool in Python, whose CLAUDE.md already held a page of decisions and gotchas. I pasted the prompt into a fresh session running a strong model at high reasoning effort. The agent read the hooks before installing them, asked four questions, and was done in about five minutes. It turned the existing decisions into ADRs marked as retroactive, ran the self-test and the project's own tests, and handed the live steps back to me.

The first live session then found a real bug. I asked for a new model preset. The agent edited an encoder file and wrote three knowledge updates with shell heredocs instead of the Edit tool. The keeper only watched tool edits, so it saw "knowledge-bearing code changed, knowledge base untouched" and held the turn. The agent checked `git diff`, said it had already captured everything, and closed with a `Knowledge check:` line. No harm done, but the reminder was wrong. The fix shipped that night: the hook now also checks the knowledge base's files on disk.

The install takes minutes. For a team, I'd budget a week or two of trial before judging it.

## What it doesn't do

- It only enforces in Claude Code. Other tools can read the knowledge base, since it's markdown in the repo, but nothing makes them write to it. Codex CLI's Stop hook accepts the same `{"decision": "block"}` and continues the turn, and Gemini's `AfterAgent` can force a retry, but Cursor's `stop` hook can only send a follow-up message. The spec includes an untested port map.
- Code changed with shell commands doesn't show up in the edit log. Knowledge written by shell does count, because the knowledge base is checked on disk, and so does a `git pull` that brings in someone else's docs change.
- The conversation markers are English only. A team that chats in another language has to adapt one regex, or that signal silently never fires.
- The explicit "no" is checked for presence, not quality.
- Precision over recall means misses. We trade recall for the hook still being enabled in month six.
- Drift checks run on demand, not in CI, and knowledge quality isn't measured.

## Keep agent memory in the repo

Our rule of thumb fits in a sentence: explained it twice, capture it; decided something, write an ADR.

Treat it as an approach rather than a format. Teams add runbooks, grow the knowledge categories their domain needs, and reshape the rules. The parts that matter are the ones the spec marks invariant.

The claim I'd defend beyond this project:

> **Agents need institutional memory, and your repo is the right place for it.** Markdown, placement rules, append-only decisions, and one hook that won't let a session end silently.

---

*The scaffold, the demo and the spec: [github.com/nisu-me/agentic-keeper](https://github.com/nisu-me/agentic-keeper). The backend numbers come from a private production repo.*

---

Markdown copy of https://nisu.me/blog/what-your-ai-coding-agent-learns-dies-with-the-session/. Site index for agents: https://nisu.me/llms.txt
