How to write instructions for AI agents (and keep them working)
Write instructions for an AI agent as a short set of rules, each in one place, loaded by the job the agent is doing, and keep the facts about customers, deals and people on their records, never in the instructions. Then maintain the rules like code: a change replaces the old rule, history goes to an archive the agent never reads, and a monthly check looks for contradictions. That is how we rewrote the rulebook our own AI agent works from, cutting it from 266 KB to 144 KB without losing one of its 918 rules.
Our case: 266 KB of rules, cut to 144 KB
At Odysi, an AI agent does much of our own CRM, email and proposal work. It logs emails, keeps deal logs and tasks up to date, drafts replies that a person sends, prepares calls and writes the weekly pipeline report. It is Claude, working through our CRM’s connector over MCP, the open standard for connecting AI applications to other systems. Everything it does follows a rulebook we call the Playbook.
In about four weeks the Playbook grew to 266 KB across 11 documents. Every correction we gave the agent became more text, and more text made the agent less reliable. On 6 October 2026 we rewrote it. The result:
| Measure | Before | After |
|---|---|---|
| Live rule text | 266 KB in 11 docs | 144 KB in 14 docs |
| Read at the start of every conversation | 36 KB | 16 KB |
| Read for one prospecting run | about 167 KB | about 60 KB |
| Contradictions between docs | 17, found by the audit | Resolved; a few clarity points left for the first monthly clean-up |
| Rules stated in more than one place | 348 | None: each rule has one home, the others point to it |
| Rules lost in the rewrite | 0 of 918 |
Odysi’s own figures for its internal Playbook, 6 October 2026.
Nothing was deleted. Most of the bulk was history, and it now sits in one archive document that the agent never reads to do work. The rest of this guide is what we learned, written so any company can use it.
Why instructions rot: a context engineering problem
A rulebook the agent reads on every run degrades the same way a long, messy conversation does. Anthropic calls the work of choosing what goes into a model’s context context engineering, and notes that as the context grows, the model’s ability to recall what is in it goes down. Chroma’s Context Rot study tested 18 models in July 2025 and found that performance falls as the input gets longer, even on deliberately simple tasks.
Contradictions make it worse. Anthropic’s documentation for Claude Code warns that when two instructions contradict each other, Claude may pick one arbitrarily, and suggests keeping each instructions file under 200 lines. Our Playbook had both problems: a lot of text that was not an instruction, and rules that disagreed.
The cost was concrete. A scheduled run that read only the entry document drafted replies from a summary that no longer matched the writing rules. Every run paid to read 100 KB or more of text that could not change what it did. And when two rules disagreed, the agent picked one without saying so.
Six ways a rulebook rots, with one example each
Our audit found six patterns behind the growth. You will probably recognise some of them in your own prompts, CLAUDE.md files or agent docs.
- No. 01Summaries that drift. The entry document restated the other docs in short form: 153 of its 261 rules were copies. When a detail doc changed, the summary did not. Example: the summary kept an older version of a writing rule after the writing doc had replaced it, and a run that read only the summary followed the old one.
- No. 02New rules added, old ones kept. A correction was appended as a new dated paragraph, “(Thomas, 18 Sep 2026)”, without removing the rule it replaced. Both versions stayed live, and the agent had to guess which one counted.
- No. 03History mixed with rules. Our prospecting doc was 70 KB, and only about a fifth of it was rules; the rest was batch lists, research and snapshots. The LinkedIn doc was 80% a table of invitations.
- No. 04Stories instead of rules. Three paragraphs on why a deck had to be sent again, with the lesson buried in the middle. The rule is one line: “Check the send copy has no speaker notes before handing it over.”
- No. 05Live data inside the rules. Lists of firms to contact next sat inside the rules and contradicted themselves after a few edits: one firm was both “dropped” and “back in the pool”.
- No. 06Fragile cross-references. “See section 4.2” pointed at sections that had been renumbered or no longer existed. Now it reads “see the reply playbook, Objections”: a heading name survives renumbering.
Where facts live: records, instructions and memory
Half of those patterns are not fixed by better writing. They are fixed by putting each kind of information where it belongs. An agent system has five parts, and instructions only work when the other four are in place:
| Part | What it is | Ours |
|---|---|---|
| Interface | Where people talk to the agent | Claude, in chat and as scheduled tasks |
| Instructions | What the agent should do and how | The Playbook |
| Records | The state of each thing being managed: who, what was agreed, the next step | Our CRM: companies, people, deals, deal logs, tasks, emails, call transcripts |
| Tools | What the agent can read and change | Connectors to the CRM, email, calendar and files |
| Triggers and review | When it runs and where a person checks | Hourly and daily scheduled tasks; drafts a person sends |
The split between instructions and records matters most. Instructions say how to work; records say what is true about each case. To let an agent act on its own, give every kind of information one home:
| Information | Where it lives | Example |
|---|---|---|
| Facts about a company, deal or person | Its record in the system of record: CRM fields, deal log, tasks, emails, transcripts | Who decides, what was agreed on the last call, the next step and its date |
| How to do the work | The instructions, one home per rule | How a follow-up is timed; what is never sent without a person |
| Values several rules share | One Facts section in the entry document | Team names and ids, page links |
| Live lists | The records, or one compact section with one line per item | Firms to contact next, segment status |
| One person’s preferences | That person’s own rules | How long a call is blocked in the calendar |
| Why a rule changed | An archive the agent never reads for work | Old change notes and batch lists |
Facts never go into instruction docs or memory files. A deal stage written into the rules is wrong by next week, and a fact in one agent’s memory is invisible to the rest of the team. On the record, the agent finds it with a tool when it needs it. Anthropic describes this as just-in-time context: the agent keeps “lightweight identifiers” and loads the data at runtime. Claude Code’s own auto memory is for learnings and patterns about how to work, which is the right use for memory.
The other half is that the agent writes to the records as it works: it logs the email, moves the deal, creates the task. The next run then starts from the facts, not from a rerun of the inbox. For a small team the records can be a spreadsheet or one file per client; a database can come later. What matters is that each fact has one home and the agent keeps it current.
How to write instructions for AI agents: ten principles
Each principle answers one of the rot patterns. The short rule comes first, then why.
- 01Start from the jobs, not from a rulebook. List every recurring job the agent does (a morning check, a reply, a proposal, a report), what triggers it and what done looks like. Then write only what the agent cannot know for each job. Why: a rulebook written by topic collects everything; one written from jobs stops at what the work needs.
- 02One home per rule, and per fact. Anywhere else, a pointer (“see the pricing rules, Deal value”), never a copy. Why: a copy is a second version waiting to drift. Most of our 17 contradictions were copies nobody updated.
- 03The entry document routes; it does not summarise. Three parts only: hard rules, a jobs table naming the docs each job reads, and the Facts. Keep it under about 10 KB. Why: the always-loaded doc is the most expensive and most trusted text in the system, and a summary there overrides the detail without anyone noticing.
- 04Every doc opens with a read-when line. One line saying when to read it, matched by the jobs table. Why: the agent loads only the rules its job needs, and a run cannot skip a doc it needs because nobody told it to read it.
- 05Split by when the work happens, not by topic. Why: the agent should never need three docs to do one step.
- 06Rules, not stories. One imperative line with the agent as the subject, and a “because” only when it changes how the rule is applied. Why: in a story, the model has to guess which sentence is the rule.
- 07One example beats a paragraph. One good line and, when useful, one “Not:” line. Why: models copy patterns more reliably than they apply descriptions. Anthropic calls examples the “pictures” worth a thousand words in its context engineering guide.
- 08Leave out what the model already knows. “Be polite” and “write clearly” add length and nothing else. Write down the taste, facts and decisions that are yours.
- 09A change replaces; it never appends. Find the existing rule and change it in the same edit. Why: dated paragraphs stacked on older ones were how two versions of the same rule ended up live.
- 10Write precedence down. Ours: the chat wins over personal rules, and personal rules win over team docs. If two docs still disagree, the agent follows the one that asks for less (no send, no delete, ask first) and says so. Why: without it, the agent picks silently.
If you want an agent built this way for your own operations, that is what our AI process automation work covers.
How to run a rewrite with separate agents
Once a rulebook has rotted, editing it in place does not get you out. We ran the rewrite as a pipeline of agents with one job each, so that no agent both changed the rules and judged its own changes. Anthropic describes the same idea as a sub-agent architecture: focused agents that hand back a short summary of their work. We made the decisions; the agents did the work.
- Step 1Snapshot. All 11 docs were saved at their current versions, so every agent worked from the same text.
- Step 2Audit and ledger, in parallel. An architect agent measured each doc (rules against stories, history and copies), named the rot patterns and designed the target: which docs, what goes in each, size budgets, the rule format. A ledger agent read every line and listed every operative rule, 918 of them, each with its source, every place it was repeated (348 rules appeared more than once) and every conflict (17).
- Step 3A written spec. One page fixed the decisions before anything was rewritten: the doc list, which doc owns which rules, the format, and how each conflict is settled (usually the newest rule wins). Two decisions were ours to make: publish live with a safety check, and archive history instead of deleting it.
- Step 4Rewriters, one per topic. Five agents each wrote their docs, moved history word for word to the archive and filled a crosswalk saying where each of their ledger rules went: kept, merged, moved, archived or cut, with a reason.
- Step 5Two independent checks. A preservation check looked for each of the 918 rules in the new docs: none was missing, and four had changed meaning and were restored. A consistency check read the 14 new docs as one system (contradictions, duplicates, broken references, routing, format), made 28 fixes and listed the questions that needed a person.
- Step 6Decide what is left. We settled the open questions that risked a wrong action and reset each doc’s size budget to its real size.
- Step 7Publish and read back. Each doc went through the CRM connector, was read back and was compared with the local text character for character. All 14 matched.
- Step 8Clean up around it. The old proposals doc was retired, with its history kept, and the scheduled tasks that cited section numbers now cite heading names.
The ledger and the crosswalk are what make it safe. “It reads well” proves nothing about what was lost; a list of 918 rules, each with a destination, does.
How to keep agent instructions clean
A rulebook stays short only if something checks it regularly. Ours now has five habits:
- Size budgets per doc, from 3 KB to 22 KB. The two docs read on every run stay under 17 KB together. When an edit goes over budget, rules are merged or tightened in the same edit; a rule is never dropped to make room.
- Change notes of ten lines at most at the end of each doc. The eleventh pushes the oldest to the archive. A change note never holds a rule.
- A monthly clean-up on the first Monday of the month. It reports contradictions, rules stated twice, stale facts, references to headings or docs that no longer exist, and docs over budget, and it changes nothing without the owner.
- References by heading name, never by section number, in the docs and in every scheduled task’s prompt, so renumbering cannot break them.
- A test with a fresh agent. Give a new session only the entry document and one job to do end to end. Every place it hesitates or opens the wrong doc is a routing or wording fault.
Rule or one-off? Most new rules come from corrections. Feedback becomes a rule only when it will apply again:
| Someone says | It is |
|---|---|
| “Never call our product a chatbot; say AI agent.” | A team rule: it applies to every future email and deck |
| “Make this email shorter.” | A one-off: it is about this draft |
| “Block calls for 45 minutes by default.” | A personal rule: it is how this person works |
| “Send this one tomorrow at 9.” | A one-off |
Before writing a rule, search for the topic and change the existing rule if there is one. Decide whether it is a team rule (only the owner of the business changes those) or a personal one. Put it in the doc whose read-when line covers the moment it is needed, write it short, remove what it replaces, add one change note and tell the person in one line what changed.
What never goes in the rules: passwords and keys, the details of one task, research results and deal state (they belong on the records), and anything a person asked not to keep. A full rewrite like ours is worth doing only when the monthly clean-ups stop keeping up.
A template for the entry document
This is the shape of our entry document, with our details removed. Copy it and fill it in with your own jobs and facts. The same shape works for a system prompt, a CLAUDE.md or an AGENTS.md file.
# Instructions Read when: always, before any job. ## Hard rules - Never send an email. Save it as a draft for a person to send. - Never delete a record. Archive it. - Write everything in the CRM in English. - Facts about a company, deal or person go on its record, never in these docs. - If two docs disagree, follow the one that asks for less (no send, no delete, ask first) and say which you followed. ## Precedence The chat wins over personal rules; personal rules win over team docs. ## Jobs | Job | When | Read | | Morning check | Daily, 08:00 | pipeline | | Reply to a lead | A reply arrives | reply-playbook, writing-rules | | Prepare a call | A call is booked | calls-and-meetings | | Proposal | Asked in chat | client-documents, pricing-rules | | Pipeline report | Mondays | reporting-rules | | Change a rule | Someone asks | editing-the-playbook | ## Facts The only place these are written. Other docs point here. - Team: names, roles, ids - Website pages we link to - Booking link ## Change notes (ten lines at most) - 6 Oct 2026: rewritten; history moved to the archive.
A detail doc is shorter still: a read-when line, rules under the headings where an agent would look for them, and its change notes.
# Reply playbook Read when: a reply to one of our emails arrives. ## Categories - Not now: log it on the deal, create a follow-up task for the date they gave, no reply unless they asked. ## Objections - Answer the objection in two sentences, then one question. Example: "..." Not: "..." ## Change notes (ten lines at most) - 6 Oct 2026: moved from the old proposals doc.
For scale, this is our Playbook after the rewrite. Two docs load on every run; the rest load by job.
| Doc | What it holds | Size |
|---|---|---|
| Instructions (entry) | Hard rules, jobs table, Facts | 10 KB |
| Personal rules | One person’s own rules; they win for that person | 6 KB |
| Pipeline | Logging email, deal logs, stages, follow-ups, drafts, tasks, the morning check | 11 KB |
| Writing rules | How an email is written | 8 KB |
| Reply playbook | Reply categories, objections, the hourly run | 10 KB |
| Calls and meetings | The call task and the preparation in its notes | 5 KB |
| Client documents | Proposals and decks: shape, style, files | 13 KB |
| Pricing rules | How a price is set, deal value | 16 KB |
| Prospecting rules | Checking and adding companies | 10 KB |
| Segments and signals | Segment rules, status, the lists of firms to contact | 22 KB |
| LinkedIn outreach | Invitations and messages | 3 KB |
| Reporting rules | The pipeline report | 16 KB |
| Drive instructions | Our shared file folders | 11 KB |
| Editing the Playbook | How rules are written, changed and cleaned up | 6 KB |
| Archive | History moved out of the rules; never read for work | 158 KB |
Sizes rounded to the nearest KB.
Sources
- Anthropic: Effective context engineering for AI agents, 29 September 2025
- Chroma: Context Rot, how increasing input tokens impacts LLM performance, 14 July 2025
- Claude Code documentation: How Claude remembers your project (CLAUDE.md and auto memory)
- AGENTS.md: a simple, open format for guiding coding agents
- Model Context Protocol: What is MCP?
Read on 6 October 2026. The Playbook figures are Odysi’s own, from the rewrite on the same day.