- ai
- agents
- context engineering
How I engineer context for coding agents
My current approach, after building a Rust game almost entirely through Claude Code: keep the always-on context small, move rules into the compiler, and give each agent only what its role needs.
For a while now I’ve been building a game in Rust almost entirely through Claude Code. I set the direction and make the design calls; agents write, review and merge most of the code. At work I do a smaller version of the same thing with Kiro on a React codebase.
The biggest lesson so far: the prompt is the weakest part of the system. What makes agents reliable is everything around the prompt. What they read and when, what they physically cannot do, and who checks their work. People call this context engineering now. This post is my current approach, which will keep changing.
1. Keep the always-on context small
Every token that loads on every task competes with the task itself. My project instructions file holds only what every task needs: how we work, a handful of architecture laws, and how to build and test. It’s under a hundred lines, and its first paragraph states that rule, so neither I nor an agent grows it by accident.
Everything else loads on demand:
- Path-scoped rules. The detailed Rust architecture and testing rules load only when an agent reads a
.rsfile. The docs rules load only for docs. - Skills with triggers. Running a story, debugging, driving the game for a screenshot: each is a skill that loads when its trigger fits.
- A reference skill for the engine. The engine moves fast and model training data lags behind it, so there is a skill with the intended way to use each API in the version I’m on. It’s the one place where I deliberately load a lot of text, because it replaces guessing.
---
paths:
- "src/**/*.rs"
---
# Rust architecture and testing rules2. One definition per thing
When the same fact lives in two places, one of them goes stale and the agent will find the stale one. So every command lives in one script, and docs never contain raw build commands, only the script’s verbs. If the build changes, one file changes.
At work the equivalent is a single frontend design document that agents treat as the source of truth for code style and architecture. The Kiro steering files build on it.
3. Move rules out of the context and into mechanisms
This is the one I’d tell every team. I rank enforcement like this, strongest first:
- The compiler or linter refuses it.
- A test refuses it.
- A hook refuses the tool call before it runs.
- A sentence in a prompt asks the agent not to do it.
The last one has failed me repeatedly, so a rule worth enforcing moves up the list. In the game, Clippy runs in pedantic mode with a deny list, and banned methods point the agent to the project’s own way of doing things:
disallowed-methods = [
{ path = "std::time::Instant::now", reason = "read time from the injected clock so tests can control it" },
{ path = "rand::rng", reason = "draw randomness from the seeded generator so tests can seed the dice" },
]The detail that matters is the reason. When the agent hits the error, the error message is the context it needed, delivered at the exact moment it needs it. The same goes for my hooks: when one refuses a command, its message names the command to use instead. A refusal that teaches is worth more than a paragraph in the instructions that the agent may or may not weigh.
The hooks themselves are code, so they get tested like code, including mutation tests. A commit that touches a guard is refused while any mutant survives. Otherwise you end up with a guard that quietly stopped guarding.
4. Give each agent only the context of its role
A story in my pipeline goes through several agents, and each gets a different slice of context on purpose:
- Second opinion. Reads the plan cold while I’m still shaping it, and says where it would design differently.
- Implementer. Gets one task, turns it into one commit, and has to show its tests can fail.
- Reviewer. Read-only. Checks each commit against the plan’s binding decisions. It stays alive across the whole story so it can spot drift between tasks.
- Final reviewer. Starts fresh and reads the whole branch once.
The fresh starts matter. A reviewer that watched the implementation being reasoned out tends to agree with that reasoning. A reviewer that only sees the result and the plan judges the result.
5. Put state where agents already look
Agents lose track of separate task lists and status files. They don’t lose track of git. So each commit carries trailers: which task it implements, where it deviates from the plan and why, and when it passed review. Progress is derived from those trailers, which means it can never drift from the actual code.
The plan document has the same split. Its frame binds, and its tasks only guide. When the code calls for a different approach than a task describes, the implementer changes course and records a deviation, instead of either following a bad plan or silently ignoring it.
6. Feed reality back in
An agent working blind on a game will happily claim a visual change works. So agents drive the running game through a small file-based command channel: set up a scene, query state, take a screenshot from the engine itself. I prefer state readouts over pixels, because a number is easier to check than an image.
When a change really does need my eyes, the pipeline builds the before and after and sends me a page I can vote on from my phone. My attention goes to the decisions that need it, and the agent doesn’t stop and wait for me in the meantime: it makes the call, ships it, and files a ticket saying what changes if I overrule it.
7. Treat lessons as code
The tempting fix for a recurring mistake is another line in the agent’s memory. That’s the weakest option again, just in a different file. My agents can’t write to their memory directly; the only route asks first why the lesson isn’t a hook, a lint rule or a ticket.
I also hold the process itself to a budget. A new review step or check has to name the failure it fixes, how often that failure happens, and what the step costs per story in agent minutes. Every story ends with a cost report from the agent transcripts, so I can see whether a mechanism pays for itself or should be deleted.
Where I’d start in your codebase
You don’t need a game or a five-agent pipeline to use any of this. If I joined your team tomorrow, I’d start here:
- Write one design document that agents and developers both treat as the source of truth, and keep the always-on instructions short.
- Take the three rules your agents break most and turn them into lint rules with good error messages.
- Add a hook that blocks destructive or expensive commands before they run.
- Never let the agent that wrote the code be the only one that reviews it.
- Measure what each piece of agent work costs, so you know what to keep.
I’ll write more about each of these. If you’re working on the same problems, I’d like to hear how you approach them.