Why an AI Agent Loses Context in a Monorepo
The agent didn't get dumber because the repo got bigger. It got dumber because I kept handing it the whole thing.
AIThe agent didn't get dumber because the repo got bigger. It got dumber because I kept handing it the whole thing.
For a while my instinct with any coding agent was the same one I'd use with a new hire: give them everything, let them figure out what matters. That works for a person on their first week. It doesn't work the same way for a model with a context window, because a person filters as they read. The model doesn't filter first and read second. Reading is filtering, or it's supposed to be, and every file you hand it that has nothing to do with the task is something it now has to reason around instead of reason with.
I've watched this pattern often enough to trust it: point an agent at a large monorepo with "here's the whole thing, go fix the bug," and somewhere between the tenth and thirtieth file it starts drifting. Not failing loudly. Drifting. It starts referencing something from a module three folders away that was never relevant, often something old, code that hasn't actually been used in months but is still sitting there looking current, and it treats it as if it matters. Or it "forgets" a constraint you mentioned at the start because that instruction is now buried under a wall of code it read on the way to finding the actual problem. The context window didn't run out. It filled up with the wrong things.
Start with what actually changed
The most obvious fix is also the correct first step: hand the agent the diff, not the repo. If the task is "fix this bug," the bug lives in specific files that recently changed, or files that are directly implicated by the bug report. A git diff, or a scoped set of recently touched files, tells the agent exactly where the ground moved.
This alone fixes a surprising amount. But it's not enough on its own, and here's where I used to stop too early. A diff tells you what changed. It doesn't tell you what that change touches. A three-line edit to a shared utility function can be invisible in the diff and load-bearing for a dozen other files that import it.
Then you need to know what's connected
This is where a dependency graph earns its cost. Not a full architectural diagram nobody keeps updated, something closer to: what does this changed file import, what imports this changed file, and how far out does that ripple before it stops mattering for the task at hand. That's the actual boundary of "relevant context" for a given change, and it's almost never the same shape as "everything in this directory" or "everything with a similar filename."
The honest version of this: a good dependency graph doesn't just add the right files, it gives you permission to leave files out. That's the harder discipline. It's easy to over-include "just in case." It's much harder to trust that a file genuinely doesn't matter for this task and leave it out of the context on purpose.
And you shouldn't rebuild that graph every time
Once you're pulling a dependency graph for every task, the next problem shows up fast: recomputing it from scratch on every single request is slow and, more to the point, expensive. Token cost scales with how much you make the model read, and rebuilding the same relationship map for a repo that hasn't meaningfully changed since the last request is paying twice for the same information.
A cached map of the repository fixes this, as long as the invalidation is honest. Not on a timer, timers go stale exactly when a change just landed and accuracy matters most, but keyed to whether the relevant part of the tree actually moved. It's the same instinct as any other cache: expensive to compute, cheap to reuse, dangerous to trust once it's stale.
None of this replaces architecture
I want to be honest about the limit here, because it's real. Minimal context tooling assumes there's a coherent boundary to find in the first place. If the codebase itself doesn't have clear module boundaries, if anything can import anything the way I wrote about with Nx tags, then the dependency graph the agent builds is just an accurate map of an incoherent territory. Good tooling can't invent boundaries that were never enforced. It can only respect the ones that exist.
What this actually buys you
Minimal context isn't an optimization you bolt on once everything else works. It's closer to a precondition for the agent being able to reason at all instead of just producing plausible-looking text from whatever happened to be in front of it. Diff first, graph second, cache to keep from paying for the same answer twice. None of it is exotic. All of it is the difference between an agent that's actually working the problem and one that's confidently guessing.
I'm building this discipline into an open source tool right now, Intentloom, still early, still rough around the edges, and I'm going deeper into the token economics of it specifically in the next piece. Worth subscribing if this kind of thing is useful to you.
The agent didn't get dumber because the repo got bigger. It got dumber because I kept handing it the whole thing. Diff first, dependency graph second, cache to avoid paying for the same answer twice.
