I kept making an AI agent more capable. More tools. More instructions. More memory. Every addition seemed sensible on its own.
Then I noticed something I hadn’t expected: the agent could do more, but I trusted it less.
Simple tasks sometimes took a detour through the wrong tool. A clear instruction became harder to follow when it sat beside twenty others. Context I had added to help seemed to make a decision harder.
My instinct as a software engineer was to patch the missing piece. Add a guardrail. Add an example. Explain the tool one more time.
That helped, until it didn’t. Eventually, I had to admit I was trying to solve complexity by adding more complexity.
So I changed the question. Instead of “What else does this agent need?”, I started asking, “What can I remove?”
A bigger toolbox isn’t the same as a better decision
Imagine an agent with one job and four tools. It still has to reason, but its choices are fairly clear. Now give it sixteen tools. Each might have a legitimate use somewhere in the system. That doesn’t mean each belongs in this particular task.
In a traditional application, adding a helper function doesn’t make every other function reconsider whether to use it. An agent works differently. Tool definitions, instructions, retrieved documents, and conversation history all shape its next decision.
When I add a capability, I also change the decision the model has to make. I give it another distinction to understand and another plausible way to go wrong.
That doesn’t mean fewer tools are always better. It means capability and reliability deserve separate attention.
When the prompt becomes a record of every past mistake
I used to treat the system prompt like a patch file. Wrong tool call? Add a sentence. Missed exception? Add a section. An unusual workflow? Explain it in more detail.
The prompt grew one failure at a time. Each addition addressed a real problem, which made the growth feel justified.
Eventually, though, the prompt read less like a set of operating principles and more like the history of every bug I had encountered. The most important instruction was still there. It was simply surrounded by everything else.
At that point, the problem wasn’t missing information. It was finding the signal inside it.
Context is working memory, not a warehouse
A large context window makes it tempting to include everything that might help: policies, examples, old decisions, user preferences, tool schemas, and intermediate results.
But “this could matter someday” is a very different standard from “this is needed for the next decision.”
My computer has plenty of storage. I don’t load the entire disk into memory before running a function. I’ve come to think about agent context in a similar way: a place for useful working information, with the rest available to retrieve when needed.
The practical rule I keep returning to is simple:
Give the model what it needs for the current decision, not everything the system knows.
For me, that means leaving irrelevant tools out of the current stage, retrieving information when it becomes useful, and letting completed work leave the active context.
Every tool is also an interface
Consider two tools named search_customer and find_customer_information. A developer can inspect the implementation to learn how they differ. The model has to choose from the interface we give it.
If the distinction is vague, we’ve created ambiguity and asked the model to resolve it. Before blaming the model, I need to ask whether the interface is clear enough.
Tool names, boundaries, descriptions, and return shapes are all part of that interface. A tool has two users: the software that executes it and the model deciding whether to call it.
Sometimes a clearer description helps. Sometimes the better fix is to expose fewer overlapping choices at that stage.
Good memory can make room to forget
“Remember more” sounds like an obvious goal for an agent. I’m increasingly interested in a different one: help it forget safely.
A long-running task may not need its full transcript at every step. It may need a compact record of what was decided, what evidence supported the decision, where that evidence lives, and when to retrieve it again.
The agent can leave breadcrumbs rather than carry the whole trail.
That approach needs care. A summary can lose a detail that matters later. Retrieval can miss the right record. But carrying everything has costs too, and those costs become harder to ignore as a workflow grows.
Simplicity helps the person debugging it
When an agent has dozens of tools, several memory paths, and a sprawling prompt, an odd result can have many explanations. Was the instruction misunderstood? Did retrieval return the wrong document? Did an old note conflict with the current state? Were two tools too similar?
Reducing complexity doesn’t make a language model deterministic. It does make the system easier for me to inspect. There are fewer places for a failure to hide.
The person operating the agent needs to understand it, too.
The question I want to test next
These are engineering observations, not benchmark results. I haven’t yet measured how much tool count, context size, or retained history changes reliability. That distinction matters.
My next step is a controlled experiment using synthetic data. I want to compare the same tasks with four, eight, and sixteen exposed tools, then compare curated context with the same useful facts surrounded by realistic distractors.
I’ll measure task success, first-tool selection, unnecessary calls, latency, and token usage. The tasks, tool implementations, model version, and evaluation rules need to stay consistent across conditions, with repeated runs to account for variation.
The hypothesis is that unnecessary choices and information can make decisions harder. The experiment could support it, complicate it, or show that the effect depends heavily on the model and task. Until it runs, I won’t attach numbers to the argument.
Complexity should earn its place
A demo asks whether an agent can do something. Operating it asks how often it does the right thing without someone steering every step.
Both questions matter. The second has changed how I build.
Tools, instructions, context, and memory can all be useful. I want each addition to justify the extra decision it creates and to be evaluated against the job the agent actually needs to do.
The next time an agent behaves badly, I’ll still check what it’s missing. But I’ll also ask what it doesn’t need.
Sometimes the most useful thing I can give an agent is less to think about.
Keep the conversation going.
Working through similar questions about AI agents or system design? I’d like to hear your perspective.
Write to me