Writing8 min read
AI agents fail at the pipeline, not the model
Nearly every disappointing agent project I have taken apart had the same cause, and it was never the model. It was what the model was allowed to see.
- Build
- Automate

I have taken apart a fair number of disappointing AI agent projects now, ours and other people's. The failure is almost never the model. It is what the model was allowed to see, and what it was allowed to do about it.
That is good news, because pipelines are an engineering problem with known answers.
The four ways it actually fails
The index is stale. Somebody loaded the documents once, six months ago. Since then the pricing changed, three processes were rewritten and the templates were replaced. The agent is confidently describing a company that no longer exists. Every single time I have found this, it was because syncing was treated as a setup step rather than as infrastructure.
Retrieval returns the wrong chunk. The answer exists in the corpus, and the retriever hands the model a neighbouring paragraph instead. The model does what models do: it produces a fluent answer from what it was given. People call this hallucination. It is closer to being handed the wrong file.
No scope, no edge. The agent will answer anything, because nobody told it what it does not know. A useful agent has a boundary and a way to say so — "I don't have that, here's who does" is a correct answer and should be a designed one.
Nobody defined what right looks like. There is no test set, so there is no way to know whether a change improved anything. The team ships prompt tweaks based on vibes and slowly makes it worse.
What I build instead
Same four problems, inverted:
- Sync on a schedule, alert when a source goes quiet. A connector that silently stops is how an agent starts lying with confidence.
- Evaluate retrieval separately from generation. Before you judge the answers, check whether the right chunk was even in the context window. Most of the time it was not, and no amount of prompt engineering fixes that.
- Permissions enforced at retrieval. Not in the prompt. A model asked nicely not to reveal something is not an access control system.
- A test set of real questions with known answers. Sixty is plenty. Score it before launch, re-run it after every change to the pipeline. This is the single highest-leverage hour in the whole build.
Cite the source, always
Every answer my agents produce carries the document it came from. Two reasons, and the second one matters more.
The obvious reason is trust: a person can check. The better reason is that citations turn a vague complaint into a debuggable report. "It gave a wrong answer" is unactionable. "It cited the 2024 pricing doc" tells me the index has a stale file in it, and I can fix that in ten minutes.
The uncomfortable summary
The interesting part of agent work is not the interesting part. It is sync jobs, chunking strategy, permission checks and an evaluation harness — ordinary engineering, done carefully, around a model that was never the bottleneck.
Anyone selling you the model is selling you the cheapest component in the system.