PRODUCTS

KEYWORDS

Agentic Engineering Failure Modes

Before I hop into the article, I would like to state my agentic engineering credentials. I’m not your average small company CEO. Or maybe I am? I picked up a vibe coding project when you could still code on vibes. I spent a week in Gas Town building a DoltLite prototype. DoltLite is a SQLite-compatible version of Dolt. I soon ditched Gas Town and the term vibe coding for a more traditional agentic engineering setup. A couple thousand agent-engineered pull requests (PRs) later, DoltLite is Beta and a core part of DoltHub’s product offerings.

I was writing another blog article that wasn’t fully baked but had a great section entitled “Agentic Engineering Failure Modes”. I thought to myself, “That article can come out of the oven immediately. Failure modes with funny names. Just ship that.” Here you go.

The Commentator#

I’ll start with failure modes I’ve written about before. The first is that agents write too many code comments. Coding agents love to write code comments even though evidence shows comments make agents worse. Anecdotally, I can confirm coding agents work worse before I cull code comments.

Agents seem tuned to write code for humans, but in write-only code projects, a human is not the main audience for the code — another agent is. So establish a process to periodically remove and shorten code comments.

The Commentator

The Historian#

This is somewhat related to The Commentator. Not only do agents love to write comments, they also love to write documentation. Beware of in-repository documentation. Agents will see it and start documenting every small change to the codebase. Documentation will grow into an unusable, uneditable mess.

The Historian

The CI Murderer#

My argument that GitHub is worth $1T, just like the labs, rested on the idea that agents need CI to function correctly. Agents need tests. $300/month in coding agent subscriptions generates $3,000/month in GitHub Actions.

Making too many changes too quickly exhausts testing resources, even if you are willing to pay more. At DoltHub, one of our employees had an agent generate testing PRs to confirm approximately fifty old Dolt issues were resolved. This occupied our GitHub Actions runners for 20 hours. Good luck getting a change in. Agents don’t have a good sense of how to bundle PRs to save CI resources. Some agents, I’m looking at you Grok, also don’t test enough locally, relying on CI to catch their errors.

The Murderer

The Shit Streak#

The Shit Streak is a special type of code bloat. Here’s the situation. Someone files an issue that may not describe a real problem. A coding agent is assigned to the fix. The coding agent assumes the issue is valid and sets out to fix it. The issue ends up being false. Now, a whole batch of changes that should never have happened in the first place is committed. Agents have a hard time deleting code so now the “fix” is stuck in your codebase. Agents didn’t know the issue they were solving was wrong or unimportant and “solved” it anyway.

This failure mode is especially prominent in “software factory”-type agentic setups because issues are often assigned to agents as soon as they are created. There is little screening of incoming work.

The Shit Streak

The Slop Spiral#

This is another failure mode that shows up more often in a “software factory”-style agentic setup where agents review each other’s work. The Slop Spiral shows up more often in issues or reviews than in code. One agent reacts to another agent’s comments as if the feedback is from a human. The agent treats the other agent’s comments as gospel and overreacts. The agent will create unnecessary new issues or follow-up commits. If an agent is reviewing those new issues or commits, the same thing can happen again for many cycles. Again, this is the result of a weak gate on new work. Humans still should be in the loop to judge importance to prevent The Slop Spiral and The Shit Streak.

The Slop Spiral

The YOLO#

We’ve written about the YOLO failure mode before in the context of agents removing Dolt lock files, calling it “a locally reasonable, globally terrible idea”. In the context of debugging an issue, an agent may take steps to resolve an issue that cause damage to a system. This is the classic “an agent deleted my database” failure mode.

This happens less with modern agents and harnesses because they are trained to be terrified of deleting anything — a reluctance that contributes to some of the failure modes listed above.

The YOLO

The Test “Fix”#

This last failure mode is one of the OGs. An agent made the tests match the new “correct” behavior instead of fixing the code. Again, I see this one less with more modern agents and harnesses, but it still happens. So it’s still included here for posterity.

The Test Fix

Conclusion#

As you can see, agentic engineering failure modes are evolving over time. Model and harness fixes are creating new failure modes. A human in the loop who recognizes these failure modes can intervene to prevent many of them. Stay vigilant and stop by our Discord if you notice any of these in DoltLite.