Spec-driven development with AI agents: the numbers from one real project

AI coding agents write code fast. The hard part is knowing that what they built is what you meant, and that it keeps working when someone does something you did not expect.

So I tried an experiment: run a small but real project end to end with AI agents doing almost all of the work, under a process strict enough that I could trust the result without reading every line. I packaged that process as a Claude Code plugin, sdd-kit, and used it to build whatif-mcp, an MCP server that answers “what if” questions about one person’s finances from a model kept in a local file.

This post covers the first milestone: one spec, ten tasks, one QA pass, and the retro that changed the process.

The process in one picture

setup -> epic -> spec -> plan -> build (per task) -> QA -> retro
  • Humans decide, agents do. Every stage that needs a decision runs in my own session, so it can ask me questions and wait. Work that needs no decision (implementing, reviewing, adversarial testing) is handed to agents.
  • Grilling before writing. A spec is written only after a structured interview: questions in rounds, each with a recommended answer, until nothing is open. Then the spec is written from the decisions, with every expected number worked out by hand.
  • Vertical slices. The plan cuts the spec into tasks that each go through every layer they need to deliver one behavior a user can see, small enough for one agent session.
  • Test boundaries agreed up front. For each acceptance criterion, the plan says where the test drives the system (a core function, the file store, the tool layer through a client) and what that boundary cannot see.
  • Every task goes through the same gate. An implementer agent builds the task test-first. A script runs lint, type checks, unit and integration tests, and checks every acceptance criterion has a test. Then two reviewer agents read the diff in parallel: one against the spec, one against the coding standards. I decide: merge, rework, or waive a finding with a reason.
  • Retro changes the kit. Problems found along the way go in a log. The retro decides each one, and the result is a new, versioned release of the process.

My job in all of this was the human gates: approve the spec, approve the plan, decide each merge. Most decisions took one line.

The numbers

Spec 0001 (the model and the projection math) ended with 34 acceptance criteria and was built in 10 tasks over three days. 148 tests at the end.

MeasureResult
Tasks10
Merged without rework (first pass)1 of 10
Rework rounds10
Findings raised by the reviewer agents176
Findings I raised myself0
Bugs found after the tasks merged (by QA)2
Median time per task, branch to mergeabout 0.4 hours

Nine of ten tasks went back for at least one more round. The reviewers were strict, and most of what they raised was real: a value that could crash a tool, a test that only matched a substring, a rule that lived in three places. The rework rounds were short, usually minutes, but nearly every task had one. My target was above 60% first pass. That is a long way off.

Zero findings from me. Two reviewer agents and a gate script stood between the agent and me, so by the time a task reached me, reading the diff would not have found anything. My time went into decisions, not code review.

What broke

After the last planned task merged, with every gate green, I ran an adversarial QA agent over the finished spec and told it where to dig: hand-edited model files, edge cases in the math, and file safety.

It found ten problems. Two were bugs against the spec:

  • A failed call could still change the file. A large enough balance crashed the step that rounds money for the reply. But the write tools saved the file before that step. So a call reported failure, the file had changed anyway, and from then on every tool failed. The spec said invalid input leaves the file unchanged. The bug got through seven tasks, each with a green gate and a finished review.
  • Some unreadable files gave a generic “Internal server error” instead of an error naming the file and the problem.

The other eight were gaps in the spec: things it never said. What happens to a symlinked model file? To unknown keys in a hand-edited file? To a path like ~/fin/model.json? Each became a spec amendment I approved, and three small fix tasks.

While reviewing those fix tasks, the reviewer agents kept finding more crashes of the same kind: an inflation rate just above -1 that divided by zero, a ~someone/ path that raised an error the server did not handle. They found four of this class in a row, each only after the previous fix.

The retro added a check: one test that throws 600 random and hostile calls at every tool with a fixed seed, and fails if any call comes back as an internal error instead of a proper tool error. A one-off run over 300 seeds, 180,000 calls, came back clean. If that test had existed from the first task, all four would have been caught the day they were written.

The AI’s own mistakes

The orchestrating session, the AI I was talking to, also made mistakes: it wrote two of the spec amendments, and both were wrong:

  • An expected number in an acceptance criterion was off by a factor of a million. The implementer agent caught it while writing the test.
  • A rule I approved said a flow’s start date could be at most 100 years before today. That seemed reasonable until a reviewer pointed out that “today” moves. A model file valid today would be refused tomorrow, by every tool, including the one you would use to fix it. We replaced it with a fixed floor, 1900-01-01.

The process caught both and now have kit rules: amendments get the same hand-checked arithmetic as new criteria, and no rule on stored data may depend on today’s date.

What the retro changed

The retro turned every logged problem and every finding above into a decision, and the result is sdd-kit 0.3.0:

  • Plain terms. I had been saying “seams” and “slices”. They are now test boundaries and vertical slices, each defined where it is first used.
  • An optional integration branch, so tasks can merge into a develop or per-spec branch and only whole, gated specs reach main.
  • Token and waiting-time metrics. Cycle time was hiding how long tasks sat waiting for me, and nothing measured what the agents cost.
  • A test that passes on its first run is not trusted until someone breaks the code on purpose and watches it fail. The agents had started doing this on their own in most tasks; now it is the rule.
  • Smaller fixes: write the gate output before the reviewers run, use a separate worktree when two tasks build in parallel.

This was the second retro. The first one, after a single task, produced 20 changes, including the biggest one: stages that need a human became skills in my own session, because agents running on their own could not ask me anything and questions came back in slow rounds through a relay.

What I take from it

  • The gate only checks what the spec says. Everything the reviewers checked, they checked well. Eight of the ten QA findings were behavior the spec had never stated, and the two bugs sat between steps that no single task owned. A QA pass that asks “what did nobody write down?” is not optional.
  • One generated test closed a class of crash. The same class slipped past careful reviewers four times.
  • My work was decisions. I approved specs and plans, ruled on findings and waived some with a reason. I did not read diffs line by line; the QA pass caught what the gate missed.
  • Nine of ten tasks needed rework. The rounds were short. Whether that rate improves is the question for part 2.

Caveats

This is one project and one spec, by one person who also wrote the process. The project is small and has no network, no database and no users yet. Token cost was not measured this time; 0.3.0 adds it. Treat this as a field report from one project.

Try it

sdd-kit is open source (MIT) at github.com/timpinard/sdd-kit, and the project built with it, whatif-mcp, is at github.com/timpinard/whatif-mcp. Version 0.x and experimental: expect the process to change, with a changelog and migration notes for each release. The grilling and test-driven development skills are adapted from Matt Pocock’s skills collection, used under the MIT license.

Part 2 will cover the next spec on 0.3.0: whether the retro changes
worked, and the first token numbers.


Comments

Leave a Reply