Testing surgery

7,705 fewer lines of tests, same coverage, one prompt?

Published on

A few years ago I worked for a company that prided itself on its extreme programming credentials.

Pair programming was the norm, every commit was deployed to production within minutes (after going through an automated CI/CD process) and extensive logging meant anything that fell over in production was jumped on by someone on the dev team.

But at one point it became clear too many defects were making it into production code.

So the order came to increase test coverage across several of our main applications.

While the intentions were sound, the reality was this achieved very little.

The entire team spent a week increasing test coverage.

But there’s real danger in chasing a single metric like this.

Lots of tests went in, and coverage increased.

But the tests were often tightly coupled to the implementation code.

Mocks being set up and configured in the tests, meaning you had to know and capture what the internals of the code under test was doing, making the tests brittle and prone to failure.

Fig. 1 · Coverage as a target
Coverage
81%
Then, a refactor
Split OrderService in two. What it does stays the same.
✗ 5 of 7 tests fail
OrderService
PricingService
Every mock pins down how the code works inside. Coverage climbs, but split the class in two without changing what it does, and most of the tests fail anyway.

It basically meant you couldn’t refactor the code under test without also changing the tests.

So adding more tests, whilst coverage did go up, put an even larger straightjacket around the development team and their ability to modify and extend the code.

Increase or maintain coverage, and remove the straightjackets

Fast-forward a decade, and I was trying to do something similar, but on a project I’m building using Claude Code.

In this case, I’m actually reasonably happy with the test coverage.

But less happy with the number of lines of test code.

I let AI go to town with adding its own tests for this application and, gone to town it has…

Before my experiment, I was staring at 33,000 lines of test code, and about 75% coverage.

Fig. 2 · The test code in play
Lines of test code
33,692
Files
159
Coverage
about 75%
Each square is one test file. The bigger the square, the more lines of test code in it.

More lines of code means more to maintain, more possible straightjackets around production code, and more likelihood of the LLM changing tests at the same time as production code (which means we’ve too many moving parts and our confidence in whether we’re preserving behaviour starts to plummet).

So I decided to try something in Claude Code, after seeing Kent Dodds try the same.

Fig. 3 · The whole prompt

GoalConstraintScope

Notice this doesn’t specify how to do testing surgery, just the goal and constraints.

First it told me I was wrong

Before deleting a single line it measured.

The whole test suite was around 550,000 lines, with about 33,700 covering this specific part of the application.

So removing “hundreds of thousands of lines” was impossible, and Claude said as much before it started the work.

Fig. 4 · Drawn to scale
Whole test suite · ~550,000 lines
This feature · ~33,700 lines
"Remove hundreds of thousands of lines"

Then it captured a baseline

Claude switched on a coverage tool and recorded exactly which lines and branches of the production code were covered by tests.

(These numbers measure something different from the 33,692 above. That was lines of test code. Coverage counts lines of production code: this feature has 10,101 of them, and the tests ran 7,740, or 76.6%.)

It ran this twice to check for variability between runs.

Fig. 5 · The safety net
Run 1 · production lines run by tests
7,740 / 10,101 lines
3,058 / 4,921 branches
Run 2
Identical
no drift between runs
THE RULECoverage is the floor, not the target.
Each cell stands for about five lines of production code, not test code. Green: run by at least one test. Any cut that turns a green cell grey fails the check.

After that it followed a rule: whatever gets deleted, every one of those lines covered by tests must still be covered afterwards.

Which is interesting in retrospect, because it’s the opposite of what the team did all those years ago, where they treated coverage as a target.

Here Claude set the coverage level as the floor. It could change anything else, but that figure (and the exact lines covered) had to remain constant.

Split the work, like a team lead

After that it divided the code into areas, and spun up agents to work on them in parallel, each in its own isolated copy (worktree) of the repo.

Fig. 6 · Four agents, four worktrees

Each agent owned its own files, so none could step on another.

Three of the four tripped the coverage check part-way through, and had to put the coverage back before they could finish.

The biggest: the authoring agent deleted a test whose assertion could never fail, and 43 lines went uncovered. It turned out to be the only test that rendered a signed engagement letter, so it went back in.

All the agents worked to this brief:

I noticed one go-to technique the agents turned to was to take similar looking tests and use a table to exercise one version of the test code with multiple scenarios (replacing lots of similar looking tests that existed before).

Fig. 7 · UsernameValidatorTest.java, a made-up example
@Test
void rejectsAnEmptyUsername() {
    var result = validator.check("");

    assertThat(result.error()).isEqualTo("Username is required");
}

@Test
void rejectsAUsernameThatIsTooShort() {
    var result = validator.check("jo");

    assertThat(result.error()).isEqualTo("Username must be at least 3 characters");
}

@Test
void rejectsAUsernameWithSpaces() {
    var result = validator.check("jon hilton");

    assertThat(result.error()).isEqualTo("Username cannot contain spaces");
}

@Test
void rejectsAUsernameWithSymbols() {
    var result = validator.check("jon!");

    assertThat(result.error()).isEqualTo("Username can only use letters and numbers");
}
@ParameterizedTest
@CsvSource({
    "'',          Username is required",
    "jo,          Username must be at least 3 characters",
    "jon hilton,  Username cannot contain spaces",
    "jon!,        Username can only use letters and numbers",
})
void rejectsAnInvalidUsername(String username, String error) {
    var result = validator.check(username);

    assertThat(result.error()).isEqualTo(error);
}
Four tests that differ only in the username and the expected error become one test, with each case as a row in a table.

To verify the changes were not lowering coverage, for every test an agent deleted it was required to note which surviving test now covered that behaviour.

The results

After about 45 minutes the results were in.

Fig. 8 · After surgery
Test code
25,987
7,705 fewer · about 23%
Coverage
76.6% → 76.6%
Lines or branches lost
0
Production code changed
0 lines
The same 159 files as Fig. 2. Amber squares were trimmed; six files were folded into others and disappear.

Not hundreds of thousands, but nearly a quarter of the test code gone and nothing lost.

So, what to do?

A lot of the reduction here came from tidying up (tables, shared setup) rather than reducing the tight coupling that made my old team’s tests so painful.

The straightjacket is lighter, but it’s still there.

So that naturally leads to the next experiment, to see if I can raise the tests up to the highest ‘seam’ without losing coverage.

Fig. 9 · The next experiment
HTTP route
Service
Class
← tests
Today most of the tests call classes directly. The next experiment is to move them up to the HTTP route, without losing coverage.

That’s for another day.

But overall, this one prompt to Opus 5.5 did a better job of optimising tests than an entire dev team did in one week a decade ago.

It established metrics, implemented a safety net, and found a way to prove it hadn’t broken anything along the way.

Go forth and experiment

If you’re struggling to ‘keep up, and get solid results from AI (and feeling like you’re losing your craft), experiments like these are a solid tactic.

Find something you want to improve, spin up Opus 5.5, Fable, Astra, and see how it handles the task.

And resist the temptation to over prescribe how the LLM should achieve its goal.

As the models get better, they need less steering, and you can always course correct once you see what they’re up to.

— Jon

P.S. we cover experiments like this, and lots more over in AI Coding Lab - come check us out over there.

As always, this article was written entirely by my own hand. AI used for some very light editing.

Your skills aren't becoming obsolete. They're becoming essential.

Practical engineering principles for building software that works - with or without AI.

    Join 6,000+ developers. Hype-free.