Testing surgery
7,705 fewer lines of tests, same coverage, one prompt?
A few years ago I worked for a company that prided itself on its extreme programming credentials.
Pair programming was the norm, every commit was deployed to production within minutes (after going through an automated CI/CD process) and extensive logging meant anything that fell over in production was jumped on by someone on the dev team.
But at one point it became clear too many defects were making it into production code.
So the order came to increase test coverage across several of our main applications.
While the intentions were sound, the reality was this achieved very little.
The entire team spent a week increasing test coverage.
But there’s real danger in chasing a single metric like this.
Lots of tests went in, and coverage increased.
But the tests were often tightly coupled to the implementation code.
Mocks being set up and configured in the tests, meaning you had to know and capture what the internals of the code under test was doing, making the tests brittle and prone to failure.
It basically meant you couldn’t refactor the code under test without also changing the tests.
So adding more tests, whilst coverage did go up, put an even larger straightjacket around the development team and their ability to modify and extend the code.
Increase or maintain coverage, and remove the straightjackets#
Fast-forward a decade, and I was trying to do something similar, but on a project I’m building using Claude Code.
In this case, I’m actually reasonably happy with the test coverage.
But less happy with the number of lines of test code.
I let AI go to town with adding its own tests for this application and, gone to town it has…
Before my experiment, I was staring at 33,000 lines of test code, and about 75% coverage.
More lines of code means more to maintain, more possible straightjackets around production code, and more likelihood of the LLM changing tests at the same time as production code (which means we’ve too many moving parts and our confidence in whether we’re preserving behaviour starts to plummet).
So I decided to try something in Claude Code, after seeing Kent Dodds try the same.
Notice this doesn’t specify how to do testing surgery, just the goal and constraints.
First it told me I was wrong#
Before deleting a single line it measured.
The whole test suite was around 550,000 lines, with about 33,700 covering this specific part of the application.
So removing “hundreds of thousands of lines” was impossible, and Claude said as much before it started the work.
Then it captured a baseline#
Claude switched on a coverage tool and recorded exactly which lines and branches of the production code were covered by tests.
(These numbers measure something different from the 33,692 above. That was lines of test code. Coverage counts lines of production code: this feature has 10,101 of them, and the tests ran 7,740, or 76.6%.)
It ran this twice to check for variability between runs.
After that it followed a rule: whatever gets deleted, every one of those lines covered by tests must still be covered afterwards.
Which is interesting in retrospect, because it’s the opposite of what the team did all those years ago, where they treated coverage as a target.
Here Claude set the coverage level as the floor. It could change anything else, but that figure (and the exact lines covered) had to remain constant.
Split the work, like a team lead#
After that it divided the code into areas, and spun up agents to work on them in parallel, each in its own isolated copy (worktree) of the repo.
Each agent owned its own files, so none could step on another.
Three of the four tripped the coverage check part-way through, and had to put the coverage back before they could finish.
The biggest: the authoring agent deleted a test whose assertion could never fail, and 43 lines went uncovered. It turned out to be the only test that rendered a signed engagement letter, so it went back in.
All the agents worked to this brief:
- Collapse near-identical tests into a single table of cases
- Pull repeated setup into shared helpers
- Delete tests that only checked whether a mock returned what it was told to return
I noticed one go-to technique the agents turned to was to take similar looking tests and use a table to exercise one version of the test code with multiple scenarios (replacing lots of similar looking tests that existed before).
@Test void rejectsAnEmptyUsername() { var result = validator.check(""); assertThat(result.error()).isEqualTo("Username is required"); } @Test void rejectsAUsernameThatIsTooShort() { var result = validator.check("jo"); assertThat(result.error()).isEqualTo("Username must be at least 3 characters"); } @Test void rejectsAUsernameWithSpaces() { var result = validator.check("jon hilton"); assertThat(result.error()).isEqualTo("Username cannot contain spaces"); } @Test void rejectsAUsernameWithSymbols() { var result = validator.check("jon!"); assertThat(result.error()).isEqualTo("Username can only use letters and numbers"); }
@ParameterizedTest @CsvSource({ "'', Username is required", "jo, Username must be at least 3 characters", "jon hilton, Username cannot contain spaces", "jon!, Username can only use letters and numbers", }) void rejectsAnInvalidUsername(String username, String error) { var result = validator.check(username); assertThat(result.error()).isEqualTo(error); }
To verify the changes were not lowering coverage, for every test an agent deleted it was required to note which surviving test now covered that behaviour.
The results#
After about 45 minutes the results were in.
- Test code: 33,692 lines down to 25,987 (7,705 fewer, about 23%)
- Coverage: 76.6% before, 76.6% after
- Lines or branches of coverage lost: zero
- Production code changed: zero
Not hundreds of thousands, but nearly a quarter of the test code gone and nothing lost.
So, what to do?#
A lot of the reduction here came from tidying up (tables, shared setup) rather than reducing the tight coupling that made my old team’s tests so painful.
The straightjacket is lighter, but it’s still there.
So that naturally leads to the next experiment, to see if I can raise the tests up to the highest ‘seam’ without losing coverage.
That’s for another day.
But overall, this one prompt to Opus 5.5 did a better job of optimising tests than an entire dev team did in one week a decade ago.
It established metrics, implemented a safety net, and found a way to prove it hadn’t broken anything along the way.
Go forth and experiment#
If you’re struggling to ‘keep up, and get solid results from AI (and feeling like you’re losing your craft), experiments like these are a solid tactic.
Find something you want to improve, spin up Opus 5.5, Fable, Astra, and see how it handles the task.
And resist the temptation to over prescribe how the LLM should achieve its goal.
As the models get better, they need less steering, and you can always course correct once you see what they’re up to.
— Jon
P.S. we cover experiments like this, and lots more over in AI Coding Lab - come check us out over there.
As always, this article was written entirely by my own hand. AI used for some very light editing.
Your skills aren't becoming obsolete. They're becoming essential.
Practical engineering principles for building software that works - with or without AI.