Leading a technical debt clean-up without stopping the business
A multi-service platform had years of shortcuts weighing the team down. Here's how we prioritised the clean-up and actually saw it all the way through.
The problem
This is a large platform made up of several services that work together, built up over a good few years by different teams under different pressures. Nobody set out to make a mess. It’s what happens naturally when a business is growing fast and every deadline matters more than the tidiness of the code behind it. Shortcuts get taken with every intention of fixing them later, and later never quite arrives.
By the time I got involved, the cost of all that was showing up everywhere. Features that should have taken days were taking weeks, because half the effort went into working around old decisions rather than building the new thing. Releases carried more risk than they should have, because nobody was entirely sure what a change in one part of the system might break in another. And the engineers, who are generally the ones who feel this most acutely, were frustrated. Good people don’t enjoy fighting the codebase every day.
The business knew something needed to happen. What it didn’t have was a clear, honest picture of where the real problems were, and it certainly didn’t have the appetite to down tools for six months to fix everything at once. That’s rarely the right answer anyway. A business still has to ship things while it cleans up.
What we did
The first step was an honest audit, and I mean genuinely honest, not a box-ticking exercise designed to justify a big rewrite. I went through the platform service by service, talking to the engineers actually working in each part, looking at where bugs kept recurring, where changes took longest, and where people were visibly working around problems rather than through them.
That gave us a real list, and the next job was ruthless prioritisation. Not everything that’s messy is worth fixing. Some technical debt is genuinely costing the business time and money every week. Some of it is just untidy and can be left alone indefinitely without anyone suffering for it. I built a shortlist based on where the pain was actually landing, not on what was theoretically the “correct” way to build things.
Then came the part that matters most and gets skipped most often: actually leading the remediation. I didn’t hand over a report and leave. I stayed on to run the programme, working alongside the existing engineering team, sequencing the fixes so the riskiest, highest-value ones went first, and making sure the clean-up happened alongside normal feature work rather than instead of it. The business kept shipping to customers the whole way through, which was non-negotiable.
Part of the job was also cultural. Technical debt creeps back in unless the team has a shared, realistic standard for what’s acceptable to ship under pressure and what isn’t. So alongside the fixes, we put lighter habits in place, ones that don’t slow the team down but stop the same problems quietly rebuilding themselves.
Where things landed
The clearest sign of progress is what the team spends its time on now. Feature work that used to involve untangling old decisions first now mostly doesn’t. That’s not a minor convenience, it’s the difference between a roadmap that moves and one that stalls.
Releases carry less anxiety too. Fewer surprises turning up from unrelated corners of the system, and fewer late nights spent working out what a small change accidentally broke elsewhere.
Perhaps the biggest shift is one you’d only notice by talking to the team: engineers who were frustrated and starting to look elsewhere are engaged again, because the daily experience of doing the work has genuinely improved. Fixing technical debt properly isn’t glamorous, but it’s the kind of work that pays the business back every single week afterwards.
Facing something similar?
If this sounds like where your business is right now, let's talk about it.
Get in touch