
A platform owner I spoke with recently was three years into a managed services contract. The service was, by everyaccount, fine. Tickets closed inside SLA, the monthly report was green, the provider was responsive, pleasant, andrenewed without argument.
So I asked what should have been an easy question: over those three years, did the platform get better?
The pause went on long enough to become the answer. He could produce SLA attainment, ticket volumes, satisfactionscores, and a stack of monthly reports going back twelve quarters. What he could not produce was any evidence that hisplatform in year three was in better shape than his platform in year one. Not worse. Just — not measurably better.Meanwhile his contract, like nearly every managed services contract signed in the last decade, listed continuousimprovement as an explicit deliverable.
He was not an outlier. He was the median.
The gap between "continuous improvement is in our SOW" and "our platform is measurably better every quarter" is not avendor-quality gap. Most managed services providers are staffed by capable people working hard. It is a measurementgap, an incentive gap, and a contract-design gap — and the third one causes the other two.
Continuous improvement appears on essentially every managed services statement of work in the market. It is also,almost without exception, the only deliverable with no defined numerator. Response times get measured. Resolution timesget measured. Ticket closure rates, availability, and satisfaction all get measured, reported, and argued about.Improvement gets asserted.
That is not an oversight. It is a structural consequence of how most managed services are priced. When a provider bills bythe hour, by the ticket, or by the seated FTE, efficiency is a revenue reduction. Researchers at the University ofTennessee named this pattern the Activity Trap: transaction-based contracts pay the provider for volume, which givesthe provider no economic reason to reduce the non-value-added volume. The uncomfortable implication is that mostmanaged services contracts are not failing to deliver continuous improvement. They are working exactly as designed.
This paper argues that continuous improvement is a measurable claim or it is nothing, and that it has to be measured onthree axes rather than one: the platform gets healthier, the delivery capacity gets more productive, and the intake getssmarter. It also takes seriously the reasons the conventional model has survived so long, and the reasons the alternativeskeep collapsing. The thesis is straightforward: stop buying continuous improvement as a promise, and start buying it as aninstrument.
Continuous improvement has strong doctrinal backing. ITIL 4 treats continual improvement as one of its foundationalpractices and prescribes a Continual Improvement Register to hold improvement opportunities in one place with ownersagainst them. The doctrine is sound. The doctrine also insists that improvement is everyone's responsibility — and inpractice, a responsibility distributed to everyone lands on no one. The register gets created during onboarding, populatedenthusiastically for two quarters, and then quietly stops being a working document.
The suspicion that process maturity and actual improvement are only loosely coupled is not new. A global survey ofexecutives conducted by Compass in 2005–2006 found that roughly seventy percent could see no link between theirprocess maturity and any measured performance improvement. The study is old and the sample was small, so treat thefigure as directional. But two decades on, ask a room of platform owners whether their managed services provider hasdemonstrably improved anything and watch how many can produce a number.
The reason is simple. Nobody agreed what the number would be. "Continuous improvement" entered the contract aslanguage, not as a metric with a baseline, a cadence, and an owner. And unmeasured deliverables do not get delivered —they get described.
Here is the mechanism, stated plainly.
If a provider is paid per hour, every hour saved is revenue lost. If a provider is paid per ticket, every recurring incidentpermanently fixed is an annuity destroyed. If a provider is paid per seated FTE, then automating a role out of existence isa self-inflicted wound. The University of Tennessee's work on outsourcing models is direct on this point — transactionbased pricing carries an inherent perverse incentive not to reduce the transactions. Cost-plus fares no better, becausewhen margin is a percentage of direct cost, reducing cost reduces margin.
Formal economics reaches the same conclusion from another direction. The academic literature models IT outsourcing asa double moral hazard problem: both parties can under-invest in ways the other cannot observe, and many real contractsinadvertently pay the vendor even when the outcome fails.
None of this requires anyone to behave badly. A provider under an hour-metered contract who invests heavily inautomation is choosing to reduce their own revenue, quarter after quarter, with no mechanism to share the gain. Some doit anyway, out of professionalism or long-game account strategy. But structuring a service so that improvement dependson the provider's willingness to lose money is not a strategy. It is a hope.
The fix is not to demand more virtue from providers. It is to stop asking them to choose between improving your platformand making their number.
The industry has an idiom for this: the watermelon SLA. Green on the outside, red at the core. It circulates widely inservice management circles and it is practitioner shorthand rather than studied phenomenon — but it names somethingreal, because conventional service metrics are structurally incapable of detecting decay.
First response time, resolution time, closure rate, and availability all measure compliance at a point in time. Every one ofthem can be met, every month, for three consecutive years, while the platform underneath quietly accumulates customization, drifts from the vendor baseline, grows a thicket of recurring incidents that get closed rather than solved,and becomes progressively more expensive and more frightening to upgrade. Nothing in a standard SLA pack will tell youthat is happening. The reports are not lying. They are answering a different question.
The cost of that blind spot is well documented. McKinsey has put technical debt at roughly forty percent of the typical IT balance sheet, with companies paying an additional ten to twenty percent on top of project costs simply to service it. TheConsortium for Information & Software Quality estimated the cost of poor software quality in the United States atapproximately $2.41 trillion in 2022. Forrester's older work put seventy to seventy-five percent of a technology budget intomerely keeping the lights on. These are broad, directional figures rather than measurements of your instance — but theydescribe the gravity every platform is subject to between go-live and the next re-platform proposal.
On ServiceNow specifically, that gravity has a schedule. Two family releases ship every year and only the current andimmediately prior versions are supported, which means at least one upgrade annually whether you planned for it or not.Every accumulated customization makes that upgrade more expensive, and eventually blocks you from adopting theplatform capabilities you are already paying for. A managed service that closes tickets inside SLA while that ratchettightens is not neutral. It is a service that is losing you money in a currency nobody on the call is counting.
If improvement is going to be a real deliverable, it has to be measured — and measured on more than one axis, becausethe three things that need to improve are genuinely different things. The useful discipline is to state, for each axis, whatthe industry promises and what should actually be on the report.
What's promised: proactive monitoring, platform health reviews, best-practice guidance.
What's measured: the ratio of configuration to customization over time. Recurring-incident rate and the problem-toincident ratio — because a service that closes the same incident forty times has closed forty tickets and solved nothing.Upgrade effort per release, in hours and regression defects, which is the most honest available proxy for how far theinstance has drifted from baseline. Performance headroom on the busiest tables. Every one of these should trend in astated direction, and every one should appear on the same report each quarter, next to the previous quarter.
What's promised: AI-enabled delivery, accelerators, reusable assets, an efficient team.
What's measured: throughput per unit of contracted capacity, tracked quarter over quarter. If the provider's tooling andplatform knowledge genuinely compound, the same money buys more delivered work in year two than in year one — andthat is a number, not an adjective. The important structural point is that this metric only survives in a contract where theprovider keeps some of the gain. Under hour-metering, productivity improvement is a revenue cut, and the number willnever be published because nobody wants to look at it.
What's promised: a ticket queue, an intake form, a monthly review.
What's measured: decision latency and rework. How long does a request wait for a decision rather than for capacity? What proportion of delivered work gets reopened because the requirement was wrong on arrival? A co-owned, jointlyprioritized queue produces different numbers than a request desk, and the difference is visible within two quarters. This is also the axis where the customer's own performance shows up honestly, which is precisely why it belongs on the report.
The encouraging development is that the excuse for not measuring any of this has expired. The instrumentation now shipswith the platform. Instance Scan, Upgrade Center, Health Log Analytics, Platform Analytics, Application PortfolioManagement, and the Automated Test Framework generate most of these numerators natively — and ServiceNow's ownrecent releases have added deflection, self-solve, and productivity KPIs to the standard reporting surface. The data isalready being produced. The only question is whether anyone has agreed to look at it on a cadence.
Anyone arguing for a different model owes the incumbent model a fair hearing, and the conventional one has real virtues.Time-and-materials is predictable, transparent, and auditable. The buyer sees exactly what they are paying for, retainscontrol over prioritization, and needs to trust the provider very little. That combination has kept it dominant for decades,and it deserves more respect than the "body shop" sneer it usually gets.
The alternatives, meanwhile, keep failing. Outcome-based pricing has been the future of IT services for twenty years andremains a rounding error: ISG's practitioners estimate it accounts for under ten percent of application development dealseven at top-tier providers. The reasons are well catalogued. Attribution collapses under scrutiny — when three partiestouch a system, proving whose work produced the improvement is genuinely hard. Baselines get disputed and metrics getgamed. Everest Group has warned that outcome metrics risk becoming aspirational dashboards rather than contractuallevers, and that in multi-vendor environments providers simply price the transferred risk back in. And when outcome contracting fails at scale, it fails expensively: the US Social Security Administration's Next Generation Telephony Project consumed over $160 million and was abandoned within roughly ten months of go-live.
There is a deeper objection, and it is the strongest one. Most of the levers that actually produce improvement do not sitwith the provider. Requirements quality, testing capacity, change management, stake holder alignment, and decisionlatency are all overwhelmingly customer-side. No provider can contractually guarantee an outcome that depends on howfast your steering committee makes up its mind. Which is the point: continuous improvement cannot be a vendor warranty.It has to be a shared instrument, with both parties' numbers on the same page.
Finally, a word on AI, because this is where the industry is currently least honest. The claim that AI makes delivery capacity compound is plausible, increasingly supported, and also not yet safe to assume. A randomized controlled trial published by METR in 2025 found that experienced developers took about nineteen percent longer on real tasks whenusing early-2025 AI tooling — while estimating afterwards that the tools had made them twenty percent faster. The study was small and its subject tooling is already historical, so it is not the last word. But the perception-versus-reality gap it documents should end the practice of selling AI productivity as a given. If a provider tells you their AI-enabled delivery gets more done per dollar, the correct response is not enthusiasm. It is: show me the throughput number, this quarter against last.
Continuous improvement is the most common promise in managed services and the least frequently instrumented. That isnot because it is unmeasurable. It is because measuring it would expose the fact that most contracts are structured toprevent it.
So before signing or renewing:
The bottom line is that a managed service is not a subscription to activity. It is a claim about the state of your platform overtime — and the only honest version of that claim is one where the numbers are agreed up front, published on a cadence,and structured so that the provider makes money by making them better.
Three years from now, someone is going to ask whether your platform improved. Decide today what you will show them.