AI is finally being judged on results. Here's how to measure yours.
There’s been a shift in the last eighteen months, and it’s a healthy one.
The first wave of AI adoption was about permission. Boards asked whether they were allowed to use it, IT asked whether it was safe, and everyone asked whether it was going to replace them. Licences were bought, pilots were run, and a lot of very enthusiastic slides were produced.
The second wave, the one we’re in now, is about proof. The question we’re asked most often has changed from “should we be doing something with AI?” to “we’re spending money on AI, so what are we actually getting back?”
It’s a much better question. It’s also one that a surprising number of organisations can’t answer, and the reason is almost never the technology.
Why AI pilots stall
In our experience, AI initiatives rarely fail because the tool underperformed. They fail because nobody wrote down what success looked like before the tool arrived.
Three patterns come up repeatedly.
No baseline. If you don’t know how long it currently takes to produce a monthly report, respond to a customer query or onboard a new starter, you can’t demonstrate improvement. Any figure you quote afterwards is an anecdote, not a measurement.
Vanity metrics. Licence assignments, weekly active users and prompt volumes tell you about activity, not value. A team can be enthusiastically using AI every day and delivering exactly the same output as last year.
Diffuse benefit. “It saves everyone about twenty minutes a day” is genuinely valuable and completely unprovable. Time saved only becomes an outcome when it’s visibly redirected into something the business cares about, more billable hours, faster quote turnaround, fewer escalations, a backlog that finally clears.
The fix isn’t more sophisticated AI. It’s better discipline about what you’re measuring, applied before you start.
The four outcomes worth measuring
We steer clients towards a small number of outcome categories, because a measurement framework nobody can remember is a measurement framework nobody uses.
- Cycle time
How long it takes to get from request to result. Quote-to-send. Ticket-to-resolution. Enquiry-to-first-response. Month-end close. Cycle time is the most persuasive AI metric because it’s already tracked in most businesses, it’s hard to argue with, and customers feel it directly.
What to capture: Measure the median and 90th-percentile duration before and after implementation. The median reflects the typical case, while the 90th percentile shows the slowest 10% of requests. A large reduction in the 90th percentile suggests AI is removing the bottlenecks and reducing the complex cases that previously required the most human effort.
- Throughput per person
Output volume against unchanged headcount. Proposals produced, cases handled, documents reviewed, campaigns shipped. This is the metric that speaks to boards, because it separates growth from hiring.
What to capture: volume per FTE per month, with a note on quality so you’re not just measuring speed. Throughput that arrives with a rework rate attached isn’t throughput.
- Quality and rework
Error rates, revision cycles, complaint volumes, first-time-right percentages. This is the outcome most people skip, and it’s the one that protects the others. Any AI programme that improves cycle time while quietly increasing rework is moving cost around rather than removing it.
What to capture: rework rate, escalation rate, customer satisfaction or NPS on the affected process.
- Cost to serve
The cost of delivering a unit of work: a support ticket, an invoice, a report, an onboarding. Not headline software spend, but the fully loaded cost of the activity. This is the number that turns AI from an IT line item into a commercial decision.
What to capture: total process cost divided by units delivered, tracked quarterly.
You do not need all four. For most organisations, one primary outcome and one guardrail metric is plenty. Cycle time with a rework guardrail is a strong default.
Baseline before you buy
The single highest-value thing you can do is measure your current state for two to four weeks before anything changes. It’s unglamorous, it feels like a delay, and it is the difference between a business case and a hunch.
A workable baseline exercise:
- Pick one process. Something high-volume, well-understood and irritating. Resist the urge to start with the strategically exciting one, start with the one people complain about.
- Define the unit. One quote. One ticket. One report. If you can’t count it, choose a different process.
- Measure for a fortnight. Duration, volume, rework. Spreadsheet is fine. Perfection is not required; consistency is.
- Agree the target before you deploy. Write down the number that would make this worth doing, and who signs off that it’s been achieved.
- Re-measure at 30, 60 and 90 days. Improvements often dip before they climb, as people learn. A single measurement at week two will mislead you in both directions.
What good looks like at 90 days
A well-run first AI deployment tends to produce a fairly consistent picture: meaningful improvement on one process, clear evidence of where the technology doesn’t help, a short list of data and access problems that were invisible until AI exposed them, and a small group of internal advocates who now understand the tool well enough to teach others.
That last one is underrated. The organisations getting genuine returns from AI aren’t the ones with the most licences. They’re the ones where a handful of people got properly good at it and pulled everyone else along.
What good does not look like: a broad rollout to every department at once, with success declared based on adoption statistics.
The uncomfortable part
Some processes won’t improve. Some will improve so much that the bottleneck simply moves downstream, and you’ll have to deal with that too. And some of the value will be real but genuinely hard to quantify, better first drafts, fewer blank-page mornings, more consistent tone across a team.
Be honest about that split. A business case that claims everything is measurable loses credibility at the first challenge. A business case that says “here are the two numbers we’re accountable for, and here’s the qualitative benefit we believe in but aren’t asking you to fund” survives contact with a finance director.
How Redsquid approaches this
Managed AI sits alongside Managed Technology, Managed Cybersecurity and Managed Connectivity for a reason: AI outcomes depend on the foundations underneath them. Poor data hygiene, over-permissive file access and inconsistent identity management all show up as disappointing AI results, because a tool that can reach everything will surface everything.
More importantly, successful AI adoption isn’t measured by how many licences you’ve deployed or how quickly you’ve rolled out a new tool. It’s measured by whether people are using AI in ways that deliver meaningful, measurable outcomes. That’s the difference between organisations that simply have AI and those that are realising its full potential.
That’s why our starting point is always a readiness and outcomes session. We work with you to understand what you’re trying to improve, establish your current baseline, identify anything that needs fixing before AI is worth deploying, and agree on the first process that will deliver tangible results. From there, AI adoption becomes purposeful, measurable and aligned to the outcomes your business cares about, not just another technology rollout. No licences are required to have that conversation.