That'sGonnaHelp
Support

Ticket Deflection Rate: Test What AI Actually Saves

AI resolution claims can overstate the work your team avoided. Use a randomized holdout, reconcile repeat contacts across channels, and compare human effort before turning ticket deflection into an ROI estimate.

Alex KhvoinitskiiNovember 11, 202517 min read

TL;DR: A ticket deflection rate does not prove how many human contacts AI prevented. Compare eligible customers with a randomized holdout, track repeat contacts, and measure the human work left behind before claiming savings.

What does ticket deflection rate actually measure?

Ticket deflection rate describes the share of support needs handled without a human agent, under a stated counting rule. It does not, by itself, establish how many tickets AI prevented. Some customers would have found an answer through your existing help center anyway; others may have left with an unresolved problem.

For a useful support ticket deflection rate, separate verified AI-only resolution from incremental reduction in human contacts. The first asks whether an AI-supported issue reached a confirmed outcome without human work or a same-issue return during the observation window. The second asks what changed because customers received the AI experience. Start with the AI customer support automation guide to choose the permitted workflow; this guide covers the measurement experiment.

A vendor success story can suggest what to investigate, but cannot supply your missing comparison group. Klarna reported 2.3 million AI assistant conversations in its first month, covering two-thirds of its customer service chats. Klarna reported a 25% drop in repeat inquiries in its first-month AI assistant announcement; this is company-reported, not a randomized result. Those claims come from Klarna's February 27, 2024 announcement; neither is a forecast for a small support team.

Where is a support deflection rate test useful?

A deflection experiment is useful when a team can identify the same customer across channels and observe a clear support outcome. Start with a frequent, low-risk question and a working human handoff. Keep the existing service available to the holdout, which is the comparison group that continues receiving the normal experience.

These five scenarios offer concrete starting points:

  • Ecommerce order status: compare customers who see an AI explanation with those who see the existing tracking page. Match later requests to the same order. The WISMO workflow guide explains how to handle delays without making unsupported delivery promises.
  • B2B software access: compare AI-guided account recovery with the normal help flow. Verify successful access, and count any agent work needed to finish it.
  • Home services appointments: compare AI confirmation help with the existing appointment page. Track whether the same customer later calls about that visit.
  • B2B product setup: answer routine configuration questions from approved documentation. Match escalation emails to the original account and setup issue.
  • Professional services onboarding: compare help finding an already-issued document with the current portal instructions. Treat a completed download as evidence only when document access was the stated need.

When it is not a good fit

A simple holdout is a poor fit when outcomes cannot be observed, customers cannot be matched, or volume is too low to distinguish a useful change from noise. Repair the measurement or extend collection before using the result for a staffing decision. For high-risk issues such as safety instructions or disputed financial actions, keep people responsible; a favorable rate does not justify expanding autonomous scope.

How to measure support ticket deflection with a holdout

To measure support ticket deflection caused by AI, randomly assign eligible customers to an AI experience or the existing support experience before either can affect the outcome. Compare the share in each group whose issue reaches the human queue within the same fixed window. This estimates reduced demand for human support, while a separate outcome check protects against customers simply giving up.

Microsoft's research on controlled experiments explains why sound randomization supports causal conclusions. A before-and-after ticket chart lacks that protection: demand can fall because a promotion ended, a bug was fixed, or order volume changed. The following six steps are a practical support application of that experiment design.

  1. Define entry before showing AI. Enroll a customer when they submit an eligible question through the shared support entry point. Record eligibility from information available then, such as order status and selected intent. Do not select the cohort from people who successfully received an AI answer.
  2. Keep assignment stable. Store an AI or holdout assignment against an existing customer ID in the support system. For a simple first test, analyze only the first eligible issue per customer, including all follow-ups to that issue. If several users share one B2B account, assign the account together and use an analysis that accounts for that grouping.
  3. Preserve the normal fallback. Give the holdout its existing help content and agent access. Keep agent access available to the AI group too. A customer choosing a person remains in their assigned group; so does a failed AI delivery.
  4. Freeze the measurement rules. Choose the enrollment dates, follow-up window, minimum worthwhile effect, quality limits, and planned sample size before launch. A 50/50 split is one straightforward planning choice, not a requirement. Check assignment and event collection with test accounts before enrolling real customers.
  5. Join the evidence. Use your help-desk export, chat event log, order or account IDs, phone dispositions, and a spreadsheet or database. Count an issue as reaching humans when a same-issue request enters a human queue, even if an agent has not replied yet.
  6. Read completed observation windows. Compare assigned groups only after each included issue has had its full follow-up period. Report missing outcomes and quality results alongside the rates. Decide at the planned review point; do not end collection simply because today's number looks favorable.

Stable assignment matters because a returning customer should not switch experiences mid-test. The online evaluation research survey discusses user-level assignment and the risks of treating related observations as independent. For this guide's simple example, every enrolled customer contributes one issue, so the customer and issue denominators match.

Human-contact rate for each group =
issues with at least one human-queue request / all enrolled issues

Absolute reduction, percentage points =
(holdout human-contact rate - AI human-contact rate) × 100

Estimated avoided human-contact issues in the AI group =
(holdout rate - AI rate) × number enrolled in the AI group

This counts issues that need human contact, not every message or ticket record. Report repeated tickets and total human effort separately. If the holdout rate is zero, a relative reduction percentage is undefined; retain the absolute difference and raw counts.

Reconcile repeat contacts before closing the cohort

Count repeat contacts across email and chat by matching customer, intent, and the relevant order or account event to the original issue. Keep every assigned issue in the denominator, including unresolved cases and failed AI attempts. A cohort is ready for analysis when its observation time has elapsed, not when every ticket happens to be marked closed.

Use one issue summary with linked events underneath it. This avoids turning three channels into three apparent customer problems. The following fields are a starting worksheet, not a requirement to collect new personal information.

Field What to record Why it matters
Issue and customer key Existing customer ID plus order, appointment, or account issue Connects follow-ups without relying on the chat session alone
Entry and assignment Eligible entry time, assigned group, reason for eligibility Prevents selection based on a successful AI outcome
Follow-up deadline Entry time plus the agreed observation window Separates pending observation from mature results
AI events Delivered answer, claimed resolution, failure, handoff Shows what the assigned experience actually did
Human-contact events Queue entry, channel, same-issue match, repeated requests Finds demand that moved to another channel
Outcome evidence Explicit confirmation or a completed action matching the need Separates a solved problem from silence
Human effort Active minutes across all linked follow-ups Measures work, including escalation and rework
Data quality Missing channel export, uncertain identity, late event Makes blind spots visible

For an illustrative seven-day window, a Monday entry becomes eligible for final reporting the following Monday. Seven days is a planning choice for routine work, not a universal standard. A different workflow may need longer, and later contacts should remain visible even when they fall outside the primary window.

Use mutually exclusive result buckets when auditing AI-only claims: verified outcome, known failure or repeat, unknown evidence, and still pending observation. An unknown outcome contributes no verified resolution. Do not silently drop it or pretend that it proves a failure; show the count and test how unresolved uncertainty could change the decision.

Missing identity is especially dangerous. If a phone export lacks usable customer links, an apparent reduction could be channel switching. Keep the enrollment count, report the coverage gap by group, and withhold a causal savings claim until matching is repaired or a defensible sensitivity analysis supports it. Preserve event timestamps and export dates so late-arriving records can restate a prior report.

A worked case: 560 AI claims become 120 avoided contacts

This invented operator composite shows how one measurement can produce several valid but different numbers; it is not a public customer claim. A small ecommerce support team has an AI order-status assistant, an existing tracking page, and a shared help desk. Its dashboard reports 560 AI resolution claims among 1,000 AI-assigned customers, but the owner wants to know what work actually disappeared.

The team enrolls 2,000 distinct customers with one eligible issue each and randomly assigns 1,000 to each experience. It logs assignment before displaying either answer path, links help-desk exports to order IDs, and records phone follow-ups in a spreadsheet. Both groups keep the same route to a person. The example assumes complete human-contact capture and independent customers, with no shared accounts.

The first export misses some phone dispositions, which makes the AI result look better than it is. The owner repairs the import and reruns the same cohort after the follow-up deadline. Of the 560 claims, 70 have a qualifying same-issue return, 30 have a known unresolved outcome, and 20 lack enough resolution evidence. The remaining 440 have confirmed outcomes, no human work, and no qualifying return: a verified AI-only resolution rate of 44% of all AI-assigned issues.

Meanwhile, 420 holdout issues and 300 AI-assigned issues reach the human queue. Human-contact rates are 42% and 30%, so the absolute reduction is 12 percentage points. Relative to the holdout's rate, that is a 28.6% reduction; applied to the 1,000 AI-assigned issues, it estimates 120 avoided human-contact issues. The 440 verified outcomes cannot all be credited as incremental savings because the existing tracking page also helps customers without agents.

The analyst also checks uncertainty. Using the large-sample difference-of-proportions method documented by NIST, these illustrative counts give an approximate 95% interval of 7.8 to 16.2 percentage points for the reduction. That interval assumes independent observations and complete contact data; it does not correct missing channels, blocked escalation, or poor service. Resolution audits and complaints still need to pass the pre-agreed quality checks before any expansion.

After all cohort work is complete, the example logs show 84 human hours for the holdout and 60 for the AI group, including repeats and escalations. The difference is 24 hours, or $960 of gross capacity value at an assumed $40 loaded hourly rate. Eight additional hours of AI review use $320 of that value; $480 of other recurring pilot costs leave a $160 capacity-based net benefit. These effort figures are separate illustrative observations, not a promise that every avoided issue saves 12 minutes.

With $1,200 of setup cost, the same recurring result would imply 7.5 months of simple capacity-based payback. It does not demonstrate cash payback: the owner must identify spending that will actually fall or funded work that will use the released hours. The full customer support automation ROI model is the next step once those measurement inputs are credible.

What does a ticket deflection pilot cost in USD?

A measurement pilot costs staff time, data preparation, AI usage, and ongoing review; there is no single market price for this scope. The worked example below assumes $1,200 upfront and $800 per comparable monthly period in recurring costs. Obtain actual license and usage terms before budgeting, and keep measurement costs separate from a full support-system replacement.

Intercom announced Fin over email at 99 US cents per resolution in its September 2024 product broadcast. This historical unit price is not a current quote or the full platform cost. The dated Intercom announcement is a pricing reference, not the basis for this example's assumed usage bill.

Cost item Illustrative USD assumption Treatment in the model
One-time measurement setup $1,200 Export mapping, assignment checks, and report setup
AI usage per monthly cohort $300 Replace with the actual bill and its billing definition
Other recurring tools and upkeep $180 Incremental reporting, integration, and knowledge costs
AI review labor 8 hours × $40 = $320 Additional effort outside the 60 human handling hours
Total recurring cost $800 $480 other costs plus $320 review labor
Gross human capacity released 24 hours × $40 = $960 A value estimate before recurring costs
Net capacity value $160 $960 minus $800; cash savings unproven
Labor-rate sensitivity $30–$50 per hour Assumed range gives $0–$320 monthly net capacity value at fixed hours and other costs

The labor rate is an assumption, not a wage benchmark. Holding the example's hours and nonlabor costs fixed, an assumed $30–$50 hourly planning range changes monthly net capacity value from $0 to $320. At zero net benefit there is no finite payback; at $320 it would be 3.75 months if the result repeats. Setup cost is held fixed for this sensitivity check.

Test the assumptions in an automation ROI calculator, especially a lower contact reduction or extra review work. Do not subtract QA twice: either value the 24 gross hours and subtract $320 of review cost, or value the 16 hours left after review. Keep the method consistent.

Build a customer service metrics dashboard from one issue log

Track customer service metrics by giving each dashboard number an issue-level source, a fixed denominator, and an observation deadline. To turn avoided contacts into saved agent hours, compare active human effort for equivalent assigned groups. Include repeats and escalations, then subtract added review work. A lower ticket count alone cannot tell you how much labor disappeared.

The useful customer service metrics to track are few enough to review in one meeting. Name the report “Support Deflection Rate: Measure What AI Answers Actually Saved Your Team,” and put the experiment dates and definitions directly below it. Each percentage should open to the cases behind it.

Dashboard line Required companion
Eligible customers assigned to each group Enrollment rule, split, and observed assignment counts
Human-contact rate and rate difference Numerators, denominators, interval, and channel coverage
Verified AI-only resolution rate Evidence rule, repeat window, unknown and pending counts
Service quality Confirmed errors, complaints, unresolved needs, satisfaction response counts
Human work saved Active handling, repeats, review, and open work still consuming time
Capacity or cash value Labor assumptions, full costs, and the named use of released capacity

Use a time and motion baseline worksheet when the help desk records elapsed time but not active work. Adjust comparisons for group size when allocation is unequal. With clustered B2B accounts or repeated issues per customer, use an analysis that respects those groups instead of applying the simple independent-customer interval above.

Common mistakes

  • Comparing different denominators. An AI resolution percentage and a whole-queue contact reduction answer different questions.
  • Removing failed AI deliveries. Keep them in their assigned group, or the report selects only the easy successes.
  • Calling silence a solved issue. Keep unknown outcomes visible, and verify completed actions or customer confirmation.
  • Celebrating a hidden phone queue. Reconcile every relevant channel before claiming avoided demand.
  • Converting every hour into cash. Released capacity matters, but payroll savings require an actual spending change.

FAQ

A defensible deflection report combines outcome evidence, a fair comparison group, and complete follow-up. These answers cover the practical choices that can change what your number means.

What is deflection in customer support?

It is resolving a customer's need through self-service or automation before human support is required. Support case deflection can happen even when software creates a background ticket record. Define the human work and outcome involved, rather than relying on whether a database row exists.

Can a closed AI chat count as a saved ticket?

It can be a candidate for verification, but closure alone is insufficient. Check the outcome, later same-issue contacts, and human work. Even a verified AI-only outcome needs a comparison group before you can estimate whether AI prevented contact that otherwise would have happened.

How long should a ticket deflection test run?

Choose the sample size and enrollment period from your expected contact rate and the smallest useful change, then allow the final entrant's full follow-up window to finish. More days alone do not guarantee a clear result. Fix the analysis point in advance and report an inconclusive estimate when uncertainty remains too large.

What if an SMB has too little support volume for a holdout?

Use a longer observation period if the workflow stays stable, or run a smaller descriptive pilot focused on verified outcomes and human effort. Label before-and-after changes as directional evidence. Do not turn a handful of clean chats into a confident annual staffing forecast.

Can we compare last month with this month instead?

Yes, as an operational comparison with explicit limits. Account for order volume, issue mix, outages, staffing, and policy changes, and state what you could not control. It is weaker evidence of AI's causal effect than a sound concurrent randomized holdout.

What if an AI customer asks for a person immediately?

Honor the request and retain the issue in the AI-assigned group. Count its human-queue contact even if the AI never answers. Excluding immediate handoffs would make the experiment describe only customers who tolerated or used the AI path.

Answer clarity notes

The calculations here illustrate measurement choices; they do not report a measured That'sGonnaHelp customer result. Public claims, example assumptions, and recommendations have different evidentiary weight.

  • Dates: The article carries a November 11, 2025 date. Klarna's figures describe its February 2024 announcement, and the historical Intercom unit price comes from September 12, 2024. Check current pricing and billing definitions before purchasing.
  • Public evidence: Microsoft and NIST support the cited experiment and statistical methods. Klarna's results remain company-reported; its announcement does not establish this guide's proposed holdout design.
  • Examples: The ecommerce case is an invented operator composite, not a public customer claim. Counts, hours, USD budgets, labor ranges, payback, and the calculated interval are illustrative inputs and outputs. Cost ranges, savings, and timelines are planning guidance, not guarantees.
  • Measurement limits: A favorable contact difference does not prove good resolution, and an interval does not repair missing data or faulty assignment. The simple interval assumes independent customers with one enrolled issue each.
  • Recommendations: The worksheet, 50/50 split, and seven-day illustration are planning choices. Follow-up windows and sample sizes must fit the actual workflow; none guarantees a conclusive result.
  • Scope: This is operating guidance for US SMBs, not financial, legal, tax, medical, compliance, or platform-policy advice. Capacity value becomes cash savings only when a corresponding expense falls.

Sources

These sources support the public claims and measurement methods used above. The budget and support experiment are separate illustrative applications.

That'sGonnaHelp can help turn support exports into a measurement plan and a testable pilot. Bring the resolution claims, channel gaps, and cost assumptions you want to verify.

A

Alex Khvoinitskii

Founder, That'sGonnaHelp

Founder of That'sGonnaHelp. Building growth and automation systems since 2021 — GTM, traction, retention, and revenue — for SaaS, FinTech, and e-commerce clients, from early-stage brands to global exchanges.

Related articles

Discuss your project