Most articles on this topic open with a percentage. Someone's vendor claims 80% deflection, someone else's analyst report says AI cuts support costs by a third, and neither number tells you anything about your own queue. The useful version of this question is arithmetic, and it has five inputs you can find in your own helpdesk this afternoon.
This piece walks through each input, the traps inside it, and what to do with the result. If you want the arithmetic done for you, the free support cost savings calculator runs the same maths in your browser.
The five numbers that decide everything
Every credible support-cost model reduces to the same chain. Ticket volume times cost per ticket gives you what support costs today. Deflection rate tells you what share of that the AI absorbs. Platform cost is what you pay to absorb it. The rest is subtraction.
- Tickets per month. Resolved conversations, not messages. If your helpdesk counts each customer reply as a ticket, you will overstate the saving by a factor of three.
- Average handle time. First touch to resolution, including follow-ups and the time an agent spends reading the thread again after a customer replies two days later.
- Fully loaded agent cost. Base salary is the wrong number. Add payroll tax, benefits, tooling licences, recruitment amortisation and the management overhead — usually 1.25 to 1.4 times base.
- Deflection rate. The share of tickets resolved without an agent. This is the contested input; the next section is entirely about it.
- Platform cost. What the AI tool costs per month. A savings figure that omits this is marketing, not a business case.
Cost per ticket: the number nobody agrees on
Cost per ticket is handle time divided by 60, multiplied by the fully loaded hourly cost. A $28/hour agent spending 12 minutes on a ticket costs $5.60 to resolve it. That is the figure that makes the rest of the model work, and it is the one most often quoted wrong — usually because someone used base salary rather than loaded cost, which understates it by 25 to 40%.
Two refinements are worth making before you trust the number. First, exclude tickets your team never touches — auto-closed spam, duplicate submissions — because including them deflates your average handle time and hides the real cost of the work that remains. Second, if your queue splits cleanly into tiers, calculate cost per ticket separately for each. Tier-1 password resets and Tier-2 integration debugging have different handle times, and averaging them across a queue that is 70% Tier-1 produces a number that describes neither.
Deflection rate is a property of your tickets, not of the vendor
This is where most business cases go wrong. Vendors quote deflection rates between 40% and 80%, and both ends of that range are honest — they are describing different queues.
A ticket deflects when its answer already exists somewhere in your documentation and the customer will accept receiving it from a bot. That gives you a practical test: pull your ticket tags for the last quarter and count the share whose resolution was “pointed the customer at an existing article”. That share is your realistic ceiling. Anything the AI achieves above it is coming from answers it synthesised across documents, which is real but not something to plan around.
Queues that deflect at the top of the range are dominated by order status, shipping and returns policy, password and access issues, plan and billing lookups, and feature availability questions. Queues that deflect at the bottom are dominated by billing disputes, account-specific configuration, anything requiring a refund decision, and complaints — where the customer's real request is to be heard by a person.
If your documentation is thin, the deflection ceiling is low for a reason that has nothing to do with AI. In that case the first project is the knowledge base, not the bot. Our guide to building a self-service knowledge base covers what to write first, and the free PDF to Markdown and webpage to Markdown converters show you what structure survives when a document is prepared for retrieval.
Worked example: a 2,000-ticket queue
Take a mid-sized support team with numbers that are unremarkable in either direction:
- 2,000 tickets per month
- 12 minutes average handle time
- $28/hour fully loaded agent cost
- 60% deflection rate
- $59/month platform cost
Cost per ticket is $5.60. Current monthly support cost is $11,200. At 60% deflection, 1,200 tickets stop reaching an agent, which is 240 agent hours freed and $6,720 of gross monthly saving. Net of the platform fee, that is $6,661 a month, or roughly $79,900 a year, with the platform paying for itself in well under a week of operation.
Now run it pessimistically. Drop deflection to 35% — a fair assumption for a queue heavy on account-specific work — and the saving falls to $3,861 a month. Still $46,000 a year. The decision does not actually turn on whether deflection is 60% or 35%; it turns on whether the queue is large enough that agent time dominates the platform fee. Below roughly 500 tickets a month it stops being a cost argument and becomes a coverage argument, which is a different case to make.
Freed hours are not automatically saved money
This is the part vendor calculators skip. Deflection produces agent hours; it does not produce a smaller payroll unless somebody decides it should. Three things actually happen to those hours:
- They absorb backlog. Teams behind on response times convert deflection into faster replies. That is a service improvement and a retention effect, but the finance line does not move.
- They prevent a hire. The most common real saving. 240 freed hours is roughly 1.5 full-time agents you no longer need to add as volume grows.
- They move to higher-value work. Proactive outreach, churn saves, onboarding calls. This usually returns more than the payroll saving would have, but it needs to be decided deliberately rather than drifted into.
Write down which of the three you intend before launch. A projection that quietly assumes payroll reduction, presented to a team that then uses the hours for backlog, is how these projects lose credibility at the six-month review.
The costs the model usually misses
- Documentation work. Almost every deployment needs a content pass first. Budget real hours for it; it is the single largest determinant of the deflection rate you will actually get.
- Integration time. Connecting the helpdesk, the order system and SSO. Usually days rather than weeks, but not zero.
- Ongoing maintenance. Docs drift. A bot grounded on last quarter's refund policy deflects less and, worse, deflects wrongly.
- Escalation handling. Deflection is never total. The 40% that still reaches an agent arrives pre-triaged, which is good, but it still needs staffing.
How to run the numbers on your own queue
Pull four figures out of your helpdesk: monthly resolved tickets, median handle time, your loaded hourly cost, and the share of last quarter's tickets whose resolution already existed in documentation. Put them into the support cost savings calculator with that documented share as your deflection rate — not the vendor's number — and read the net monthly saving and the payback period.
Then run it a second time at half that deflection rate. If the case still holds, it is a decision. If it only works at the optimistic number, it is a hope.
Where CustomerGPT fits
CustomerGPT indexes your website, help centre and documents and answers from them with citations, so deflection is measurable on your own queue rather than estimated from a case study. Pricing starts at $59/month, with 14 days free and no card required — enough to replace the estimated deflection rate in your model with a measured one before you commit.
For the mechanics of getting your documentation in, see the knowledge base integration guide. For what the platform costs at each tier, see pricing.