A&A INSIGHTS
Before letting AI receive support inquiries, define the conditions for handing them back: measure wrong resolutions, not just answer rate
For owners of small companies: use Intercom's Fin case and Anthropic's agent-design guidance to plan an AI support setup that measures wrong resolutions and handoff burden alongside answer rate.
日本語で読むTHE STARTING POINT
Increasing the share of inquiries an AI can answer does not by itself reduce the work of catching wrong resolutions or reconstructing cases that come back to a person. Using Anthropic's published Intercom case and its engineering essay on building effective agents as a lens, this article suggests a small-company setup in which the escalation conditions and handoff format are defined before an AI product is chosen. A hypothetical online shop illustrates two burdens to measure alongside answer rate.
"Percent answered by AI" hides the wrong-resolution cost
Hypothetical example
Imagine a five-person online shop considering an AI chat for its inquiries. The starting motivation is to cover replies outside business hours and reduce time spent on repetitive questions. This is a hypothetical example. Vendor materials tend to place "first-response rate" and "resolution rate" at the center of comparison. Even when those rates are high, an inquiry in which a return was wrongly declared out of scope, one in which shipping was waived without authority, or one in which a change of address was acknowledged but never propagated to shipping does not appear on the same table.
A&A perspective
A&A's suggestion for owners of small companies is to judge an AI introduction by the movement of cases a person has to pick up later, not by the increase in the count of AI responses. Answer rate only moves upward, but the count of cases picked up later can move in either direction. Before choosing a product, decide on a way to measure which direction it moves in. This is not a fight over the numbers themselves; it is a way to separate which direction actually helps the business.
Read the Intercom case as one part of the outcome, not the whole
From the sources
Anthropic's Intercom case page describes the company's AI agent Fin as capable of an up-to-86% resolution rate together with human-quality support responses. The same page also notes that Fin achieves an average 51% resolution rate across customers.
A&A perspective
The important qualification is that the case is centered on the share Fin returns and how that share was operated upward. Under what conditions the roughly half that Fin does not resolve is routed to a person, and how many of the returned answers turn out to be wrong and how they are detected, are not the subject of that document. Reading the case is useful for enumerating what your own company should measure, not for adopting the same rate as a promise. Smaller shops are more exposed to a headline rate because the denominator is small.
A&A perspective
A&A's reading is that what a smaller company should carry away from a large case like Intercom's is the design view that "AI answered," "AI handed to a person," and "a person had to pick this back up" can and should be tracked as three separate signals rather than folded into one resolution rate. If they are folded, the effort of picking back up disappears from the report.
Start with the simplest procedure, and define the handoff conditions first
From the sources
Anthropic's engineering essay "Building effective agents" advises verifying that the smallest possible procedure suffices and adding multi-step agent constructions only when a simpler solution falls short. The same essay describes a routing pattern that classifies inputs and directs them to a specialized follow-up task, and treats returning to a person for information or judgement as a normal part of the loop.
A&A perspective
Translating this into a small company's inquiry desk, the first task is to write down, before choosing an AI product, which kinds of inquiries must go to a person. Read one to two weeks of the actual inquiry history and mark the cases the answering employee treated as beyond the first line. Refunds, cancellation fees, communication that touches regulation, changes to an existing customer's contract terms, and requests that require identity verification are typical. For these, decide that AI performs only intake and acknowledgement, and is out of scope for the reply itself.
A&A perspective
With the handoff conditions decided in advance, comparison of AI products narrows to "within the scope you actually delegate." A "90% resolution" product that raises its number by answering across categories you did not intend to delegate is not a useful number for you. If you plan to delegate only simple FAQ, it suffices to check two things: response quality within that scope, and the accuracy of the handoff when a case leaves that scope for a person.
Count wrong resolutions separately from the burden of handoff
A&A perspective
An inquiry received by AI ends in one of three ways. First, AI answers, the customer accepts, and the reply is consistent with internal rules. Second, AI answers, the customer accepts in the moment but a discrepancy becomes visible later. Third, AI does not decide and passes the case to a person. What is good for the business is the first case growing, the second case shrinking, and the third case being one where "the cases that ought to come back reliably come back." Vendor resolution rates are often defined so the first and second are summed in the numerator and the third is removed from the denominator; without checking the definition, the invisibility of the second case remains.
Hypothetical example
Concretely in the hypothetical online shop: for an inquiry about returns, AI replies "out of scope," and the customer accepts on the spot. A few days later the customer surfaces the same case through another channel, for example the payment processor's inquiry desk or a social post, and only then is the mistake noticed. If "cases detected after the fact" are not counted, surface figures such as reduced inquiry volume and reduced handling time can take on a life of their own. Compare before and after by watching related contacts through the payment processor and social channels as well.
A&A perspective
The third case, "cases passed to a person," cannot be judged by count alone either. Even if their count is low, if any of the customer's own words, the content AI presented, or the internal source AI consulted is missing at handoff, the person has to ask the customer the same question again. From the customer's perspective this is a duplicated contact; from the handler's perspective the time per case can be longer than it would have been without AI. Record "average handling time for handed-back cases" and "number of re-questions to the customer," alongside your pre-AI baseline.
Decide the handoff content before the contract
A&A perspective
Before signing with an AI product, put on a single sheet what must accompany a handoff. As a minimum, six items should arrive as one unit: the customer identifier, the inquiry text, the full text of what AI replied, the ID of any knowledge article AI referenced, the reason AI decided to pass the case to a person, and candidate categories for the human response. With these six items present, the handler can reconstruct the situation from internal records before asking the customer any additional question.
Hypothetical example
In the same shop example, suppose a returns question is passed from AI to a person. If the passed information lacks "the ID of the knowledge article AI referenced," the handler cannot rule out that AI was reading an outdated policy version, and ends up searching for the article themselves. Conversely, if the full AI reply is preserved, the handler can decide on the spot whether the first response was wrong and, if it was, send a correction to the customer first. Decide the handoff content before you decide the contract terms.
A&A perspective
With a handoff format decided, AI-product comparison can proceed on "does it export in that format?" and "when a field is missing, which item is dropped?" Many vendors have their own standard UI, but check whether records can be written in the shape the handler actually reads, not the shape of the standard UI. This check separates into cases resolvable by AI-side configuration and cases that require a receiving surface in your internal customer-management tool. If the estimate does not include the latter where relevant, you will end up in a state where "cases return but cannot be read."
A checklist for a small-company trial
A&A perspective
The trial that supports the introduction decision does not need to be a large A/B test. Restrict AI to a narrow delegated scope for one or two weeks and count that period's inquiries under the three-way split above. In parallel, take a random sample of inquiries that arrived in the non-delegated scope and have a handler re-read them to confirm none were mishandled through AI. Ten to twenty percent of the total is enough to reveal a tendency while keeping the handler's added workload bounded.
A&A perspective
At the end of the trial, do not look only at the change in response rate. Cases detected after the fact should be at zero or no higher than the pre-AI baseline. Average handling time for cases handed back should be shorter than the baseline, or at least equal. The random sample from the non-delegated scope should show no mishandling. If any one of these fails, either narrow the delegated scope or rebuild the handoff format. The decision to stop the introduction is also made at this stage, not later.
A&A perspective
When A&A is asked to help with inquiry design, the same trial design and the same articulation of the delegate-versus-return boundary are the starting points. The boundary is not there to reject a given AI product; it is there to keep the product within the range where it is actually good. At the three-month and six-month marks after introduction, check whether the burden of handed-back cases is quietly eating into the gross margin on the business side. The related article "When identical questions stop your team's hands: making support capacity with the Intercom case" handles the same Fin case from the side of building capacity.
When you decide whether to let AI receive inquiries, the number worth comparing is not just the share it answers. Put wrong resolutions and the burden of handing cases back to a person into a form you can count before you introduce anything. Naming the delegated scope in advance, and defining the handoff content in advance, is the starting point that lets a small company get real business value from an AI support setup.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Intercom provides customer service tech that delivers up to 86% resolution rates with Claude
Anthropic · Publication date not stated on page
Accessed 2026-09-17 - Building effective agents
Anthropic · 2024-12-19
Accessed 2026-09-17
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-17