Butter Labs Request demo

Containment versus resolution: why the metric you track shapes the system you build

A containment rate counts conversations that never reach a person. A resolution rate counts customer tasks that are verified as complete. One measures what the system avoids; the other measures what the system achieves. The metric you choose during procurement shapes the design of the system you buy, the incentives of the team that runs it, and the experience of the customer who calls.

What containment rate actually measures

Containment rate is the proportion of inbound contacts handled without transfer to a human agent. It rose as the standard metric for IVR and chatbot programmes because it directly maps to agent labour cost. A higher containment rate means fewer calls reach the queue.

The problem is that containment says nothing about whether the customer's task was completed. A caller who gives up after three failed attempts counts as contained. A caller who hangs up and calls back counts as a new contact, not a failure. A caller routed to a dead-end self-service path counts as contained until they find another channel.

When containment is the primary target, the system is optimised to keep calls away from agents. When resolution is the primary target, the system is optimised to finish the customer's task.

What resolution rate measures instead

A resolution rate starts with an eligible task, checks whether the expected completion event occurred in the connected business system, and then watches for repeat contact within an agreed window. It treats a confirmed outcome as the unit of value, not a prevented transfer.

This requires more instrumentation. Your team needs a task identifier, a completion event, integration with the system of record, and a repeat-contact matching rule. The metrics scorecard describes these definitions in detail.

The return is a measure that answers the question your customer asked: did this get done?

How the metric shapes the buying decision

Teams that evaluate voice AI by containment rate will tend to select systems that are good at deflecting calls. Teams that evaluate by resolution rate will tend to select systems that are good at completing tasks. These are different designs with different integration requirements, different quality review methods, and different cost structures.

ComparisonContainmentResolution
Primary success metricCalls not reaching an agentTasks verified as complete
Integration depthMinimal telephony and routingDeep business systems and task records
Transfer handlingMinimise all transfersTransfer when system cannot verify completion
Repeat contact treatmentSeparate metric or invisibleBuilt into resolution definition
Quality review focusTransfer rate and abandonmentCompletion accuracy and handoff quality
Cost modelCost per call avoidedCost per resolution delivered

Neither model is wrong in every case. Containment is appropriate for simple informational queries where there is no task to complete. But for transactional workflows where the caller expects something to happen, resolution is the more reliable measure.

The repeat-contact test

The simplest way to test which metric your organisation actually tracks is to ask: what happens after a contained call?

If your reporting does not match a contained call to a subsequent contact about the same issue, then your containment number includes an unknown proportion of unresolved tasks. The customer who calls back enters your count as a new contact, and the original interaction still counts as contained.

A resolution framework, by definition, holds the outcome open until the observation window closes. It counts the task once and follows it to a verified end. The human-handoff guide explains how to define the transfer criteria that keep unresolvable tasks visible instead of burying them in a containment number.

Choosing the right metric for your pilot

If you are running a voice AI pilot, agree the primary metric before the first live call. The pilot playbook walks through the process of setting a baseline, choosing thresholds, and assigning evidence owners.

For a pilot scoped to a transactional intent such as a payment, appointment, or account change, verified task completion is the clearer measure. For a pilot scoped to information delivery with no downstream task, containment may be appropriate, but your team should still define what success looks like for the caller.

Use the call-intent worksheet to classify your candidate intents by task type before selecting the measurement approach.

If your team is evaluating voice AI that measures resolution rather than containment, discuss your pilot scope. Bring one candidate intent and the outcome records your team can access; confirm measurement, integration, and commercial requirements during that discussion.

Prepared with AI assistance and reviewed by Butter Labs.

Reviewed 22 September 2026