Contact-centre AI metrics: an outcome scorecard
Measure a voice AI pilot against a verified customer outcome, then check repeat contact, handoff quality, and the total cost of completing the work. Agree the definitions before the first live call so operations and finance can judge the same result. For enterprise teams evaluating Butter Labs for selected high-volume telecom or utilities workflows, this scorecard provides a proposed pilot method; confirm product capabilities and data access during scoping.
Define completion before you count it
Choose the system event that proves the requested task happened. For a payment flow, specify whether your completion event means authorisation or settlement. For an appointment flow, require the scheduling system to record the agreed change. Give each eligible task a stable identifier so retries don't inflate the completion count. Keep unverified outcomes in an unresolved category until your team reconciles them. Record exclusions and their counts beside the result so a change in eligibility remains visible. Agree how you'll handle calls with several requests before comparing versions.
A finished conversation can still leave the customer's task unresolved.
A proposed metric dictionary
Use these definitions as a starting point for your pilot agreement. Your analytics owner should record each data source and its limitations, including missing events and the period needed to observe repeat contact.
| Measure | Proposed calculation or record | Evidence and owner |
|---|---|---|
| Verified task completion | Unique eligible tasks with the agreed completion event divided by all eligible tasks started in the cohort | Business-system event matched to task ID; operations and analytics |
| Repeat contact after recorded completion | Completed tasks followed by a matched contact about the same issue within the agreed window, divided by completed tasks with a full observation window | Cross-channel contact records and matching rules; analytics |
| Transfer rate | Eligible calls transferred to a person divided by eligible calls started; split customer-requested and system-required transfers | Telephony events and transfer reason; operations |
| Handoff quality | Audited transfers that meet every agreed handoff criterion divided by audited transfers | Receiving-agent review of context, reason, and next action; QA |
| Abandonment | Eligible calls ended by the caller before completion or successful transfer divided by eligible calls started | Disconnect events reconciled with task outcome; operations |
| Policy and action errors | Count confirmed wrong actions, inappropriate disclosures, and policy breaches by severity; report audited-call rates separately | Documented review sample and incident records; QA and risk owner |
| Tool failure | Failed tool requests divided by all tool requests, with retries identified separately | Request logs and outcome codes; engineering |
| Cost per resolution | Total cost attributable to the cohort divided by unique eligible tasks with verified resolution | Reconciled cost ledger and outcome records; finance |
Report counts beside rates and mark missing evidence as unknown. A zero denominator means the rate is unavailable.
The repeat-contact measure above tests whether a recorded completion holds up. Track repeat contacts after unresolved tasks separately so those customers remain visible. Agree when the observation window starts and report how many tasks still await a full window. If your team can’t match contacts across channels, state that coverage limit beside the rate.
Keep the scorecard comparable
Separate results by intent and workflow version, and keep the baseline on the same eligibility rules. Report the number of reviewed calls and how you selected them. Review successful calls as well as failed calls because a success label can conceal a wrong action. Where lawful and approved, examine relevant caller cohorts and call conditions, and describe gaps in coverage. Your risk owner should set pause criteria for serious incidents before live traffic starts.
Use the call-intent worksheet to define the cohort and the human-handoff guide to agree transfer criteria.
Reconcile the economics
Include platform and telephony charges, implementation effort, support, quality review, and human handling after transfer in the costs attributable to the cohort. State the accounting period and how finance allocates shared costs. Compare that total with the baseline cost for the same work. Track released capacity separately and describe a cash saving only when finance can identify the cost removed. If the cohort has no verified resolutions, show its total cost and leave cost per resolution unavailable. The cost-per-resolution guide explains how to document the assumptions.
Agree the next decision
Use the pilot playbook to record the baseline, threshold, evidence owner, and response for each measure. During early live testing, agree a review frequency that matches the volume and risk; reduce it only after the owners have enough evidence to do so. Record a continue, pause, or expand decision with the unresolved questions and the person responsible for each one.
If you're evaluating Butter Labs for a measured voice AI pilot alongside your existing contact-centre stack, discuss your pilot scope. Bring one candidate intent and the outcome records your team can access; confirm integration, security, and commercial requirements during that discussion.
Reviewed 18 September 2026