Birchy
Outsourcing & BPO
Research
How many customer service outsourcing providers publish a control over their AI?
Five of the 32 customer service outsourcing providers assessed in this research publish any stated control over what their AI produces: TELUS Digital, Simply Contact, HGS, TaskUs and Conectys. Three of the 32 — Covisian, iSON Xperiences and Mplus — could not be reached in this pass and are counted with the silent, so the finding rests on the 29 whose published materials were searched. A stated control means a published description of how AI output is checked by people — review sampling, escalation triggers, an honest containment rate, or a position on which decisions stay human. Evidence was collected between 12 and 31 August 2026 from each provider's own published materials.
This counts what providers publish, not what they do. A provider absent from the five may operate careful oversight and describe none of it; this research reads public sources and can only report what was found published, in the sources searched, on the dates searched. It is not a census of any provider's practice.
These figures are not part of any score. The AI evidence behind this page sits in the evidence ledger and is reflected in no ranking on this site. The scoring criterion that touches AI was applied on deployment evidence, which was collected systematically, while the control evidence this page reports was not — so the counts here describe what is published and do not move anyone's position.
What counts as a stated control
The definition is the published scoring anchor's, and the finding turns entirely on it. The anchor requires deployed AI described concretely — which workflows, which tools — and a stated human-in-the-loop quality control design. Its own examples of that design are:
- review sampling of AI outputs — what proportion a person checks
- escalation triggers — what sends a case to a human, and when
- containment versus deflection stated honestly — resolved, as against merely not reaching a person
- a stated position on what stays human — which decisions are never automated
The distinction that decides most cases is the direction the sampling runs. Several providers publish AI that samples agent interactions — the AI checking the humans. That is not the anchor: the anchor is humans checking AI outputs. An assist architecture is a delivery model, not a control, and the anchor says so itself, listing "agent assist" as an example of a tool merely mentioned. A quality-assurance system that increases sampling volumes by automating the sampling is evidence of deployment, not of oversight.
Publishing an oversight design is one disclosed fact. It is not a quality verdict, and this page does not rank the five above anyone: they sit across the whole range of this research's scores, from third overall to seventeenth.
What each of them publishes
| Provider | What is published |
|---|---|
| TELUS Digital | Confidence thresholds route only what AI cannot resolve to humans rather than maximising deflection, with "a person, not a system, stays accountable for the moments that matter" |
| HGS | Human-in-the-loop oversight stated at every decision point carrying regulatory, financial or clinical consequence, with runbooks, controls and rollback paths mapped to GDPR, HIPAA and the EU AI Act |
| Conectys | "Human-in-the-loop oversight, complex decision making, high-empathy interactions" set out as the human half of a stated division of labour against a machine half |
| TaskUs | A Global AI Policy dated 3 June 2026: an AI Advisory Board, a human-in-the-loop model proportional to risk and complexity, client data excluded from internal AI initiatives, and mandatory PII redaction before AI processing |
| Simply Contact | AI tools are tested before implementation, with agents evaluating how the AI handles requests and identifying errors — a review of AI outputs by people, published in a named client case |
The partial case, which is the more useful one
Probe Group publishes a Responsible AI Policy v1.0 dated 1 July 2024, establishing a cross-functional AI Risk and Ethics Committee that meets quarterly and mandating "effective human oversight of AI systems with clear accountability across the AI lifecycle".
It does not meet the anchor, and what it omits is the instructive part: the policy delineates no decisions as human-only and specifies no escalation procedure. A commitment to oversight without a stated trigger and without a boundary is a governance intention rather than an operating control, and the gap between those two is precisely what a buyer is trying to establish. It is published, dated and versioned, which is more than most of this field offers.
One further hedge worth preserving, because it is the substance rather than the wording: Conectys states alignment with "the principles of the EU AI Act", not compliance with it.
What was searched, and what was not
A negative finding is only worth quoting if the method is stated.
AI material was collected for 29 of the 32 providers: each provider's own AI or technology pages, its published policies, and its client cases, read on dated visits. Where a provider named products but described no oversight, that is recorded as evidence read and no control found.
three providers were not reached in this pass — Covisian, iSON Xperiences and Mplus. No AI material was collected for them at all. That silence is ours, not theirs, and they are named here rather than counted as publishing nothing.
The count above is therefore five of 32 assessed, drawn from 29 searched. The next edition collects the control half as a category for every provider, so the same question can be asked of a complete set.
What a buyer should ask
Every question below is the scoring anchor turned around. They cost nothing to ask and a provider that has the design will answer them quickly.
- What proportion of AI outputs does a person review, and how is that sample selected? A stated rate is the difference between oversight and the intention of it.
- What triggers an escalation to a human, and who set the threshold? Ask for the condition, not the reassurance.
- What is your containment rate as against your deflection rate? Containment is resolved. Deflection is only that the customer did not reach a person. A provider that reports one number for both is reporting the wrong one.
- Which decisions stay human, always? The answer should be a list, not a principle.
- Who signs off a model change, and what happens to it if the change degrades quality? Rollback paths are cheap to describe and rarely described.
If a provider cannot answer these, that is not proof of a problem. It is the same finding this page reports, arriving in a room where you can ask a follow-up question.
Check this, and correct it
The evidence behind every line above is on each provider's data sheet, with its source, its evidence grade and the date it was read. Both scoring tables are published as CSV on the data and downloads page, under CC BY 4.0 — cite the edition year with any figure, since these are edition-specific.
If you are a provider named here and publish a control this research did not find, send it with a public URL. Findings are corrected on evidence, and a correction that moves a count moves it on this page and in the data at the same time.
Disclosure
Birchy operates commercially in the outsourcing sector and may have worked with, or may in future work with, any of the providers evaluated here. That has no bearing on the evaluation: every provider is scored against the same published anchors, from evidence documented in its public data sheet, and there is no scoring input that cannot be independently re-derived and challenged.