Birchy
Outsourcing & BPO
Research
CALIBRATION ADDENDUM — how "evidenced" is read
BPEF v1.2-c · The framework's published reading of “evidenced”
The anchors say a capability must be "evidenced" without defining what satisfies it. This framework resolves that question as follows, with the worked examples below. They are the calibration, and they are written out so that the same question resolves the same way in every edition.
The evidence GRADES and CAPS are unchanged and still bind. A 5 still requires E1 or E2. Entirely-E3 evidence still caps at 3. Absence of evidence still scores 0 or 1. What follows governs only what counts as evidence of a capability — not how trustworthy a source is.
The general rule
A published, specific service description from the provider is evidence that the capability exists. It is not evidence that the capability is good, and it does not lift a score past 3 on its own. Corroboration — a named client, an analyst evaluation, a review, a certification whose scope reaches the work — is what carries a sub-criterion to 4 or 5.
The strict reading treats a described-but-uncorroborated capability as absent and scores it 1. This framework treats it as present but unproven and scores it 3. Score it 3.
What this does not license: inferring a capability nobody claims, crediting a capability the provider's own materials contradict, or reading scale as capability. Those still score as the anchors direct.
Sub-criterion rules, with the worked example each rests on
B1.2 · Technical support depth
- 3 — a technical support line is published with any concrete operational detail: tiering, an escalation model, or knowledge-base practice. Escalation SLAs are not required.
- 4 — that, plus a technical client evidenced by name, review, or analyst assessment.
- 5 — that, plus named tooling and a documented escalation model.
Worked example: Foundever and Konecta both scored 4 on a published tiering description with client evidence and no published SLA. Conectys, Mplus and Armatis scored 3 on the description alone.
B2.3 · Coverage model
- 3 — 24/7 or extended coverage published as an offering, with a verified multi-hub footprint capable of supporting it.
- 4 — that, plus 24/7 evidenced for at least one client, or a footprint spanning EU+US or wider.
- 5 — 24/7 across a verified multi-region footprint offered as standard.
- Shift-handoff practice is not required at any level. Its absence caps nothing.
Worked example: Teleperformance, Foundever, Konecta and Transcom all scored 5 without publishing shift-handoff practice. Armatis scored 3. No provider in the field scored below 3.
B3.1 · QA methodology
- 3 — a QA method described in operational terms, or an external validation (COPC, ISO 18295) held.
- 4 — either one at depth: certification held, or methodology described with sampling, rubric or calibration named.
- 5 — described methodology and external validation.
Worked example: Armatis scored 4 on ISO 18295-1 certification alone. Teleperformance and Foundever scored 4. Konecta, Transcom, Conectys and Mplus scored 3 on description alone.
B3.3 · Workforce management
- 3 — forecasting or scheduling practice described, or volume variability handled with numbers. Named WFM tooling is not required.
- 4 — a ramp or seasonal flex evidenced with numbers.
- 5 — multiple evidenced ramps, or ramp plus seasonal elasticity.
Worked example: Teleperformance, Foundever, Konecta and Armatis scored 4 without naming WFM tooling.
B4.1 · Access control
- 3 — concrete controls named (any of: role-based access, MFA, VDI, device management, clean-desk, encryption, monitoring).
- 4 — those controls sitting inside a certification the provider holds, even where the certificate's site-level scope is not published.
- 5 — controls plus published scope covering the delivery hubs, or client-audit evidence.
Worked example: Teleperformance and Foundever scored 4 with scope-per-entity explicitly recorded as unverified.
B4.3 · Controls on sensitive workflows
- 3 — the provider serves regulated or sensitive work and holds certifications that reach it. A published maker-checker or four-eyes mechanism is not required at this level.
- 4 — that, plus a regulated engagement evidenced by name, review, or analyst assessment.
- 5 — multiple regulated engagements evidenced, with the client-retained authority boundary stated.
Worked example: Foundever, Konecta, Transcom and Conectys all scored 4 without publishing a control mechanism. No provider serving regulated work scored below 3.
B7.2 · Regulated-workflow capability
- 3 — regulated work evidenced, under certifications that cover it.
- 4 — regulated work across two or more verticals, evidenced.
- 5 — an externally recognised position in a regulated vertical.
Worked example: field floor of 3 for every provider serving regulated work; Foundever and Konecta scored 4; Teleperformance 5.
Who authored it, not who hosts it
Two sources corroborate each other only if two different parties wrote them. The test is authorship, not the platform the text sits on, and the distinction decides several grades.
A G2 or Clutch seller profile is written by the provider: its description, its minimum project size, its rate band. A provider's own site and its own seller profile are one author twice, and grading that pair E2 asserts an agreement that does not exist. The same applies to a company's second document — a key-figures page against its own CSR report is one author, not two.
A review on those same platforms is written by a customer. A rating is computed by the platform from those reviews. A LinkedIn headcount range is derived from individual employees' own profiles, and hiring volume from posted activity. None of these is a company assertion, even though the company controls the page they appear beside. That is why A2.2 accepts "site + LinkedIn headcount range or hiring volume" as corroboration while E2 would reject "site + seller profile": the anchor points at figures the provider does not author.
It is the same distinction the live-posting convention rests on. A marketing page is the provider asserting a capability; a posting, a headcount range and a customer review are all artifacts the provider did not write, or did not write for that purpose.
A4.1 · Employee-side signal
The operative rule sits in the anchors appendix and is summarised here because it decides which number is read before any anchor is applied. Score the employer profile with the largest review volume; treat a difference of 0.5 or more between sources as material, that being the width of one rating band and so the finest distinction this criterion resolves; and let a material divergence cost one interpolated anchor step rather than a whole band.
The rule is needed because an employer rating is not a single well-defined number. A platform may carry several unmerged entities for one company, two platforms may disagree, and both platforms offer city and country filters whose figures look identical in form to the company-wide one. None of these is resolved by knowing the corporate history: Glassdoor merged Webhelp into Concentrix and Sitel into Foundever, and carries Konecta, Mplus and Transcom as separate unmerged entities. The profiles are therefore counted, not assumed.
Worked example: Transcom's employer footprint resolves to three unmerged Glassdoor entities — 2,072, 217 and 131 reviews. The largest governs, giving 2.9 and anchor 0. Mplus resolves to three entities of 22, 17 and 5 reviews, where the largest is also the lowest; the rule is indifferent to which direction the sample points.
Live job postings as evidence
Eleven claims across both editions rest on live job postings, spanning eight providers. Because those claims are graded E2 rather than E3, the reasoning is set out here rather than left as an unstated exception to the evidence scale.
E1 and E2 describe evidence generated by someone other than the provider, and a provider writes its own job adverts. The exception is deliberate and rests on a distinction the scale does not otherwise draw: a marketing page is the provider asserting a capability; a job posting is the provider acting on one. Reading a posting is not accepting the provider's account of itself — it is observing an operational artifact produced for a different purpose and drawing an inference from it. It is also independently checkable by any reader for as long as it is live.
What qualifies. A live posting, on the provider's own careers portal or on a third-party job platform, naming the site and the language or skill at issue. A posting that names neither does not support a footprint claim.
What it evidences. That the provider is staffing that capability at that location. It does not evidence that a marketed total is correct, and the method is as valuable for what it fails to corroborate as for what it confirms: where a provider markets thirty languages and postings evidence nine, both numbers are reported and the gap is the finding.
The limitation. Postings expire. The read date on the source line is what makes such a claim checkable afterwards, and a reader re-running the search later may find the posting gone without the claim having been wrong when it was made. This is the one evidence type on which the read date is load-bearing rather than administrative.
A4.2 · Hub and business continuity
- 3 — two or more verified hubs plus a documented continuity approach.
- 5 — multi-hub delivery with continuity provisions, where the footprint itself demonstrates redundancy. A named disruption is required only where continuity is the provider's specific claim to distinction.
Worked example: Teleperformance, Foundever, Konecta, Transcom, Conectys, Mplus and Armatis all scored 5. Simply Contact's 5 rests on a named disruption; the others' do not.
Unchanged — scored strictly in both readings
These are scored strictly. Do not loosen them.
- A3.1 / A3.2 / A3.4 — certificate identifiers, issuing bodies and validity dates. Anchor 3 is the field ceiling where identifiers are unpublished; 4 requires two verified standards. No provider in the European field reached 5.
- A4.1 — employer rating bands applied literally, including where the band is unflattering. Konecta and Conectys scored 1; Transcom scored 0 on a company-wide rating below 3.0.
- B4.2 — public DPA, subprocessor list, residency options. Field range 1–3.
- B5.1 / B5.3 — entry terms and commercial transparency. Teleperformance, Foundever, Konecta, Transcom, Mplus and Armatis all scored 1 on B5.3.
- A1 — client-review evidence, with the analyst-equivalence provision for enterprise providers unchanged.
Why this direction
The choice is not between a strict rubric and a lenient one. It is between two readings of an underspecified word, one of which is published here and applied consistently across both fields. Adopting the other would move scores to gain nothing that a written definition does not also gain.
The cost is real and belongs on the methodology page: this framework scores what a provider publishes about its capability, corroborated where corroboration exists. It does not audit delivery. A provider that describes a control model it does not operate will score 3 here and fail in procurement — which is why every data sheet carries its E5 gaps as due-diligence questions, and why the limitations section says private delivery data must be validated in procurement.
That was always the honest description of what this instrument does. It is now written down.