Birchy
Outsourcing & BPO
Research
METHODOLOGY APPENDIX — Scoring Anchors
BPEF · Birchy Provider Evaluation Framework · 2026 Edition, v1.0
This appendix is the rater's instrument. It is published in full so that any reader can see exactly how scores are produced, and so the framework can be re-run by anyone. Licensed CC BY 4.0.
0. Rater instructions
Scoring unit. Each of the 12 criteria decomposes into sub-criteria. Raters score each sub-criterion on a 0–5 scale by matching documented evidence to the written anchors below. Anchors are defined at 0, 1, 3, and 5; scores of 2 and 4 are awarded when evidence clearly exceeds the lower anchor but does not meet the higher one. Sub-criterion points = (anchor score ÷ 5) × sub-criterion weight, rounded to one decimal. Criterion score = sum of its sub-criteria.
Evidence rules.
- Score only what is in the provider's data sheet. If it isn't documented with a source URL and access date, it doesn't exist for scoring purposes.
- Evidence-grade caps: a sub-criterion score of 5 requires at least one E1 or E2 source. Evidence that is entirely E3 (company-stated, single source) caps the sub-criterion at 3, regardless of how impressive the claim. E4 (industry-compiled) evidence never scores an individual provider — it is context only. This cap is the mechanism that makes transparency materially worth points and makes unverifiable boasting worthless.
- Absence of evidence scores 0 or 1 as anchored — raters do not infer capability from company size, reputation, or vibes.
- Conflicting sources: higher-grade source wins; the conflict is noted in the data sheet.
- Collection window: evidence dated after the window closes is out of scope for this edition.
Single-rater scoring, reproducible by design. This edition is scored by one rater against the anchors below. The control on that is not a second opinion but reproducibility: every sub-criterion score traces to a documented, evidence-graded datapoint in the provider's data sheet, and the weighting arithmetic is published, so any reader can re-derive a total and challenge a specific anchor match on the evidence. Disagreements are resolved through the corrections process, in public, rather than in an internal reconciliation. This applies identically to every provider, including any with which the publisher has a professional relationship.
Rounding & ties. Total scores are reported to one decimal. Providers within 1.5 points are treated as tied for tiering narrative purposes even if the table shows an order.
AXIS A — DELIVERY PROOF (45 points)
A1 · Independent client evidence — 14 pts
A1.1 Verified review volume & recency — 5 pts
- 0 — No presence on independent review platforms (Clutch, G2, or equivalent).
- 1 — Platform profile exists but has fewer than 3 reviews, or no review in the past 36 months.
- 3 — 5–14 reviews with at least 3 posted in the past 24 months; aggregate rating ≥4.0.
- 5 — ≥15 reviews with ≥5 in the past 24 months; aggregate rating ≥4.5; review flow is continuous rather than clustered in one burst (clustering suggests a solicitation campaign and caps this sub-criterion at 4).
A1.2 Review substance — 5 pts
Scored from reading the reviews, not counting them.
- 0 — Reviews are anonymous, generic, or content-free ("great team, recommend").
- 1 — Some reviews name a reviewer role but describe no concrete engagement (scope, duration, team size, outcome all absent).
- 3 — Most recent reviews identify reviewer role and company type and describe concrete scope (channels, languages, team size, duration); at least 2 mention measurable outcomes.
- 5 — Reviews consistently include named/verified reviewers, concrete scope, measurable outcomes (CSAT movement, response-time change, cost effect), and at least one describes a problem and how the provider handled it. Critical detail in reviews is a positive signal, not a negative one.
A1.3 Named case studies with dated, quantified outcomes — 4 pts
- 0 — No case studies, or case studies with unnamed clients only ("a leading airline").
- 1 — Named-client logos exist but case content is undated or outcome-free.
- 3 — ≥2 case studies naming the client, dating the engagement, and quantifying at least one outcome.
- 5 — ≥4 named, dated, quantified case studies across ≥2 industries, at least one less than 24 months old, and at least one corroborated outside the provider's own site (client statement, press, award, review). E1/E2 corroboration is required for a 5 per the evidence cap.
A2 · Market presence & tenure — 8 pts
A2.1 Operating tenure — 2 pts
- 0 — <3 years (fails inclusion; recorded for the screened list).
- 1 — 3–5 years.
- 3 — 6–10 years.
- 5 — >10 years of continuous operation under substantially the same identity (renames/mergers documented, not disqualifying).
A2.2 Delivery scale band — 3 pts
- 0 — Scale unverifiable from any source.
- 1 — 100–299 FTE band, corroborated (site + LinkedIn headcount range or hiring volume).
- 3 — 300–1,999 FTE band, corroborated.
- 5 — ≥2,000 FTE band, corroborated. Note: scale is scored once, here, at 3 points of 100 — deliberately small. Scale advantages that matter to buyers (languages, coverage, flexibility) are scored where they show up in capability, so size is not double-counted.
A2.3 External recognition & growth signals — 3 pts
- 0 — No third-party recognition of any kind.
- 1 — Directory listings and purchased-tier awards only.
- 3 — Substantive third-party recognition: juried industry awards, top-tier directory rankings with stated methodology, or sustained hiring growth across ≥2 hubs.
- 5 — Coverage in a major analyst evaluation (Everest, ISG, NelsonHall, Gartner) at any tier, or equivalent independent market validation. Any analyst-report presence anchors this sub-criterion at 5 automatically.
A3 · Certifications & compliance — 12 pts
Global rule for A3: "certified" means the certificate is current and verifiable (registry entry, certificate number confirmed, or auditor attestation on record). Self-declared "compliant with" or "aligned to" language scores at most 40% of the certified anchor, and the profile prints the distinction.
A3.1 Information security certification — 5 pts
- 0 — No information-security certification and no stated program.
- 1 — Self-stated ISO 27001/SOC 2 "alignment," no certificate evidence.
- 3 — ISO 27001 certified or SOC 2 Type II attested, verified, covering the relevant delivery entities.
- 5 — ISO 27001 certified and SOC 2 Type II attested (or ISO 27001 + a second verified attestation), current, with scope covering the delivery locations serving clients.
A3.2 Payment & data-handling compliance — 3 pts
PCI DSS note: no "PCI certificate" formally exists. Valid evidence is a current Attestation of Compliance (AoC) — QSA-validated for Level 1 service providers, SAQ-based for Level 2 — and/or listing on the card networks' compliant service provider registries. "Certified" and "compliant" in marketing copy are treated identically: both score on the AoC/registry evidence behind them, and stating the service-provider level is itself a specificity signal.
- 0 — Handles payment or regulated data with no stated compliance posture.
- 1 — PCI DSS "compliant/certified/ready" claimed with no level stated and no AoC or registry evidence.
- 3 — Current AoC confirmed (level stated; via trust page, RFI, or registry listing) where payment flows are served; documented GDPR processor posture (DPA offered, subprocessor list exists).
- 5 — QSA-validated Level 1 AoC or registry listing verified + demonstrably mature GDPR posture (public DPA, named DPO, data-residency options) + any applicable extra regime (e.g., HIPAA BAA capability) evidenced. If the provider serves no payment flows, score on the GDPR/data-handling half and note N/A on PCI; do not award free points.
A3.3 Operations standards — 2 pts
- 0 — No operations-standard certification or framework reference.
- 1 — References COPC or ISO 18295 concepts in marketing without certification.
- 3 — ISO 18295-1 certified, or COPC-trained staff/registered coordinators documented.
- 5 — COPC CX Standard certification (site-level) or ISO 18295-1 certification, current and verified.
A3.4 Compliance transparency — 2 pts
This sub-criterion scores verifiable identifiers, not published certificate documents. Standard practice — and the 5-anchor — is a trust page with standard, certificate/attestation number, issuing body, validity dates, and scope, with audit reports (SOC 2 report, PCI AoC) available under NDA. ISO certification bodies maintain public registries, so number + body = independently checkable.
- 0 — Certifications claimed with no identifying detail; requests for identifiers unanswered.
- 1 — Standards name-dropped or badge-displayed across the site without number, body, version, or scope anywhere.
- 3 — Certificate identifiers, issuing body, or scope provided (public page or RFI); versions current.
- 5 — Trust/security page listing each standard with number, issuing body, validity dates, and scope; NDA availability stated for underlying reports; renewals reflected.
A4 · Delivery stability — 6 pts
A4.1 Employee-side signal — 3 pts
Proxy for attrition and management quality. Read distribution and recency, not just the average; ≥25 ratings required for full confidence — below that, cap at 3 and note thin data.
- 0 — Employer rating <3.0, or credible public evidence of systemic workforce problems (wage disputes, mass-departure reports).
- 1 — Rating 3.0–3.4, or ratings healthy but reviews describe chronic churn on client accounts.
- 3 — Rating 3.5–3.9 with recent reviews broadly consistent; no red-flag patterns.
- 5 — Rating ≥4.0 across ≥25 reviews, recent reviews included, with management-response activity — and no contradiction between employer reviews and client reviews on team stability.
A4.2 Hub continuity & business continuity — 3 pts
- 0 — Delivery locations unstable or unverifiable; hubs claimed but no local hiring or address evidence.
- 1 — Hubs verified but single-site concentration with no continuity provisions stated.
- 3 — ≥2 verified delivery hubs; documented BCP approach (site redundancy, remote-shift capability, client-facing continuity commitments).
- 5 — Multi-hub delivery with demonstrated continuity through a real disruption (documented client-service continuity during a named event — outage, natural disaster, wartime operation), corroborated beyond the provider's own claim. This anchor was written with full awareness that Ukraine-based providers have the strongest real-world evidence in the industry here; the anchor rewards the documented evidence, not the geography.
A5 · Transparency & verifiability — 5 pts
v1.2 note: this criterion scores public verifiability only. There is no provider questionnaire in this edition; the methodology is public-source, and this criterion rewards providers whose claims can be independently checked without asking them.
A5.1 Claim verifiability — 3 pts
- 0 — Headline claims (headcount, clients, languages, metrics) contradicted by observable evidence, or internally inconsistent across the provider's own materials (e.g., conflicting attrition figures).
- 1 — Claims plausible but nothing independently checkable.
- 3 — Most material claims checkable and consistent across ≥2 independent sources.
- 5 — Claims systematically verifiable: public materials cite their own evidence (named clients, dated figures, certificate identifiers), consistent across site, review platforms, and third-party listings.
A5.2 Organizational & commercial transparency — 2 pts
- 0 — Leadership unidentifiable; no imprint; contact routes dead.
- 1 — Generic team page; leadership names without verifiable profiles.
- 3 — Named leadership with active professional profiles; working contact channels; legal entity identifiable.
- 5 — All of the above plus published pricing model (structure, not necessarily rates) and clear entity/ownership structure.
AXIS B — CAPABILITY FIT (55 points)
B1 · Service scope & channels — 8 pts
B1.1 Channel coverage — 3 pts
- 0 — Single channel only.
- 1 — Voice + email; digital channels claimed without evidence of live delivery.
- 3 — Voice, email, chat, and at least one of social/in-app delivered, evidenced in case studies or reviews (not just a service-page list).
- 5 — Full omnichannel set delivered and evidenced, including at least one advanced channel (in-app, video, community/marketplace messaging) in production for a named or reviewed client.
B1.2 Technical support depth — 3 pts
- 0 — No technical support offering.
- 1 — "Technical support" listed; no tiering or evidence.
- 3 — L1/L2 tiering described concretely (escalation model, knowledge-base practice) with at least one supporting case or review.
- 5 — L1/L2 (±L3 coordination) with named tooling, documented escalation SLAs, and client evidence from a technical product (SaaS, app, device).
B1.3 Omnichannel orchestration & adjacent scope — 2 pts
- 0 — Channels operate as silos; no orchestration story.
- 1 — Orchestration claimed generically.
- 3 — Unified queue/CRM operation across channels described with named platforms; or evidenced back-office adjacency (order management, content moderation, KYC support) under the same QA framework.
- 5 — Both: platform-level orchestration evidence and support+back-office delivery under one documented quality framework.
B2 · Language & geographic reach — 10 pts
B2.1 Native-language delivery, verified — 4 pts
Verification standard: a language counts as verified native delivery when the language claim co-occurs with corroborating hiring data (live or recent postings for native/C1+ speakers in a delivery hub) or client evidence naming the language.
- 0 — Language list published; zero languages verifiable.
- 1 — 1–4 languages verified.
- 3 — 5–9 languages verified, including ≥3 non-English European languages.
- 5 — ≥10 languages verified, including high-demand/hard-to-staff coverage (e.g., German, French, Italian, Dutch, or Nordic languages) evidenced through hiring or client proof.
B2.2 Delivery footprint — 3 pts
- 0 — Single unverified location.
- 1 — Single verified hub.
- 3 — 2–3 verified hubs in ≥2 countries.
- 5 — ≥4 verified hubs across ≥3 countries, or verified hub + structured remote model covering additional geographies with stated controls.
B2.3 Coverage model — 3 pts
- 0 — Business-hours single-time-zone only, where the market expects more.
- 1 — Extended hours claimed; no operational detail.
- 3 — 24/7 or follow-the-sun delivered for at least one evidenced client, with shift-handoff practice described.
- 5 — 24/7 multi-hub follow-the-sun as a standard offering, evidenced across multiple clients/reviews, covering EU+US or broader time-zone spans.
B3 · Quality management system — 10 pts
B3.1 QA methodology — 4 pts
- 0 — No QA description beyond the word "quality."
- 1 — QA team exists; method unstated (no sampling, rubric, or calibration detail).
- 3 — Concrete methodology: stated sampling approach, scoring rubric, calibration practice, feedback loop to agents. (COPC/ISO 18295 vocabulary used correctly is supporting evidence, not a requirement.)
- 5 — Full methodology as above plus external validation: COPC/ISO 18295 certification, client reviews explicitly praising QA/reporting, or published QA artifacts (sample scorecards, calibration cadence).
B3.2 CX measurement practice — 3 pts
- 0 — No mention of CSAT/NPS/quality metrics.
- 1 — Metrics name-dropped; no practice description.
- 3 — Measurement practice described (survey mechanism, cadence, target-setting) and referenced in at least one case or review.
- 5 — Practice described and results published with client attribution (named-client CSAT/NPS outcomes), or metric outcomes corroborated in independent reviews.
B3.3 Workforce management maturity — 3 pts
- 0 — No forecasting/scheduling capability described.
- 1 — WFM claimed generically.
- 3 — Forecasting and scheduling practice described with tooling named; evidence of handling volume variability (seasonality, launches).
- 5 — Demonstrated WFM through evidenced seasonal/spike ramps (case with numbers: ramp size, timeline, service-level held), or WFM-specific client praise in reviews.
B4 · Security & data protection operations — 8 pts
A3 scores the certificates; B4 scores the operational reality.
B4.1 Access control & environment security — 3 pts
- 0 — No description of how agent access to client systems and data is controlled.
- 1 — Generic "secure environment" language.
- 3 — Concrete controls described: role-based access, least-privilege, clean-desk/clean-room options, VPN/VDI delivery, device management.
- 5 — Controls described and independently anchored (within a verified ISO 27001/SOC 2 scope covering those hubs, or client-audited per reviews/cases), including options for restricted environments (PCI zones, no-mobile floors) where relevant.
B4.2 Data residency & subprocessor transparency — 2 pts
- 0 — No data-handling information; DPA unavailable.
- 1 — GDPR mentioned; no residency or subprocessor detail.
- 3 — Data-residency options stated (EU processing available); DPA offered; subprocessor disclosure on request.
- 5 — Public DPA, published subprocessor list, EU residency guaranteed on request, named DPO/contact.
B4.3 Controls on sensitive workflows — 3 pts
- 0 — Serves regulated/sensitive workflows with no stated control model.
- 1 — Controls claimed without mechanism.
- 3 — Mechanisms described: maker-checker/four-eyes on sensitive steps, segregation of duties, audit trails, defined client-retained decision authority.
- 5 — Mechanisms described and evidenced in a regulated engagement (named or reviewed client in fintech, health, insurance, aviation security context), with the client-retained authority boundary stated explicitly. If the provider does not serve sensitive workflows and does not claim to, score 3 as neutral-N/A and note it.
B5 · Flexibility & commercial model — 7 pts
B5.1 Entry model & minimums — 3 pts
- 0 — Entry terms opaque; engagement only via enterprise sales process.
- 1 — Minimums exist and are high relative to the mid-market (>20 FTE or >12-month lock-in), or unstated.
- 3 — Pilot or small-team entry available (≤10 FTE or ≤3-month initial term), stated publicly or in RFI.
- 5 — Structured pilot offering with defined success criteria described, low entry minimums, and evidence of pilots converting (case or review references a pilot-to-scale path).
B5.2 Ramp & elasticity — 2 pts
- 0 — No ramp evidence.
- 1 — "Fast ramp" claimed without numbers.
- 3 — Ramp evidenced with numbers once (case: N agents in W weeks), or seasonal flex documented.
- 5 — Multiple evidenced ramps or documented seasonal elasticity across clients (scale up and down), with training-time transparency.
B5.3 Commercial transparency — 2 pts
- 0 — Pricing model undisclosed even in structure; exit terms unknown.
- 1 — Pricing model named (per-FTE/per-hour/per-contact) but nothing else.
- 3 — Pricing model + what's included/excluded explained (public or RFI); notice/exit structure stated.
- 5 — Model, inclusions, indicative rate bands or rate logic, and exit/knowledge-transfer terms all on the record. (Publishing actual rate cards is not required for a 5 — almost no one does — but rate logic is.)
B6 · Technology & AI operations — 7 pts
B6.1 Platform capability — 2 pts
- 0 — Tooling unstated; implied manual operation.
- 1 — Generic tool logos.
- 3 — Named platform stack (helpdesk, telephony, QA, WFM tools) with evidence of operating in client tools as an option.
- 5 — Stack named + integration/reporting capability evidenced (dashboards, API-level reporting, client-tool operation in cases/reviews).
B6.2 AI-human hybrid operations — 3 pts
2026 market context: analyst coverage now treats AI-enabled delivery as the category's main axis of change. This sub-criterion scores operational reality, not slideware.
- 0 — No AI position at all, or pure buzzword usage with no mechanism.
- 1 — AI tools mentioned (agent assist, bots) without deployment evidence or quality-control design.
- 3 — Deployed AI in production described concretely (which workflows, which tools) with human-in-the-loop quality control design stated (review sampling of AI outputs, escalation triggers, containment vs. deflection honesty).
- 5 — Production AI-human operations evidenced with outcomes (containment/AHT/CSAT effects with numbers, client-attributed or reviewed), plus a stated position on what stays human. Vendors reporting AI limits honestly score higher here than vendors claiming AI does everything.
B6.3 Roadmap credibility — 2 pts
- 0 — No forward view.
- 1 — Trend-following statements ("we embrace AI").
- 3 — Specific, dated initiatives (named pilots, partnerships, hires) consistent with observable activity (job postings, releases).
- 5 — Roadmap with shipped evidence over the trailing 12 months — announced initiatives that verifiably became production capability.
B7 · Vertical expertise — 5 pts
B7.1 Named vertical depth — 3 pts
- 0 — Claims "all industries"; evidence in none.
- 1 — Verticals listed; no named evidence in any.
- 3 — ≥1 vertical with real depth: multiple named clients or substantive cases, vertical-specific process knowledge visible (e.g., IATA/amadeus-adjacent workflows for travel, chargeback flows for e-commerce).
- 5 — ≥2 verticals at that depth, at least one with ≥3 named clients or an externally recognized position in the vertical (award, analyst note, client testimony).
B7.2 Regulated-workflow capability — 2 pts
- 0 — Claims regulated work (KYC, claims, health admin) with zero supporting evidence.
- 1 — Regulated capability listed; controls and evidence absent.
- 3 — Regulated workflows evidenced with the control model from B4.3 applied and the client-retained authority boundary stated.
- 5 — Multiple regulated engagements evidenced across compliance-relevant certifications (A3) that actually cover them. If the provider does not claim regulated work, score 3 as neutral-N/A and note it.
Appendix notes
N/A handling. Two sub-criteria (B4.3, B7.2) carry neutral-N/A at 3 for providers that legitimately do not operate in scope. All others score on evidence as anchored — a provider that doesn't do something scores what the anchors say, because the framework measures fit for the category, not effort.
Anti-gaming provisions. Review-burst clustering caps A1.1 at 4; E3-only evidence caps any sub-criterion at 3; scale is isolated in A2.2 at 3 points to prevent size from leaking into quality criteria; A5 makes stonewalling expensive and B5.3 makes opacity expensive. If a provider disputes an anchor as biased, the corrections process accepts arguments against anchors for the next edition (logged), but never mid-edition score negotiations.
Known bias, restated. The evidence-grade caps mean transparent providers outscore equally capable secretive ones. This is disclosed on the methodology page as a deliberate design choice: in a market where buyers must verify everything anyway, willingness to be verified is part of quality.
Version. v1.0, 26 August 2026. Changes to any anchor require a version increment and a changelog entry; anchors are frozen for the duration of an edition's scoring.
Criterion index
| Code | Criterion | Axis | Points |
|---|---|---|---|
| A1 | Independent client evidence | A · Delivery Proof | 14 |
| A2 | Market presence & tenure | A · Delivery Proof | 8 |
| A3 | Certifications & compliance | A · Delivery Proof | 12 |
| A4 | Delivery stability | A · Delivery Proof | 6 |
| A5 | Transparency & verifiability | A · Delivery Proof | 5 |
| B1 | Service scope & channels | B · Capability Fit | 8 |
| B2 | Language & geographic reach | B · Capability Fit | 10 |
| B3 | Quality management system | B · Capability Fit | 10 |
| B4 | Security & data protection operations | B · Capability Fit | 8 |
| B5 | Flexibility & commercial model | B · Capability Fit | 7 |
| B6 | Technology & AI operations | B · Capability Fit | 7 |
| B7 | Vertical expertise | B · Capability Fit | 5 |