Judge a CRO agency by whether it creates a traceable sequence of trustworthy experiments, verified deployments, and cumulative improvement to a meaningful business outcome.

Do not use win rate, number of tests, or headline uplift as the primary verdict. Those metrics describe activity or isolated results. They do not prove that the customer experience changed, that the change remained live, or that the business became more valuable.

A defensible value chain looks like this:

Business problem -> measurement -> hypothesis -> experiment -> deployment -> verified impact -> next decision

If an agency cannot show that chain, its case studies may be persuasive marketing, but they are not enough to evaluate the program.

Why win rate is a weak agency KPI

Win rate is the percentage of experiments classified as winners. It can help diagnose an experimentation program, but it is easy to misread and easier to game.

An agency can raise its reported win rate by:

  • testing only safe, incremental ideas;
  • choosing sensitive proxy metrics instead of business outcomes;
  • changing the winner definition after results appear;
  • stopping tests opportunistically;
  • excluding inconclusive or losing experiments from the denominator; or
  • presenting only favorable clients and periods.

A lower win rate is not automatically virtuous either. Repeated weak hypotheses, poor implementation, or invalid analysis are not evidence of scientific courage.

The correct question is not “How many tests won?” It is:

Which changes produced reliable evidence, were deployed, and improved the outcome the business actually values?

Microsoft researchers have documented how incorrect metric interpretation can turn apparently positive experiment results into harmful business decisions. A result is only as useful as the metric definition, guardrails, data quality, and interpretation behind it. A Dirty Dozen: Twelve Common Metric Interpretation Pitfalls in Online Controlled Experiments

Test volume and isolated uplift are also incomplete

Number of tests measures throughput. Throughput matters because a slow program may not learn enough to justify its cost. But launching twenty weak tests is not more valuable than solving one important bottleneck.

Average winning lift describes selected winners. It ignores losses, inconclusive tests, undeployed changes, incompatible metrics, limited populations, and whether the effect remained relevant after rollout.

Site-wide conversion rate reflects the whole business. It can move because of traffic quality, seasonality, price, inventory, promotions, device mix, or other changes unrelated to the CRO program.

Deployed cumulative impact is closer to the program’s actual output. It tracks the combined relative effect of comparable changes that were validated, implemented, and verified against one consistent business outcome.

It is still not financial ROI. ROI requires incremental profit, eligible traffic, margin, refunds, duration, implementation cost, and agency cost.

The six-evidence test for a CRO agency

Evaluate evidence, not promises. For each criterion, ask for an artifact from real work.

CriterionWhat a strong agency can demonstrateEvidence to requestRed flag
1. Measurement integrityThe primary outcome, eligibility, attribution window, and guardrails are defined and the critical events reconcile with business systems.Tracking or measurement plan, QA report, event definitions, reconciliation sample.The agency accepts dashboard numbers without validating how they were produced.
2. Hypothesis qualityEach test connects observed evidence to a plausible mechanism and a decision the business needs to make.Two recent experiment briefs, including the evidence, mechanism, treatment, metric, and decision rule.The backlog is mainly copied best practices, redesign opinions, or isolated UI ideas.
3. Experiment trustworthinessAssignment, exposure, sample ratios, primary metrics, guardrails, stopping rules, and uncertainty are reviewed before a result is acted on.One complete winner report and one complete losing or inconclusive report with raw rates and quality checks.Reports show only platform screenshots, confidence badges, or a lift number.
4. Build and deployment ownershipThe agency can design, build, QA, launch, deploy, verify, and roll back changes, or defines a precise handoff with the client team.Responsibility matrix, launch checklist, deployment register, median deployment lag.“Winning” changes remain in presentations because nobody owns implementation.
5. Business-impact reportingThe agency distinguishes experiment evidence from realized impact and reports comparable deployed changes against a consistent outcome.Deployed-impact register, assumptions, qualified and all-deployed results, separate profit model where reliable.It adds lifts arithmetically, counts undeployed winners, ignores negative deployments, or labels uplift as ROI.
6. Learning retention and decision cadenceResults update the opportunity map, prevent repeated weak tests, and determine what happens next.Decision log, experimentation knowledge base, revised backlog, monthly or quarterly review.Every test is reported as an isolated campaign with no retained reasoning.

Do not average away a failure in the first, third, or fourth criterion. If the measurement is untrusted, the analysis is invalid, or nobody deploys the result, the program cannot produce defensible business impact.

A result does not create realized business impact until it is deployed

An experiment result can create evidence. A deployed and verified change alters the customer experience and can create realized business impact.

That distinction sounds obvious, but it changes agency reporting materially. A Microsoft case study of experimentation at Bing treated the process as a full cycle from code change through testing, iteration, and deployment to users, not as a report-generation step. Characterizing Experimentation in Continuous Deployment

A useful deployment register should record:

  • the approved primary-outcome lift;
  • whether the result was sufficiently reliable to act on;
  • relevant guardrail outcomes;
  • deployment status and date;
  • verification that the intended change is live;
  • rollback or expiry status; and
  • whether the deployed version became the baseline for the next experiment.

The 99ways CRO Compounded Uplift methodology and spreadsheet template provide a simple way to maintain this register for comparable sequential experiments.

The method keeps two outputs:

  • Qualified compounded uplift: deployed changes whose approved relative lift clears a chosen meaningful-effect threshold.
  • All-deployed compounded uplift: every deployed change entered, including smaller and negative effects.

The second result matters because it limits selective reporting. A program should not advertise substantial wins while silently excluding underperforming changes that were also shipped.

Use the register only when experiments share a comparable primary outcome, population, conversion window, and sequential baseline. It is a directional program-performance summary, not statistical proof or causal financial ROI.

What to request before hiring or renewing an agency

Ask for five artifacts:

  1. A measurement specification showing how the primary business outcome and guardrails are constructed.
  2. A complete experiment brief and report for a winner, including raw rates, uncertainty, and downstream effects.
  3. A complete report for a loss or inconclusive test showing how the agency interpreted it and changed the next decision.
  4. A deployment register showing which validated changes actually reached users and how deployment was verified.
  5. A cumulative-impact report that includes assumptions, all deployed effects, and a separate financial model only where the inputs support one.

These documents reveal more than a polished case-study deck. They show whether the agency operates a repeatable decision system.

The final decision rule

Choose the agency that can accept end-to-end accountability for evidence, not the agency with the largest isolated lift or the highest claimed win rate.

A strong partner should be able to say:

  • what was observed;
  • what was inferred;
  • what was tested;
  • what could invalidate the result;
  • what was deployed;
  • what business outcome changed;
  • what remains uncertain; and
  • what the evidence says to do next.

The objective is not to manufacture winners. It is to improve the business through a sequence of trustworthy decisions and verified changes.

Frequently asked questions

What is a good win rate for a CRO agency?

There is no universal good win rate.

Win rate depends on how ambitious the hypotheses are, what counts as a winner, the traffic and detectable effect, the chosen metric, and whether all results are included. Use win rate to investigate the program’s testing mix and reporting behavior, not as the primary measure of value.

How many A/B tests should a CRO agency run?

Enough to maintain a useful learning cadence without outrunning traffic, implementation capacity, or decision quality.

Test count measures throughput. It should be considered alongside hypothesis quality, cycle time, deployment rate, and cumulative deployed impact.

Should losing experiments count as agency value?

A losing experiment does not contribute positive deployed impact if the losing change was not shipped.

It can still create learning when the experiment was trustworthy, the hypothesis was explicit, and the result changes a future decision. A poorly designed loss is not automatically valuable.

How should a CRO agency report results?

Each report should state the hypothesis, eligible population, primary outcome, control and treatment rates, relative and absolute change, uncertainty, guardrails, data-quality checks, segment findings, decision, and deployment status.

Program reporting should then separate operational performance, experiment quality, and deployed business impact.

Is compounded uplift the same as CRO ROI?

No.

Compounded uplift measures the combined relative effect of comparable deployed changes. CRO ROI compares incremental profit attributable to the program with total program cost. ROI requires additional financial assumptions and should be modeled separately.

What if the client owns implementation?

The agency should still maintain deployment status and verify that the represented change went live correctly.

The contract should define the handoff owner, expected implementation time, QA responsibility, and what happens when a validated winner is delayed, modified, or never deployed.

Evaluate the complete system

If you are comparing CRO partners, ask each one to show the chain from measurement to deployment and cumulative impact.

See how 99ways approaches end-to-end CRO and experimentation. The relevant question is not how many tests an agency can launch. It is whether the work produces decisions and deployed improvements you can audit.

Author

  • Iman Nazari

    Iman combines user psychology, business strategy, and experimentation to uncover what drives action and improves performance. He focuses on hypothesis development, evidence-based decision making, and turning insights into changes that can be confidently tested and scaled.

Related blog posts

Leave a Reply

Your email address will not be published. Required fields are marked *