Choose ongoing A/B testing when your business has trustworthy measurement, enough eligible outcomes, and the capacity to run, deploy, and learn from a sequence of experiments.
Choose a one-time CRO redesign when the current experience is structurally unfit or a substantial change is unavoidable. Keep the old experience as the control, test the new system against it, and require the redesign to prove itself.
Sometimes the right answer is neither. If you cannot trust the data or identify the actual bottleneck, start with diagnosis and measurement.
The main difference is not project length. It is what the work can teach you.
A bundled redesign can tell you whether the complete alternative beat the current experience. An ongoing experimentation program can also improve the organization’s model of why customers respond—and use that knowledge to make the next test better.
There are three valid starting points
The usual choice is presented as one-time CRO versus ongoing CRO. In practice, there are three paths.
| Starting condition | Best first step | Primary output |
|---|---|---|
| Measurement, bottlenecks, or opportunity size are unclear | CRO and tracking audit | Trusted baseline and ranked hypotheses |
| The experience is structurally unfit or a major rebuild is unavoidable | One-time redesign tested against the current experience | Ship-or-reject decision about the bundle |
| Data, traffic, implementation, and decision cadence support repeated tests | Ongoing experimentation program | Deployed impact plus cumulative learning |
Do not run experiments simply because an A/B testing tool is available. Do not commission a redesign simply because the current site looks old.
The starting point should match the uncertainty.
Diagnose first when the uncertainty is what is broken.
If the uncertainty is whether a coherent replacement is better, test the replacement.
If the uncertainty is which mechanisms drive the outcome and how to improve them over time, build an ongoing program.
What a bundled redesign can prove
A redesign test compares two systems:
- Control: the current experience.
- Treatment: the new experience containing several coordinated changes.
If the treatment produces a trustworthy improvement in the primary business outcome, without unacceptable guardrail regressions, the business has evidence to ship the new system.
That is a valid and useful experiment.
The result estimates the effect of the whole treatment as experienced. It does not estimate the effect of each component. This is the same discipline used in 99ways’ broader Growth Lever Analytics framework: define the controllable mechanism, the intervention, and the decision the evidence must support.
Suppose a redesign changes:
- navigation;
- visual hierarchy;
- headline and offer framing;
- product education;
- social proof;
- pricing presentation; and
- the checkout path.
If it wins, some changes may have helped, some may have been neutral, and some may have hurt while being outweighed by stronger improvements.
If it loses, the reverse may be true.
The test answers:
Should we replace the current experience with this complete alternative for this population, period, and outcome?
It does not answer:
Which of the seven changes caused the result?
A bundled test is appropriate when the unit of decision is the bundle. It is insufficient when the next decision requires component-level understanding.
When a one-time redesign is rational
A substantial one-time treatment is reasonable when at least one of these conditions applies.
The current experience is structurally unfit
Incremental changes cannot repair an incoherent information architecture, broken mobile experience, obsolete checkout, severe performance problem, or funnel designed for a previous business model.
A large change is unavoidable
A rebrand, platform migration, new offer, regulatory requirement, pricing model, product architecture, or market shift may make the current experience impossible to preserve.
The mechanisms are tightly coupled
Some treatments need to be experienced as one system.
A premium repositioning may require coordinated changes to offer framing, design, proof, pricing, and qualification. Testing each component in isolation could create unnatural intermediate experiences and miss interactions.
The business needs one binary decision
Sometimes the immediate question is simply whether to replace version A with version B.
The organization may accept lower learning resolution because the shipping decision itself is valuable.
The company cannot yet support an ongoing program
Traffic, engineering capacity, or operational maturity may permit one meaningful old-versus-new test but not a continuous testing cadence.
That does not make the redesign automatically correct. It makes the one-time test a pragmatic evidence strategy.
How to test a redesign without turning it into a gamble
A redesign should still follow experimental discipline.
- Preserve the old experience as a valid control.
- Define the eligible population before launch.
- State one bundle-level hypothesis.
- Choose a primary business outcome, not only engagement.
- Define downstream and technical guardrails.
- Verify assignment, treatment delivery, exposure, and event tracking.
- Keep the treatment stable during the analysis period.
- Plan rollback before launch.
- Interpret the result at the bundle level.
- Record what the result changes about the next decision.
A useful bundle-level hypothesis might be:
For new paid-traffic visitors, replacing the fragmented product journey with one coherent value, proof, pricing, and checkout experience will increase refund-adjusted revenue per visitor because users can understand the offer and complete the purchase with less uncertainty.
That hypothesis permits a bold treatment. It also prevents the team from claiming that a particular testimonial, color, or headline caused the result.
If the redesign loses, do not enter redesign roulette
The most dangerous response to a failed redesign is another unrelated redesign.
The pattern looks like this:
- Replace the current experience with a broad alternative.
- The alternative loses or produces an unclear result.
- Explain the failure with a new story.
- Build another broad alternative around that story.
- Repeat without narrowing the uncertainty.
The team is running experiments, but it is not accumulating knowledge.
Each new treatment changes so many assumptions that the previous result cannot meaningfully constrain the next one. The organization is searching the design space without updating a model of the customer.
After a bundled loss, ask:
- Did the implementation and measurement work correctly?
- Which planned intermediate metrics moved?
- Where did the variants begin to diverge?
- Did important segments react differently for a pre-existing reason?
- Which mechanism was essential to the bundle hypothesis?
- Which components can be tested as one coherent follow-up?
- What did qualitative evidence reveal about the mechanism?
- Is the treatment worth revising, or should the team move to another ranked lever?
A loss should reduce uncertainty. If it only produces a new aesthetic opinion, the experiment was expensive but not informative.
Test one idea at a time; not necessarily one UI element
“Change one thing at a time” is useful advice when it means one causal proposition.
It becomes harmful when interpreted as one pixel, sentence, or component.
NIST’s design-of-experiments guidance notes that one-factor-at-a-time approaches fail when factors interact. In digital experiences, a headline may work only with the supporting proof; pricing clarity may require coordinated changes across several sections; a new onboarding model may need both sequence and copy changes. (NIST: One variable at a time)
The two dimensions are separate:
| Interpretable causal idea | Multiple unrelated ideas | |
|---|---|---|
| Small treatment | Change CTA wording to reduce ambiguity | Change CTA, badge, menu label, and image for unrelated reasons |
| Bold treatment | Rebuild the pricing narrative so total cost is clear before checkout | Redesign navigation, branding, pricing, proof, and checkout around several unconnected theories |
The best experiment is not always in the top-left cell. A bold, interpretable treatment can produce more useful movement than a trivial element test.
Use this rule:
Every material change in the treatment should serve the same mechanism or the same bundle-level shipping decision.
Microsoft’s Experimentation Platform described splitting a substantial infrastructure migration into separate experiments because it contained two distinct hypotheses. The first test exposed an unexpected latency mechanism; the team revised the design and reran it before moving to the next architectural question. The useful boundary was the hypothesis, not the number of files or services changed. (A/B Testing Infrastructure Changes at Microsoft ExP)
If the team wants to estimate separate component effects or interactions, a factorial or multivariate design may be appropriate when traffic and complexity permit it. NIST’s factorial-design guidance distinguishes designs intended to estimate main effects and interactions. For most CRO teams, however, a sequence of coherent tests is easier to operate and interpret than a large factorial design. (NIST: Fractional factorial designs)
Why ongoing experimentation can create more long-term value
A one-time redesign has one primary output: a decision about the redesigned experience.
An ongoing experimentation program can produce three forms of value:
Total program value = deployed business impact + avoided downside + better future decisions
This is a conceptual decomposition, not an accounting formula.
Deployed business impact
Validated changes are implemented and improve purchases, qualified leads, revenue per visitor, activation, retention, or another meaningful outcome.
Avoided downside
Losing tests prevent plausible but harmful ideas from being deployed broadly.
A prevented regression is real value even though it does not appear as uplift.
Better future decisions
Each interpretable experiment changes the team’s beliefs about the customer, product, mechanism, or metric. That learning improves hypothesis selection and treatment design.
Microsoft’s experimentation research describes controlled tests as a way to evaluate ideas scientifically and reports that results are often humbling enough to challenge expert prioritization. (Online Experimentation at Microsoft)
Ongoing experimentation therefore has option value: the next decision begins with more evidence than the previous one.
But that value is not automatic.
Every experiment should have two outputs
A mature experiment ends with two records.
Output 1: the product decision
- Deploy.
- Reject.
- Revise and retest.
- Investigate a defined uncertainty.
- Stop because the effect is too small to matter.
- Preserve the control while another lever takes priority.
Output 2: the customer-model update
- Which behavioral claim gained support?
- Which claim lost support?
- For which population, journey, and outcome?
- What alternative explanation remains?
- What should the next test preserve, remove, or isolate?
An experiment that produces a ship decision but no learning can still be valuable.
An ongoing program that repeatedly produces no learning is not compounding. It is merely processing tests.
Maintain a Living Customer Model
A static persona says what the customer is supposedly like.
A Living Customer Model records what the organization currently has evidence to believe about customer behavior—and preserves the boundaries and uncertainty of each claim.
Microsoft recommends archiving hypotheses, test metadata, metric movements, and ship decisions as a living history that can inform future ideas and prevent teams from repeating contradictory changes. (Patterns of Trustworthy Experimentation: Post-Experiment Stage)
The model should contain entries such as:
| Field | Example |
|---|---|
| Behavioral claim | Application friction appears to perform a qualification function in this funnel |
| Population and context | New webinar leads applying for a high-consideration sales call |
| Evidence | Easier form increased submissions but reduced downstream quality and purchase performance |
| Status | Supported in this funnel; not assumed universal |
| Competing explanation | Question wording may have changed the information collected, not only effort |
| Product implication | Optimize qualifying effort, not minimum effort |
| Next question | Can structured questions reduce effort while preserving qualification signal? |
| Last reviewed | Date and experiment IDs |
| Counterevidence | Future tests that support or contradict the claim |
Use evidence statuses such as:
- Tentative: one interpretable result with important uncertainty.
- Supported: credible evidence for a bounded population and context.
- Replicated: consistent evidence across more than one suitable test or method.
- Contradicted: material counterevidence requires revision.
- Retired: no longer relevant because the product, market, or funnel changed.
Do not write “customers hate X” after one test.
Write:
In this acquisition context, treatment X reduced the downstream outcome; the evidence is consistent with mechanism Y, but explanations Z remain.
The model should become more useful by becoming more precise, not more confident-sounding.
A public example: friction became part of the customer model
99ways tested an application form that replaced descriptive questions with easier multiple-choice questions.
Form submissions increased by 56.7%, which looked like a clear conversion win. Downstream lead quality, attendance, and purchase performance deteriorated. (Form Design That Sells)
The superficial conclusion would be:
Multiple-choice forms are bad.
The reusable learning was narrower:
In this high-consideration application funnel, effort and descriptive response may perform a qualification function. Optimizing form completion without downstream quality can damage the business outcome.
That claim improves the next decision.
The team can now test structured questions, scoring, conditional logic, or reduced effort while explicitly guarding qualification, attendance, and sales.
A losing treatment produced valuable learning because the experiment was interpretable and the outcome was measured far enough downstream.
An anonymized 99ways sequence: how learning changes the next treatment
One 99ways program began with a full-page alternative against the existing experience. The result was inconclusive.
The next test did not commission another unrelated page. It examined one mechanism: whether the alternative direction needed more opportunities to enter the funnel. That narrower treatment improved the final purchase outcome even though the immediate entry metric did not improve.
Later, a broad replacement of an early explanatory section lost heavily. The result suggested that the original “too long” content was performing trust and reassurance work that the new product-mechanics section removed.
A subsequent focused treatment clarified cost and process before the user entered the transaction flow. It improved purchases.
The learning sequence was:
- More funnel entry opportunities can matter without changing the nearest metric.
- Early explanation may be persuasion, not merely friction.
- Cost and process uncertainty appear commercially important.
- The primary outcome must remain downstream purchase quality, not only flow entry.
The later treatment was stronger because earlier tests constrained the hypothesis space.
This is what compounding learning looks like: not a perfect ladder of winners, but a progressively better model of the decision.
Ongoing experimentation needs a learning operating system
Experiment velocity alone does not create compounding value.
The program needs five persistent artifacts.
1. Experiment brief
Before launch, record:
- business problem;
- evidence;
- causal hypothesis;
- treatment mechanism;
- eligible population;
- primary and guardrail metrics;
- decision rule; and
- expected learning if the treatment wins, loses, or remains inconclusive.
2. Decision log
After analysis, record:
- result and uncertainty;
- data-quality checks;
- ship decision;
- deployment status;
- unresolved questions; and
- next action.
3. Living Customer Model
Update only the behavioral claims supported or contradicted by the result.
4. Deployment register
Track whether validated changes were permanently implemented and verified.
A winning test left inside a temporary variation has not created durable value.
5. Cumulative-impact register
Track the combined effect of comparable changes that were actually deployed. The PostHog implementation for Growth Lever Teams article explains the measurement, identity, exposure, revenue, and QA dependencies required before that register can be trusted.
The detailed methodology belongs in How to Measure the Cumulative Impact of Conversion Rate Optimization. Do not use test count or win rate as a substitute for deployed impact.
Starting-point decision matrix
| Question | If yes | If no |
|---|---|---|
| Can the primary business outcome and critical funnel steps be trusted? | Continue | Start with a conversion tracking audit |
| Is the current experience structurally incompatible with the present offer, platform, or user journey? | Consider a tested bundled redesign | Prefer a ranked sequence of hypotheses |
| Do several changes need to work together to create a coherent treatment? | Test the bundle, with bundle-level claims | Separate unrelated mechanisms |
| Is there enough eligible outcome volume to distinguish useful effects? | Use controlled experiments | Use research, staged rollout, or fewer high-conviction decisions with explicit uncertainty |
| Can the team implement and permanently deploy validated changes? | Ongoing testing can create realized impact | Fix delivery capacity before increasing test velocity |
| Does every experiment update a decision log and customer model? | Learning can compound | Build the learning system before calling the work an experimentation program |
| Is the experimentation system itself trustworthy? | Scale the cadence | Run an experimentation system audit |
What each engagement should deliver
One-time CRO redesign
Require:
- current-state evidence;
- bundle-level hypothesis;
- old experience retained as control;
- primary outcome and guardrails;
- implemented treatment;
- experiment or staged validation;
- analysis and ship decision;
- rollback plan; and
- a documented set of follow-up hypotheses.
Do not accept a design file and a list of best practices as proof of value.
Ongoing experimentation program
Require:
- trustworthy measurement;
- ranked hypothesis backlog;
- research and diagnostic cadence;
- treatment design and implementation;
- experiment QA;
- downstream primary and guardrail metrics;
- deployment ownership;
- decision and learning reviews;
- Living Customer Model;
- deployment register; and
- cumulative business-impact reporting.
For partner selection, use How to Evaluate a CRO Agency: Measure Deployed Business Impact, Not Win Rate.
Final decision rule
Use a one-time tested redesign when the business needs a coherent structural reset.
Start with ongoing experimentation when the business can support a sequence of decisions and wants both performance improvement and cumulative customer understanding.
Start with an audit when neither option has a trustworthy problem definition.
Then apply one rule to every test:
Make the treatment bold enough to matter, coherent enough to interpret, measured far enough downstream to reflect business value, and documented well enough to improve the next decision.
Frequently asked questions
Is ongoing CRO always better than a one-time redesign?
No.
Ongoing CRO creates more long-term value only when the company has trustworthy measurement, enough eligible outcomes, implementation capacity, and a process for retaining and applying learning.
A tested one-time redesign can be rational when the current experience is structurally unfit or a large change is unavoidable.
Should an A/B test change only one element?
No.
It should usually test one coherent hypothesis or one bundle-level decision. Several UI elements can change together when they express the same mechanism. Literal one-element testing can miss interactions and produce treatments too weak to matter.
Can a complete website redesign be A/B tested?
Yes, when the old and new experiences can run concurrently for comparable eligible users and the assignment, exposure, outcomes, and guardrails can be measured correctly.
The result estimates the effect of the complete redesign, not each component.
What can you learn from a bundled redesign test?
You can learn whether the bundle improved the chosen outcome enough to ship.
Diagnostic metrics and qualitative evidence may suggest mechanisms, but the experiment does not isolate the effect of each changed component.
What should we do if the redesign loses?
Keep the control, verify the experiment, identify where the variants diverged, examine the bundle hypothesis, and design a narrower follow-up around the most important unresolved mechanism.
Do not immediately replace it with another unrelated redesign.
What is the difference between one hypothesis and one change?
A hypothesis is a causal proposition about why a treatment should affect an outcome.
One hypothesis may require several coordinated changes. One change may also contain several unrelated hypotheses. Interpretability depends on the mechanism, not the number of edited elements.
What is a Living Customer Model?
It is a versioned register of supported behavioral claims, their context, evidence, uncertainty, counterevidence, implications, and next questions.
It is not a static persona and should not convert one experiment into a universal statement about customers.
Can a losing experiment still create value?
Yes.
It can prevent a harmful deployment and improve future decisions. That value depends on trustworthy measurement, an interpretable treatment, and a documented model update.
How should an experimentation program be measured?
Measure trustworthy experiment execution, deployment of validated changes, cumulative impact on a consistent business outcome, avoided regressions, and whether evidence improves subsequent decisions.
Do not use win rate or test count as the primary verdict.
When should we start with a CRO audit?
Start with an audit when the data cannot be trusted, the business outcome is unclear, bottlenecks have not been ranked, traffic or conversion volume is uncertain, or implementation constraints are unknown.
A serious audit should produce ranked, testable hypotheses rather than a generic UX checklist. When the main uncertainty is metric selection, use the planned guide to A/B testing metrics closest to business value rather than choosing the easiest event to move.
The next useful decision
If you are deciding between one major CRO intervention and an ongoing testing program, begin by clarifying the uncertainty:
- Do you need a structural replacement?
- Do you need a reliable opportunity map?
- Or do you need a repeatable system that improves both the funnel and the team’s understanding of the customer?



