Scaling Strategic Experimentation: Lessons from Apurva Sandbhor on Building Enterprise-Grade Product Testing Ecosystems

In the evolving landscape of digital commerce and product development, the transition from tactical A/B testing to a holistic experimentation culture remains a significant hurdle for many Fortune 500 organizations. Apurva Sandbhor, Manager of Platform and Product Experimentation at The Home Depot, provides a blueprint for this transition in the 25th installment of the CRO Perspectives series. Based in Atlanta, Georgia, Sandbhor oversees a complex experimentation infrastructure for one of the world’s largest retailers, focusing on how organizations can scale their testing programs across diverse product lines and teams without succumbing to operational friction or data silos.

The core of Sandbhor’s philosophy rests on the premise that scaling a team without a corresponding evolution in platform architecture is an exercise in futility. As companies expand their digital footprints, the traditional impulse to increase headcount often leads to a "centralized bottleneck" where a small team of experts becomes overwhelmed by requests from across the enterprise. Sandbhor argues that true scale is achieved not through human capital alone, but through a platform-first strategy that empowers individual product squads while maintaining rigorous centralized governance.

The Architectural Foundation of Scaled Experimentation

The history of experimentation in the corporate sector has often followed a linear path: from a single "optimization" specialist to a centralized department, and eventually to a decentralized model where every product manager is an experimenter. However, this final stage often results in "metric chaos" where different teams use different statistical methods or definitions of success. To mitigate this, Sandbhor introduces the "Dual-Lane" framework, a structural operating model designed to decouple day-to-end enablement from core innovation.

Chasing Just Win Rates Incentivizes Your Team to Test Low-Reward Ideas

In the first lane, "Core Platform Innovation," a centralized team focuses on the infrastructure itself—building the tools, statistical engines, and automated pipelines that the rest of the company uses. In the second lane, "Squad Enablement," this same central team acts as a consultant, providing the training and frameworks that allow decentralized product squads to run their own tests autonomously. This model prevents the central team from becoming a reactive service desk and instead positions them as architects of a "force multiplier" platform.

The implications of this shift are significant. By automating the "low-value" aspects of experimentation—such as QA, traffic allocation, and basic data analysis—organizations can reduce the end-to-end testing lifecycle by as much as 50%. This increased velocity does not come at the cost of quality; rather, it is supported by "embedded statistical guardrails" that prevent common errors like Sample Ratio Mismatch (SRM) or the misinterpretation of false positives.

Moving Beyond the "Win Rate" Fallacy

Perhaps the most provocative aspect of Sandbhor’s perspective is the rejection of the "win rate" as a primary KPI for experimentation programs. In many organizations, success is measured by the percentage of tests that yield a statistically significant positive result. Sandbhor argues that this creates an insidious incentive structure. When teams are judged on win rates, they become risk-averse, focusing on minor "cosmetic" changes—like button colors or font sizes—that are likely to succeed but unlikely to drive transformative business value.

Citing research by Harvard Business School Professor Stefan Thomke, who noted that in many high-performing organizations, nearly 90% of experiments do not yield the expected positive result, Sandbhor suggests a paradigm shift toward "learning value" and "decision integrity." Under this model, an experiment that fails to move the needle is not a failure if it provides a high-integrity insight that prevents a costly strategic misstep.

Chasing Just Win Rates Incentivizes Your Team to Test Low-Reward Ideas

To replace the win rate, Sandbhor proposes three alternative metrics:

  1. Experimentation Velocity: The speed at which a team can move from a hypothesis to a validated insight.
  2. Learning Density: The volume of actionable insights generated per experiment, regardless of whether the result was a "win" or a "loss."
  3. Strategic Risk Avoidance: The estimated revenue loss prevented by identifying flawed features through testing before a full-scale rollout.

Data Integrity and the Anomaly Decision Flow

In a high-traffic environment like The Home Depot’s digital platform, data integrity is the primary currency of the experimentation team. Sandbhor emphasizes that "velocity without safeguarding data integrity is an existential risk." To manage this, she advocates for an automated "Anomaly Decision Flow." This system uses a proprietary statistical engine to monitor tests in real-time. If a test triggers high metric volatility or an SRM alert, it is automatically paused.

The decision to resume, iterate, or scrap the test is then governed by a strict risk-reward matrix. If the anomaly is traced to an infrastructure issue, the test is rerun. If the data suggests a fundamental flaw in the hypothesis, the test is scrapped to protect the customer experience. This rigorous approach ensures that executive decisions are based on "ironclad insurance" rather than statistical noise.

The Case for Server-Side Experimentation

As organizations mature, they often hit the limits of client-side testing—which can impact site performance and is often limited to UI changes. Sandbhor outlines the prerequisites for moving to server-side experimentation, a more robust method where tests are executed on the server before the page is even sent to the user’s browser.

Chasing Just Win Rates Incentivizes Your Team to Test Low-Reward Ideas

However, server-side testing requires a much higher level of technical maturity. Organizations must have a robust data culture and a "modernized architecture" that supports the decoupling of deployment from activation. This allows developers to "dark launch" features and activate them for specific test segments without a full code deployment. Sandbhor warns that moving to server-side testing too early, without the proper organizational alignment and training, can lead to increased complexity without a corresponding increase in ROI.

A Behavioral Case Study: The Perils of Choice Paralysis

To illustrate the power of strategic experimentation, Sandbhor shares a case study involving a high-visibility checkout optimization at a major retailer. The team introduced an algorithmic recommendation engine designed to cross-sell accessories during the checkout process. Initial data showed a surprising and significant drop in overall cart conversion rates.

Rather than simply rolling back the feature, the team conducted deep-dive segment analysis. They discovered that while the recommendations were accurate, presenting multiple individual product choices at a high-intent stage of the funnel triggered "decision paralysis." Customers were so overwhelmed by the choices that they abandoned the purchase entirely.

The refined hypothesis was to replace the individual choices with a single, pre-configured "bundle" of accessories. A follow-up test validated this approach, converting the initial loss into a net revenue lift. This example underscores the importance of the "insight-to-hypothesis loop," where the conclusion of one test becomes the data-driven foundation for the next.

Chasing Just Win Rates Incentivizes Your Team to Test Low-Reward Ideas

Managing Executive Expectations and the Role of AI

Maintaining executive buy-in is a constant challenge for experimentation leaders, particularly when tests do not result in immediate revenue gains. Sandbhor recommends a "Strategic Risk Avoidance" reporting style. By showing leadership exactly how much money was saved by not launching a feature that the data proved would have been detrimental, the experimentation program is framed as a safeguard for corporate capital.

Looking toward the future, Sandbhor addresses the growing role of Artificial Intelligence in the field. While AI can commoditize the execution layer—automating data pulls, building test variants, and enforcing guardrails—she maintains that human judgment remains irreplaceable. Leadership, according to Sandbhor, involves choosing which "mountain" to climb, while AI helps optimize the path. Strategic boundaries, ethical considerations, and the translation of complex user psychology into novel hypotheses remain the domain of the human strategist.

Conclusion: Experimentation as a Long-Term Capability

The insights provided by Apurva Sandbhor suggest that the most successful experimentation programs are those that view the practice not as a series of isolated projects, but as a permanent, automated infrastructure system. By focusing on platform architecture, moving away from vanity metrics like win rates, and prioritizing data integrity through automation, enterprise organizations can build a culture of continuous learning.

As the digital landscape becomes increasingly competitive and AI-driven, the ability to uncover business truths quickly and safely becomes a primary competitive advantage. For leaders like Sandbhor, the goal is clear: build a system that acts as a force multiplier, allowing the organization to innovate at scale while protecting the most valuable asset of all—customer trust.

Related Posts

How to build SaaS comparison pages buyers actually trust (with 4 examples + a free template)

The B2B software-as-a-service (SaaS) industry is currently navigating a significant paradigm shift in how buyers evaluate and purchase solutions. As market saturation increases and procurement cycles become more scrutinized, the…

Crazy Egg Enhances Conversion Optimization Suite with New Inline Embedded Survey Functionality

Crazy Egg, a pioneer in website optimization and user behavior analytics, has announced a significant update to its feedback collection capabilities by introducing inline embedded surveys to its platform. This…

You Missed

The Progress of Global Health Initiatives and the Evolving Landscape of Maternal Mortality Reduction through the Goalkeepers 2017 Report

  • By
  • August 28, 2026
  • 3 views
The Progress of Global Health Initiatives and the Evolving Landscape of Maternal Mortality Reduction through the Goalkeepers 2017 Report

Optimizing LLM Inference: The Evolution of PagedAttention and RadixAttention in High-Performance Serving Engines

  • By
  • August 28, 2026
  • 3 views
Optimizing LLM Inference: The Evolution of PagedAttention and RadixAttention in High-Performance Serving Engines

ADM Communications Director Darcie Rosenthal Redefines Manager Engagement and AI Integration in Modern Corporate Strategy

  • By
  • August 28, 2026
  • 3 views
ADM Communications Director Darcie Rosenthal Redefines Manager Engagement and AI Integration in Modern Corporate Strategy

AI-Driven Misinformation and the Rise of Recycled Truth: How a Restaurant Brand Faced a Modern Crisis Management Challenge

  • By
  • August 28, 2026
  • 3 views
AI-Driven Misinformation and the Rise of Recycled Truth: How a Restaurant Brand Faced a Modern Crisis Management Challenge

The Future of Strategic Communication Measurement: Four Essential Metrics for the 2026 Business Landscape

  • By
  • August 28, 2026
  • 5 views
The Future of Strategic Communication Measurement: Four Essential Metrics for the 2026 Business Landscape

Keeping Cool When Everything Is on Fire: Producer Marta Ravin’s Lessons from Her Celebrity-Filled TV Career

  • By
  • August 28, 2026
  • 4 views
Keeping Cool When Everything Is on Fire: Producer Marta Ravin’s Lessons from Her Celebrity-Filled TV Career