Scaling Strategic Experimentation: Lessons from Apurva Sandbhor on Building Enterprise-Grade Product Testing Ecosystems

In the evolving landscape of digital commerce and product development, the transition from tactical A/B testing to a holistic experimentation culture remains a significant hurdle for many Fortune 500 organizations. Apurva Sandbhor, Manager of Platform and Product Experimentation at The Home Depot, provides a blueprint for this transition in the 25th installment of the CRO Perspectives series. Based in Atlanta, Georgia, Sandbhor oversees a complex experimentation infrastructure for one of the world’s largest retailers, focusing on how organizations can scale their testing programs across diverse product lines and teams without succumbing to operational friction or data silos.

The core of Sandbhor’s philosophy rests on the premise that scaling a team without a corresponding evolution in platform architecture is an exercise in futility. As companies expand their digital footprints, the traditional impulse to increase headcount often leads to a "centralized bottleneck" where a small team of experts becomes overwhelmed by requests from across the enterprise. Sandbhor argues that true scale is achieved not through human capital alone, but through a platform-first strategy that empowers individual product squads while maintaining rigorous centralized governance.

The Architectural Foundation of Scaled Experimentation

The history of experimentation in the corporate sector has often followed a linear path: from a single "optimization" specialist to a centralized department, and eventually to a decentralized model where every product manager is an experimenter. However, this final stage often results in "metric chaos" where different teams use different statistical methods or definitions of success. To mitigate this, Sandbhor introduces the "Dual-Lane" framework, a structural operating model designed to decouple day-to-end enablement from core innovation.

Chasing Just Win Rates Incentivizes Your Team to Test Low-Reward Ideas

In the first lane, "Core Platform Innovation," a centralized team focuses on the infrastructure itself—building the tools, statistical engines, and automated pipelines that the rest of the company uses. In the second lane, "Squad Enablement," this same central team acts as a consultant, providing the training and frameworks that allow decentralized product squads to run their own tests autonomously. This model prevents the central team from becoming a reactive service desk and instead positions them as architects of a "force multiplier" platform.

The implications of this shift are significant. By automating the "low-value" aspects of experimentation—such as QA, traffic allocation, and basic data analysis—organizations can reduce the end-to-end testing lifecycle by as much as 50%. This increased velocity does not come at the cost of quality; rather, it is supported by "embedded statistical guardrails" that prevent common errors like Sample Ratio Mismatch (SRM) or the misinterpretation of false positives.

Moving Beyond the "Win Rate" Fallacy

Perhaps the most provocative aspect of Sandbhor’s perspective is the rejection of the "win rate" as a primary KPI for experimentation programs. In many organizations, success is measured by the percentage of tests that yield a statistically significant positive result. Sandbhor argues that this creates an insidious incentive structure. When teams are judged on win rates, they become risk-averse, focusing on minor "cosmetic" changes—like button colors or font sizes—that are likely to succeed but unlikely to drive transformative business value.

Citing research by Harvard Business School Professor Stefan Thomke, who noted that in many high-performing organizations, nearly 90% of experiments do not yield the expected positive result, Sandbhor suggests a paradigm shift toward "learning value" and "decision integrity." Under this model, an experiment that fails to move the needle is not a failure if it provides a high-integrity insight that prevents a costly strategic misstep.

Chasing Just Win Rates Incentivizes Your Team to Test Low-Reward Ideas

To replace the win rate, Sandbhor proposes three alternative metrics:

  1. Experimentation Velocity: The speed at which a team can move from a hypothesis to a validated insight.
  2. Learning Density: The volume of actionable insights generated per experiment, regardless of whether the result was a "win" or a "loss."
  3. Strategic Risk Avoidance: The estimated revenue loss prevented by identifying flawed features through testing before a full-scale rollout.

Data Integrity and the Anomaly Decision Flow

In a high-traffic environment like The Home Depot’s digital platform, data integrity is the primary currency of the experimentation team. Sandbhor emphasizes that "velocity without safeguarding data integrity is an existential risk." To manage this, she advocates for an automated "Anomaly Decision Flow." This system uses a proprietary statistical engine to monitor tests in real-time. If a test triggers high metric volatility or an SRM alert, it is automatically paused.

The decision to resume, iterate, or scrap the test is then governed by a strict risk-reward matrix. If the anomaly is traced to an infrastructure issue, the test is rerun. If the data suggests a fundamental flaw in the hypothesis, the test is scrapped to protect the customer experience. This rigorous approach ensures that executive decisions are based on "ironclad insurance" rather than statistical noise.

The Case for Server-Side Experimentation

As organizations mature, they often hit the limits of client-side testing—which can impact site performance and is often limited to UI changes. Sandbhor outlines the prerequisites for moving to server-side experimentation, a more robust method where tests are executed on the server before the page is even sent to the user’s browser.

Chasing Just Win Rates Incentivizes Your Team to Test Low-Reward Ideas

However, server-side testing requires a much higher level of technical maturity. Organizations must have a robust data culture and a "modernized architecture" that supports the decoupling of deployment from activation. This allows developers to "dark launch" features and activate them for specific test segments without a full code deployment. Sandbhor warns that moving to server-side testing too early, without the proper organizational alignment and training, can lead to increased complexity without a corresponding increase in ROI.

A Behavioral Case Study: The Perils of Choice Paralysis

To illustrate the power of strategic experimentation, Sandbhor shares a case study involving a high-visibility checkout optimization at a major retailer. The team introduced an algorithmic recommendation engine designed to cross-sell accessories during the checkout process. Initial data showed a surprising and significant drop in overall cart conversion rates.

Rather than simply rolling back the feature, the team conducted deep-dive segment analysis. They discovered that while the recommendations were accurate, presenting multiple individual product choices at a high-intent stage of the funnel triggered "decision paralysis." Customers were so overwhelmed by the choices that they abandoned the purchase entirely.

The refined hypothesis was to replace the individual choices with a single, pre-configured "bundle" of accessories. A follow-up test validated this approach, converting the initial loss into a net revenue lift. This example underscores the importance of the "insight-to-hypothesis loop," where the conclusion of one test becomes the data-driven foundation for the next.

Chasing Just Win Rates Incentivizes Your Team to Test Low-Reward Ideas

Managing Executive Expectations and the Role of AI

Maintaining executive buy-in is a constant challenge for experimentation leaders, particularly when tests do not result in immediate revenue gains. Sandbhor recommends a "Strategic Risk Avoidance" reporting style. By showing leadership exactly how much money was saved by not launching a feature that the data proved would have been detrimental, the experimentation program is framed as a safeguard for corporate capital.

Looking toward the future, Sandbhor addresses the growing role of Artificial Intelligence in the field. While AI can commoditize the execution layer—automating data pulls, building test variants, and enforcing guardrails—she maintains that human judgment remains irreplaceable. Leadership, according to Sandbhor, involves choosing which "mountain" to climb, while AI helps optimize the path. Strategic boundaries, ethical considerations, and the translation of complex user psychology into novel hypotheses remain the domain of the human strategist.

Conclusion: Experimentation as a Long-Term Capability

The insights provided by Apurva Sandbhor suggest that the most successful experimentation programs are those that view the practice not as a series of isolated projects, but as a permanent, automated infrastructure system. By focusing on platform architecture, moving away from vanity metrics like win rates, and prioritizing data integrity through automation, enterprise organizations can build a culture of continuous learning.

As the digital landscape becomes increasingly competitive and AI-driven, the ability to uncover business truths quickly and safely becomes a primary competitive advantage. For leaders like Sandbhor, the goal is clear: build a system that acts as a force multiplier, allowing the organization to innovate at scale while protecting the most valuable asset of all—customer trust.

Related Posts

Instapage Unveils Advanced Campaign Scheduling and Dedicated Website Templates to Streamline Digital Marketing Workflows

The digital marketing landscape is undergoing a significant transformation as brands seek more efficient ways to bridge the gap between high-cost advertising and conversion-optimized user experiences. In a strategic move…

The Rise of India as a Global Hub for AB Testing Software A Comprehensive Guide for D2C and SaaS Enterprises

India has officially emerged as the world’s second-largest ecosystem for A/B testing and experimentation software, trailing only the United States in the total number of specialized firms and market penetration.…

You Missed

Unlocking Digital Reach: How Hosted Signup Forms Empower Businesses Without a Traditional Website

  • By
  • August 8, 2026
  • 1 views
Unlocking Digital Reach: How Hosted Signup Forms Empower Businesses Without a Traditional Website

How to Turn Your Comms Team into AI Builders: A Strategic Framework for Navigating the Velocity Gap in Public Relations

  • By
  • August 8, 2026
  • 1 views
How to Turn Your Comms Team into AI Builders: A Strategic Framework for Navigating the Velocity Gap in Public Relations

The End of Marketing Drudgery: AI Empowers Marketers to Reclaim Their Time

  • By
  • August 8, 2026
  • 1 views
The End of Marketing Drudgery: AI Empowers Marketers to Reclaim Their Time

Harnessing the Power of Emotion: How a Marketing Psychology Consultant Revolutionized Direct-to-Consumer Ad Performance

  • By
  • August 8, 2026
  • 1 views
Harnessing the Power of Emotion: How a Marketing Psychology Consultant Revolutionized Direct-to-Consumer Ad Performance

Adalysis Launches Comprehensive PPC KPI Monitoring Series to Empower Advertisers

  • By
  • August 8, 2026
  • 1 views
Adalysis Launches Comprehensive PPC KPI Monitoring Series to Empower Advertisers

Mastering the Digital Pulse: Crafting and Executing an Effective Social Media Posting Schedule for Optimal Engagement

  • By
  • August 8, 2026
  • 1 views
Mastering the Digital Pulse: Crafting and Executing an Effective Social Media Posting Schedule for Optimal Engagement