Sequential Hypothesis Testing for Digital Marketing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional fixed-horizon hypothesis testing in digital marketing is inefficient and inaccurate, particularly when dealing with continuous non-binary data, as it requires pre-specifying a minimum detectable effect and running until a set number of samples is collected, which can lead to data inefficiency and increased risk of Type I errors, and is limited by its reliance on binary data distributions.
Innovation Solution
Sequential hypothesis testing is employed, which involves defining a data distribution model, such as an ensemble model, to estimate parameters and generate a decision boundary, allowing for real-time monitoring and flexible execution, and continuing the test beyond initial accuracy guarantees to achieve higher accuracy, without the need for a fixed sample horizon.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If fixed-horizon hypothesis testing is used with pre-specified minimum detectable effect, then the test can be run until a set number of samples is collected, but this leads to data inefficiency and increased risk of Type I errors
Solution Approach 1:
The patent transitions from fixed-horizon to sequential hypothesis testing, making the testing duration dynamic rather than static. The test continues until a pre-computed spending curve is reached, allowing the testing to adapt to the actual data accumulation rate and significance level achievement, thereby reducing unnecessary testing time while maintaining accuracy
Solution Approach 2:
The patent implements continuous monitoring of the p-value against a pre-computed spending curve during the testing process. This feedback mechanism allows real-time adjustment of the testing duration based on the actual significance level achieved, preventing both premature termination and excessive testing, thus resolving the contradiction between reliability and time loss
2Adaptability or versatility
If fixed-horizon hypothesis testing requires pre-specifying minimum detectable effect, then the test design is simplified, but this limits adaptability to actual data distributions
Solution Approach 1:
The patent changes the fundamental parameter from pre-specified minimum detectable effect to pre-computed spending curve based on actual data distribution. By computing the spending curve from the observed data characteristics rather than assuming a fixed effect size, the test becomes adaptable to continuous non-binary data while the pre-computation aspect maintains design simplicity
Solution Approach 2:
The patent segments the testing process into two distinct phases: (1) pre-computation phase where the spending curve is calculated based on expected data characteristics, and (2) execution phase where the test follows this curve adaptively. This segmentation allows complex adaptability to be achieved through pre-computation rather than real-time complex calculations
3Measurement precision
If sequential hypothesis testing continues beyond initial accuracy guarantees, then higher accuracy is achieved, but computational resources and time are increased
Solution Approach 1:
The patent performs preliminary action by pre-computing the spending curve before the actual hypothesis testing begins. This pre-computation establishes the exact path to achieve the desired significance level, preventing unnecessary continued testing beyond the point of achieving accuracy guarantees, thus avoiding wasted computational resources while maintaining high precision
Data Source
AI summary
Sequential hypothesis testing in a digital medium environment is described using continuous data. To begin, a model is received that defines at least one data distribution. Testing data is also received that describes an effect of user interactions with the plurality of options of digital content on achieving an action using continuous non-binary data. Values of parameters of the model are then estimated for each option of the plurality of options based on the testing data. In one example. A variance estimate is then generated based on the estimated values of the parameters of the model for each option of the plurality of options. From this, a determination is made as to a decision boundary based on the variance estimate and an estimate for a mean value of each option of the plurality of options based on the testing data.


