Anytime Valid Confidence Sequence for A/B Test Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

A/B testing tools are often misused by practitioners with limited statistical knowledge, leading to inflated type-I errors due to 'peeking' or early stopping, which inflates type-I error rates and results in flawed conclusions.

Innovation Solution

An anytime analysis system that determines an effect metric or lift metric during an A/B test, allowing for continuous monitoring and accurate estimation of digital content variations by generating confidence intervals and effect intervals, reducing the risk of type-I errors through precise bounding of the effect metric.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If practitioners perform multiple comparisons before reaching the pre-specified sample size, then they can monitor tests continually and make early decisions, but the type-I error is drastically inflated

Engineering Contradiction:
Improvetest monitoring efficiencyVSAvoidtype-I error rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent changes the statistical parameter from fixed-horizon p-value to sequential p-value that adapts to the current sample size. The sequential p-value is calculated as the ratio of the current p-value to the product of adjustment factors, where each adjustment factor accounts for the information gained from additional observations. This dynamic parameter adjustment allows continuous monitoring while maintaining proper type-I error control.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements a feedback mechanism where the sequential monitoring process continuously updates the p-value based on accumulated data. The adjustment factors are computed from the observed data stream, creating a feedback loop that adapts the significance threshold to the current state of the experiment. This allows practitioners to peek at results without inflating error rates, as the feedback mechanism automatically compensates for multiple looks.

Inventive Principle:
Principle #23Feedback

2Loss of information

If practitioners collect more data and perform additional comparisons, then they can gather more evidence, but the type-I error rate increases significantly

Engineering Contradiction:
Improveevidence accumulationVSAvoidstatistical conclusion validity
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-specifying the maximum sample size and computing the sequence of adjustment factors in advance based on the planned sampling scheme. This preliminary configuration allows flexible monitoring during the experiment while guaranteeing type-I error control. The adjustment factors are predetermined functions of the sample size, so practitioners can collect data continuously without worrying about error inflation, as the correction mechanism is already in place.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If practitioners use fixed-horizon methodologies with pre-specified sample sizes, then type-I error is controlled, but they cannot stop tests early or monitor tests continually

Engineering Contradiction:
Improvetype-I error controlVSAvoidtest flexibility
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent introduces dynamics by transforming the static fixed-horizon testing framework into a dynamic sequential monitoring system. The sequential p-value changes over time as data accumulates, and the significance threshold adapts to the current sample size. This dynamic approach maintains the reliability of fixed-horizon methods while adding the flexibility to monitor and stop tests at any point, making the testing process adaptive rather than rigid.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250022006A1Anytime valid confidence sequence for relative increment in online tests
Publication Date: 2025.01.16 ADOBE INC
  • US20250022006A1 patent drawing
  • US20250022006A1 patent drawing
  • US20250022006A1 patent drawing

AI summary

A method, a system, and a computer program product for analyzing data collected during a randomized controlled experiment to determine an effect of variations of digital content. Determination of the effect includes execution of first and second testing sequences that prompt responses to first and second digital contents, respectively, from users. The testing sequences execute during a predetermined duration of time. Responses to the first and second testing sequences generate first and second test data, respectively. One or more confidence intervals for each first and second test data are generated at a randomly selected time during the predetermined duration of time. A testing metric indicating the effect of the second digital content over the first digital content is determined at the randomly selected time. The testing metric is determined at any time before expiration of the predetermined duration of time.