Pre-approved A/A Data Buckets for Online Experiment Validation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current online experimentation methods face challenges in reducing the time required for validating and ensuring the accuracy of data buckets, identifying gaps between expected and actual data bucket sizes, and detecting inconsistencies, which can lead to inaccurate results due to user misplacement and uneven group sizes in A/B testing.

Innovation Solution

A system and method for providing pre-approved data buckets by generating user engagement parameters, ranking, and excluding values to create a homogenous set for bucket allocation, along with monitoring layers to detect discrepancies and inconsistencies, ensuring accurate user placement and experiment validity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional A/A validation process is used to ensure data bucket accuracy, then measurement precision is improved, but loss of time increases due to 4-5 days validation period

Engineering Contradiction:
Improvedata bucket validation accuracyVSAvoidvalidation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary hashing of user identifiers and pre-approval of data buckets before the actual A/B experiment begins. By pre-computing hash values and validating data bucket assignments in advance, the system eliminates the need for time-consuming post-hoc validation, reducing validation time from 4-5 days to a minimal preprocessing period while maintaining measurement precision through pre-verified bucket integrity.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If multiple data buckets are opened for A/A validation to account for potential failures, then reliability is improved, but device complexity increases due to bucket selection overhead

Engineering Contradiction:
Improveexperiment validityVSAvoidbucket management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements self-verification mechanisms where data buckets automatically validate their own integrity through pre-computed hash values and metadata. Each bucket contains verification data that allows automatic detection of inconsistencies without requiring external validation processes or complex selection logic, thereby maintaining reliability while reducing management complexity.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If manual monitoring of data bucket consistency is performed, then measurement precision is improved, but loss of time increases due to continuous verification requirements

Engineering Contradiction:
Improvedata bucket consistencyVSAvoidmonitoring time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system implements automated feedback loops that continuously monitor data bucket consistency by comparing actual user assignments against expected hash value distributions. Verification layers automatically detect discrepancies and trigger alerts or corrections without requiring manual intervention, maintaining measurement precision while eliminating time-consuming human monitoring activities.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12141097B2Method and system for providing pre-approved A/A data buckets
Publication Date: 2024.11.12 YAHOO ASSETS LLC
  • US12141097B2 patent drawing
  • US12141097B2 patent drawing
  • US12141097B2 patent drawing

AI summary

The present teaching generally relates to detecting providing pre-validated data buckets for online experiments. In a non-limiting embodiment, user activity data representing user activity for a first plurality of user identifiers may be obtained. A first set of values and a second values, representing first and second user engagement parameters, respectively, may be generated for each user identifier based on the user activity data. A first ranking and a second ranking may be determined for the first and second sets, respectively. A first exclusion range including a first number of values to be removed from the first and second sets may be determined. A homogenous value set may be generated by removing the first number of values from the first and second sets, where each value from the homogenous value set corresponds to a user identifier available to be placed in a data bucket for an online experiment.