Monitoring Layer for Detecting Data Bucket Discrepancies

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current online experimentation methods face challenges in reducing the time required for validating and ensuring the accuracy of data buckets, identifying gaps between expected and actual data bucket sizes, and detecting inconsistencies between data buckets, which can lead to inaccurate results due to insufficient user population and overlapping user assignments.

Innovation Solution

A system and method for providing data buckets for online experiments that involves obtaining user activity data, generating sets of user engagement parameters, determining rankings, and identifying exclusion ranges to create a homogenous value set for user assignment, as well as implementing a monitoring layer to detect discrepancies and inconsistencies within the experimentation platform.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional A/A validation methods are used to validate data buckets, then validation thoroughness is improved, but validation time increases to four to five days

Engineering Contradiction:
Improvevalidation thoroughnessVSAvoidvalidation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent pre-computes hash values for user identifiers and stores them in a hash table before the validation process begins. This preliminary action allows the validation system to quickly retrieve and compare hash values without performing computationally intensive hashing operations during the actual validation, thereby reducing validation time from four to five days to a much shorter duration while maintaining thoroughness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a copy of the data bucket assignments and stores them in a separate validation database. This copying allows the validation process to work with the copied data without affecting the original experiment data, enabling thorough validation while using efficient data structures that reduce processing time.

Inventive Principle:
Principle #26Copying

2Reliability

If multiple data buckets are opened for A/A validation to account for potential failures, then validation reliability is improved, but decision complexity and delay increase

Engineering Contradiction:
Improvevalidation reliabilityVSAvoiddecision complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements an automated feedback mechanism that monitors the validation results of multiple data buckets in real-time. When a data bucket successfully passes the A/A validation criteria, the system automatically provides feedback to stop the validation process and select that bucket for the experiment. This feedback-driven approach maintains high reliability by validating multiple buckets while reducing decision complexity through automation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The validation system automatically selects and configures appropriate data buckets based on predefined validation criteria without requiring manual intervention. The system self-services by autonomously determining which data buckets have passed validation and are ready for experimentation, thereby reducing the complexity of manual decision-making while maintaining validation reliability.

Inventive Principle:
Principle #25Self-service

3Productivity

If data bucket size is reduced to speed up experimentation, then productivity is improved, but measurement precision deteriorates due to insufficient user population

Engineering Contradiction:
Improveexperimentation speedVSAvoidresults accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the parameter of data bucket size dynamically based on the available traffic and validation requirements. Instead of using fixed large data buckets that slow down experimentation, the system adjusts the data bucket size to be just sufficient for statistical validity given the current traffic conditions. This parameter optimization allows faster experimentation while maintaining measurement precision by ensuring each bucket has enough users for reliable results.

Inventive Principle:
Principle #35Parameter changes

4Device complexity

If user assignment is not monitored, then system complexity is reduced, but data consistency deteriorates due to overlapping user assignments

Engineering Contradiction:
Improvesystem complexityVSAvoiddata consistency
Core Design Contradiction:
Device complexityVSStability of the object's composition

Solution Approach 1:

The patent introduces an intermediary monitoring layer that sits between the user assignment process and the experiment execution. This intermediary layer tracks user assignments across multiple data buckets and experiments, detecting and preventing overlapping assignments without adding significant complexity to the core assignment logic. The intermediary maintains data consistency by ensuring each user is properly assigned to only one group within an experiment.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11227256B2Method and system for detecting gaps in data buckets for A/B experimentation
Publication Date: 2022.01.18 YAHOO AD TECH LLC
  • US11227256B2 patent drawing
  • US11227256B2 patent drawing
  • US11227256B2 patent drawing

AI summary

The present teaching generally relates to detecting data bucket discrepancies associated with online experiments. In a non-limiting embodiment, a monitoring layer may be generated within an online experimentation platform that includes at least a first layer, and where a first online experiment is associated with the first layer, the monitoring layer includes a monitoring layer data bucket, and the first layer includes at least a first data bucket. First data representing user activity associated with a first plurality of identifiers may be obtained, the user activity being associated with the first layer. Second data representing at least one user engagement parameter may be generated, and a first discrepancy between the first and second data may be determined. The first discrepancy indicating a first amount of identifiers that include a first metadata tag associated with the first layer and lack a second metadata tag associated with the monitoring layer.