Change Point Detection in Multi-Armed Bandit Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional multi-armed bandit techniques fail to accurately model the dynamic changes in user interaction probabilities over time, leading to inefficient use of computational and digital content resources, as they either overrepresent or underrepresent changes in reward distributions.

Innovation Solution

Employing change point detection to identify shifts in reward distributions, allowing the recommendation system to regenerate statistical models and adapt recommendations, thereby addressing the middle ground between adversarial and stochastic models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If adversarial bandit model is used to model reward distributions, then changes in user interaction patterns are detected, but computational resources are inefficiently used due to overrepresentation of changes

Engineering Contradiction:
Improvedetection of changes in user interaction patternsVSAvoidcomputational efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies parameter changes by transitioning from assuming time-invariant reward distributions (adversarial model) to piecewise constant reward distributions with change points. This allows the system to adapt the model parameters dynamically - maintaining stability during constant periods and detecting changes only when significant shifts occur, thereby resolving the contradiction between detecting changes and maintaining computational efficiency

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces dynamics by making the reward distribution model adaptive rather than static. The piecewise constant model with change points allows the system to transition between different states (constant vs. changing) based on actual user interaction patterns, enabling efficient resource usage during stable periods while maintaining detection capability during transition periods

Inventive Principle:
Principle #15Dynamics

2Productivity

If stochastic bandit model is used to model reward distributions, then computational resources are efficiently used, but accuracy is reduced due to underrepresentation of changes

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidaccuracy of change detection
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent resolves this contradiction by introducing change point parameters into the stochastic bandit framework. The reward distribution is modeled as piecewise constant with parameters that remain stable during constant periods (maintaining computational efficiency) but can change at detected change points (improving accuracy), thus combining the advantages of both adversarial and stochastic models

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If reward distributions are assumed time-invariant, then model simplicity is maintained, but accuracy is reduced due to failure to capture real-world dynamics

Engineering Contradiction:
Improvemodel simplicityVSAvoidaccuracy of user interaction modeling
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies segmentation by dividing the time horizon into multiple constant segments separated by change points. Each segment maintains a simple time-invariant reward distribution model, while the overall piecewise structure captures temporal dynamics. This segmentation allows the system to maintain model simplicity within segments while accurately representing real-world changes across the entire time period

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamics into the model by allowing the reward distribution parameters to change at specific change points while remaining constant between them. This piecewise constant approach maintains simplicity during constant periods while capturing dynamic changes when they occur, resolving the contradiction between model simplicity and accuracy

Inventive Principle:
Principle #15Dynamics

4Measurement precision

If change point detection is implemented, then recommendation accuracy is improved, but system complexity increases

Engineering Contradiction:
Improverecommendation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent manages system complexity by carefully controlling which parameters change - only the reward distribution parameters at change points, while maintaining constant parameters between change points. This selective parameter change approach improves recommendation accuracy by adapting to real-world dynamics while avoiding unnecessary complexity by maintaining simplicity during stable periods

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10878451B2Change point detection in a multi-armed bandit recommendation system
Publication Date: 2020.12.29 ADOBE INC
  • US10878451B2 patent drawing
  • US10878451B2 patent drawing
  • US10878451B2 patent drawing

AI summary

Recommendation systems and techniques are described that employ change point detection to generate recommendations for digital content. In one example, a change point detection technique is employed by a recommendation system to identify when a change point has occurred at a respective time step of a series of time steps. Detection of this change point may then be used by the recommendation system to reset the statistical model to address this change as well as generate a subsequent recommendation configured for exploration of reward distributions of the items of digital marketing content.