Simulation Augmented Reinforcement Learning for Content Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional online inventory management platforms face challenges in optimizing content selection processes, as the data provided is often deficient, leading to suboptimal content selection and limited user customization options, which hinders the achievement of inventory providers' objectives such as maximizing user satisfaction and revenue.

Innovation Solution

The implementation of a process selection service and a request optimization service that utilize simulation augmented reinforcement learning and machine learning models to determine the optimal content selection process between direct and indirect content, optimizing request data to enhance response relevance and return, and employing multi-layered learning and knowledge transfer to deploy models effectively in real-time scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional content selection processes are used with available data, then the selection can be made with current system capabilities, but the content selection becomes suboptimal and fails to maximize inventory provider objectives

Engineering Contradiction:
Improvecontent selection accuracyVSAvoiddata deficiency
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system performs preliminary data augmentation by generating synthetic data samples that supplement deficient real-world data before the content selection process. This pre-processing step enriches the input data with synthesized information that mimics real data distributions, enabling more accurate selection decisions even when actual data is limited.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

A simulation environment acts as an intermediary between the deficient real data and the content selection process. The simulation generates synthetic scenarios and outcomes that bridge the data gap, allowing the selection algorithm to learn from both real and simulated experiences without requiring large amounts of actual data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple content selection processes are made available to users, then user customization and flexibility are improved, but the complexity of selecting the optimal process increases

Engineering Contradiction:
Improveuser customization flexibilityVSAvoidprocess selection complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements self-service by automatically evaluating and selecting the optimal content selection process based on current conditions and provider objectives. The automated evaluation mechanism assesses multiple processes and their suitability for given scenarios, eliminating the need for users to manually navigate complex selection criteria while still providing access to multiple processes when needed.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback loops that continuously evaluate the performance of different content selection processes and adjust selections based on outcomes. This feedback mechanism learns from past selections and results, automatically optimizing the choice of process for each situation without increasing user burden.

Inventive Principle:
Principle #23Feedback

3Productivity

If real-time content selection optimization is implemented, then inventory provider objectives are maximized, but computational complexity and processing latency increase

Engineering Contradiction:
Improveinventory optimization efficiencyVSAvoidcomputational latency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary training of machine learning models using historical data and simulation environments before real-time deployment. This pre-training phase captures complex patterns and relationships offline, so that during real-time operation, the system only needs to apply pre-learned knowledge through faster inference operations, reducing real-time computational burden.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses approximation and sampling techniques that provide sufficiently accurate results without computing all possible outcomes. By selecting representative samples and using probabilistic methods, the system achieves near-optimal content selection with significantly reduced computational requirements compared to exhaustive analysis.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11847670B1Simulation augmented reinforcement learning for real-time content selection
Publication Date: 2023.12.19 AMAZON TECH INC
  • US11847670B1 patent drawing
  • US11847670B1 patent drawing
  • US11847670B1 patent drawing

AI summary

Systems, devices, and methods are described herein for improving inventory management. As used herein, “inventory” refers to digital space at an inventory providers webpage at which content can be delivered. The disclosed techniques utilize reinforced machine learning and an offline training process to train various models with which a content request corresponding to the inventory can be classified according to historical requests and a selection process identified for the request (e.g., a direct or an indirect selection process). If an indirect selection process is chosen, the content request may be optimized for that process utilizing additional machine learning models trained using reinforced machine learning and the offline training process. The disclosed techniques enable the inventory provider to optimize content selections according to a preferred objective. The training operations are performed offline, in a training system configured to simulate the run time environment.