Training Action Selection Policy Using Competency-Weighted Offline Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models require extensive online training data to achieve acceptable performance, which can be time-consuming and resource-intensive, and lack robustness and generalizability due to the absence of effective utilization of offline training data and competency measures.

Innovation Solution

A system that trains a target action selection policy using both online and offline training data, where the offline training data is conditioned on a measure of competency of a baseline agent, allowing the policy to adapt and improve performance with less online data, enhancing robustness and generalizability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If extensive online training data is used to train machine learning models, then model performance is improved, but training time and resource requirements increase

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by collecting and storing offline training data in advance from various sources (simulations, logs, expert demonstrations) before the actual model training process. This pre-collected data serves as a foundation that reduces the amount of online training data needed, thereby decreasing training time while maintaining model performance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent merges offline training data with online training data to create a comprehensive training dataset. By combining these two data sources and using appropriate weighting mechanisms, the system achieves better model performance with less online data required, thus reducing training time and computational resources.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If extensive online training data is used to train machine learning models, then model performance is improved, but resource requirements increase

Engineering Contradiction:
Improvemodel performanceVSAvoidcomputational resources
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary data collection and processing offline, storing training data in advance. This reduces the computational burden during actual model training by minimizing the amount of online data that needs to be processed, thereby decreasing energy consumption and computational resource requirements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent combines offline and online training data with intelligent sampling and weighting strategies. This merging approach allows the model to learn from diverse data sources efficiently, achieving high performance with reduced computational resources compared to using only extensive online data.

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If offline training data is not effectively utilized, then training speed is improved, but model robustness and generalizability deteriorate

Engineering Contradiction:
Improvetraining speedVSAvoidmodel robustness
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent merges offline training data (from simulations, logs, expert demonstrations) with online training data using a unified training framework. This combination enables the model to learn from diverse scenarios and edge cases present in offline data, improving robustness and generalizability while maintaining training efficiency through selective sampling and weighting.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent employs parameter changes by dynamically adjusting the weighting and sampling rates of offline versus online data during training. This allows the system to optimize the balance between training speed and model quality, ensuring that offline data contributes meaningfully to robustness without significantly slowing down the training process.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If offline training data is not effectively utilized, then training efficiency is improved, but model generalizability deteriorates

Engineering Contradiction:
Improvetraining efficiencyVSAvoidmodel generalizability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent merges offline training data representing diverse real-world scenarios with online data, using stratified sampling and importance weighting to ensure comprehensive coverage. This approach enhances model generalizability across different conditions and distributions while maintaining training efficiency through optimized data selection and processing pipelines.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent applies parameter changes by dynamically adjusting data sampling rates, weighting factors, and selection criteria for offline versus online data during training. These parameter optimizations ensure that the model learns from the most informative offline examples without overwhelming the training process, thereby maintaining efficiency while improving generalizability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240232643A9Leveraging offline training data and agent competency measures to improve online learning
Publication Date: 2024.07.11 GDM HOLDING LLC
  • US20240232643A9 patent drawing
  • US20240232643A9 patent drawing
  • US20240232643A9 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a target action selection policy to control a target agent interacting with an environment. In one aspect, a method comprises: obtaining a set of offline training data, wherein the offline training data characterizes interaction of a baseline agent with an environment as the baseline agent performs actions selected in accordance with a baseline action selection policy; generating a set of online training data that characterizes interaction of the target agent with the environment as the target agent performs actions selected in accordance with the target action selection policy; and training the target action selection policy on both: (i) the offline training data, and (ii) the online training data, wherein the training of the target action selection policy on the offline training data is conditioned on a measure of competency of the baseline agent.