Robot Control Policy Training with Adaptive Simulation Fidelity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training robot control policies using traditional methods is resource-intensive and time-consuming, especially when simulating robot behavior in high-fidelity environments, which requires significant computational resources and prolongs the engineering cycle time.

Innovation Solution

Implementing a training method that initially uses high-speed, low-fidelity simulation to generate training data, gradually increasing the proportion of high-fidelity simulation data as learning slows, and eventually incorporating real-world data, to conserve computational resources while maintaining performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high fidelity simulation is used for training robot control policies, then the accuracy and realism of training data is improved, but computational resource consumption and training time increase dramatically

Engineering Contradiction:
Improveaccuracy of training dataVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The patent segments the training process into distinct phases: initial training using low-fidelity simulations for rapid learning, followed by fine-tuning using high-fidelity simulations. This segmentation allows the system to benefit from both speed and accuracy at appropriate stages, reducing overall computational resource consumption while maintaining training data accuracy where most needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by using low-fidelity simulations to perform initial training and establish baseline performance before transitioning to computationally intensive high-fidelity simulations. This preliminary training phase prepares the robot control policies in advance, so that subsequent high-fidelity training requires fewer resources and less time to achieve the same level of accuracy.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If high fidelity simulation is used for training robot control policies, then the realism and quality of training data is improved, but engineering cycle time is prolonged

Engineering Contradiction:
Improvequality of training dataVSAvoidengineering cycle time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The training process is segmented into multiple stages with different fidelity requirements. Low-fidelity simulations handle the bulk of training iterations quickly, while high-fidelity simulations are applied selectively in later stages for fine-tuning. This segmentation dramatically reduces engineering cycle time compared to using high-fidelity simulations throughout the entire training process.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by using high-fidelity simulations only for the portion of training that requires maximum realism (fine-tuning phase), rather than applying them excessively throughout all training stages. This selective application maintains training data quality where critical while minimizing the time loss associated with computationally intensive simulations.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If low fidelity simulation is used for training, then computational resources are conserved and training speed is improved, but the accuracy and realism of learned policies may be adversely affected

Engineering Contradiction:
Improvetraining speedVSAvoidaccuracy of learned policies
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the training workflow into two distinct phases: a rapid learning phase using low-fidelity simulations that conserves computational resources and achieves high training speed, followed by an accuracy refinement phase using high-fidelity simulations that corrects and enhances policy accuracy. This segmentation ensures that the benefits of fast training are not sacrificed for accuracy, nor is accuracy compromised by relying solely on low-fidelity data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses low-fidelity simulations as an intermediary training mechanism that bridges the gap between theoretical learning and realistic performance. These simulations provide a computationally efficient intermediate step that prepares policies for subsequent refinement with high-fidelity data, mediating between the extremes of speed and accuracy requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4410499A1Training with high fidelity simulations and high speed low fidelity simulations
Publication Date: 2024.08.07 GDM HOLDING LLC
  • EP4410499A1 patent drawingFigure 1A
  • EP4410499A1 patent drawingFigure 1B
  • EP4410499A1 patent drawingFigure 2

AI summary

Implementations are provided for training a robot control policy for controlling a robot. During a first training phase, the robot control policy is trained using a first set of training data that includes (i) training data generated based on simulated operation of the robot in a first fidelity simulation, and (ii) training data generated based on simulated operation of the robot in a second fidelity simulation, wherein the second fidelity is greater than the first fidelity. When one or more criteria for commencing a second training phase are satisfied, the robot control policy is further trained using a second set of training data that also include training data generate based on simulated operation of the robot in the first and second fidelity simulations, which has a ratio therebetween lower than that in the first set of training data.