Robot Policy Training with Mixed-Fidelity Simulation Phases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training robot control policies is costly and time-consuming due to the computational resources required for simulating robot behavior in high fidelity environments, which can be mitigated by using a combination of high speed, low fidelity and high fidelity simulations.

Innovation Solution

Initial training with high speed, low fidelity simulation followed by incremental transition to high fidelity simulation, with real-world data integration, to conserve computational resources and reduce training time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high fidelity simulation is used for training robot control policies, then the accuracy and realism of the simulation increases, but the computational resources and training time required increase dramatically

Engineering Contradiction:
Improvesimulation accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The training process is segmented into multiple phases with different fidelity levels. Initially, low fidelity simulation is used for bulk training, then high fidelity simulation is introduced incrementally. This segmentation allows the system to benefit from both fast low-fidelity training and accurate high-fidelity training without paying the full computational cost throughout the entire training process.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The simulation fidelity is made dynamic rather than static. The system transitions from low fidelity to high fidelity simulation during training based on performance criteria. This dynamic adjustment optimizes the balance between computational efficiency and training accuracy at different stages of the learning process.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If high fidelity simulation is used for training robot control policies, then the simulation realism improves, but the computational resources consumed increase

Engineering Contradiction:
Improvesimulation accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

Instead of applying high fidelity simulation throughout the entire training process, the system applies it partially - only during specific phases when the policy has already learned basic behaviors from low fidelity simulation. This partial application of high fidelity simulation reduces computational resource consumption while still achieving the desired training accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The computational resources are segmented and allocated differently across training phases. Low fidelity simulation consumes fewer resources and is used for the majority of training, while high fidelity simulation consumes more resources but is used only when necessary for fine-tuning and achieving high accuracy.

Inventive Principle:
Principle #1Segmentation

3Reliability

If more training data is generated from high fidelity simulation, then the training quality improves, but the training time and computing power required expand dramatically

Engineering Contradiction:
Improvetraining qualityVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary training using low fidelity simulation before introducing high fidelity simulation. This preliminary action allows the robot control policy to learn basic behaviors and patterns efficiently, so that when high fidelity simulation data is later introduced, the model can leverage this pre-learned knowledge and require fewer high-fidelity training examples to achieve high quality performance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12533801B2Training with high fidelity simulations and high speed low fidelity simulations
Publication Date: 2026.01.27 GDM HOLDING LLC
  • US12533801B2 patent drawing
  • US12533801B2 patent drawing
  • US12533801B2 patent drawing

AI summary

Implementations are provided for training a robot control policy for controlling a robot. During a first training phase, the robot control policy is trained using a first set of training data that includes (i) training data generated based on simulated operation of the robot in a first fidelity simulation, and (ii) training data generated based on simulated operation of the robot in a second fidelity simulation, wherein the second fidelity is greater than the first fidelity. When one or more criteria for commencing a second training phase are satisfied, the robot control policy is further trained using a second set of training data that also include training data generate based on simulated operation of the robot in the first and second fidelity simulations, which has a ratio therebetween lower than that in the first set of training data.