Domain Randomization Distribution Learning for Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current reinforcement learning (RL) systems face challenges in effectively training agents to handle diverse real-world situations due to the 'reality gap' between simulated and real-world environments, with domain randomization (DR) methods being sensitive to the selection of randomization distributions.

Innovation Solution

A method and system that simultaneously learn a domain randomization distribution and an agent policy using reinforcement learning, where the agent policy is trained over a range of simulated environmental parameters to optimize performance, allowing for better generalization to real-world scenarios with fewer resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If domain randomization is used to train RL agents in simulation, then the agents can be trained with fewer real-world data requirements, but the success is highly dependent on correct selection of randomization distribution

Engineering Contradiction:
Improvereal-world data requirementsVSAvoidsuccess dependence on DR distribution selection
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent automatically adjusts domain randomization parameters (distribution types, ranges, and combinations) during the training process based on agent performance feedback. This dynamic parameter adaptation eliminates the need for manual DR distribution selection while maintaining training effectiveness, directly resolving the contradiction between reduced real-world data needs and reliability concerns.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system implements a feedback loop where agent performance in simulation is continuously monitored and used to refine the domain randomization distribution. This closed-loop approach ensures that the DR parameters are optimized based on actual training outcomes, making the process reliable without requiring expert manual configuration.

Inventive Principle:
Principle #23Feedback

2Ease of operation

If traditional domain randomization with fixed distributions is used, then training can be simpler, but the agent may not generalize well to diverse real-world situations

Engineering Contradiction:
Improvetraining simplicityVSAvoidgeneralization to real-world situations
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent transforms the static, fixed domain randomization distribution into a dynamic system that adapts during training. The DR parameters evolve based on agent performance, allowing the training process to automatically explore diverse environmental conditions that improve generalization while maintaining operational simplicity through automation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs self-optimization of domain randomization parameters without external intervention. The training process automatically identifies which environmental variations are most beneficial for agent generalization and adjusts the DR distribution accordingly, eliminating the need for complex manual configuration while achieving superior adaptability.

Inventive Principle:
Principle #25Self-service

3Reliability

If extensive real-world training data is collected, then agent performance can be improved, but the cost and time requirements increase significantly

Engineering Contradiction:
Improveagent performanceVSAvoidtraining time and resource costs
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates virtual copies of diverse real-world environments through domain randomization in simulation. By training agents on these synthesized virtual experiences with automatically optimized DR parameters, the system achieves real-world performance without the time and cost of collecting extensive actual real-world data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary training in simulation with automatically adapted domain randomization before any real-world deployment. This pre-training phase prepares the agent for real-world conditions without requiring extensive real-world data collection, significantly reducing overall training time and resources while maintaining high performance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11501167B2Learning domain randomization distributions for transfer learning
Publication Date: 2022.11.15 HUAWEI TECH CANADA CO LTD
  • US11501167B2 patent drawing
  • US11501167B2 patent drawing
  • US11501167B2 patent drawing

AI summary

Method or system for reinforcement learning that simultaneously learns a DR distribution ϕ while optimizing an agent policy Π to maximize performance over the learned DR distribution; method or system for training a learning agent using data synthesized by a simulator based on both a performance of the learning agent and a range of parameters present in the synthesized data.