Robotic Control Policy Training via Simulation Parameterization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing processes for creating free-moving robots are time-intensive and processing-intensive, often requiring new control processes for each hardware change, making it difficult to design and improve robotic systems quickly.

Innovation Solution

A method involving parameterization of inputs to robotic devices, generating samples within defined value ranges, training control policies using reinforcement learning, and deploying these policies to onboard controllers for stable and creative robotic system design.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If model-based processes are used to create free moving robots, then control accuracy is improved, but development time and processing time increase significantly

Engineering Contradiction:
Improvecontrol accuracyVSAvoiddevelopment time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses simulation environments to create virtual copies of robotic devices for training control policies. These simulated models allow extensive testing and iteration without physical hardware, dramatically reducing development time while maintaining control accuracy through realistic physics modeling.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent trains control policies in advance using simulation environments before deploying them to physical robots. This preliminary training in virtual space allows the system to learn optimal control strategies without requiring time-consuming physical experimentation, thereby reducing overall development time.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If hardware changes are made to robotic devices, then system capabilities are improved, but control processes must be completely redeveloped

Engineering Contradiction:
Improvesystem capabilitiesVSAvoidcontrol process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent parameterizes hardware characteristics in the simulation environment, allowing the control policy to learn across variations in mass, torque, force, speed, and range of motion. This enables the same control framework to adapt to different hardware configurations without requiring complete redevelopment, as the policy learns to handle parameter variations.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a universal control policy through reinforcement learning that can operate across multiple hardware configurations. The simulated training exposes the policy to diverse hardware parameters, enabling it to function effectively on different robotic devices without requiring device-specific control processes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If extensive training data is generated for control policies, then control performance is improved, but processing requirements increase

Engineering Contradiction:
Improvecontrol performanceVSAvoidprocessing power
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

The patent replaces physical experimentation with computational simulation for generating training data. By using virtual environments with physics engines, the system can generate extensive training datasets without the resource constraints of physical hardware, achieving high control performance without proportionally increasing processing power requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250108505A1Rapid design and animation of freely-walking robotic devices
Publication Date: 2025.04.03 DISNEY ENTERPRISES INC
  • US20250108505A1 patent drawing
  • US20250108505A1 patent drawing
  • US20250108505A1 patent drawing

AI summary

A method of training a robotic device includes: parameterizing, via a processing element, an input to the robotic device. The parameterizing comprises defining a range of values of the input. The method further includes generating, via the processing element, a plurality of samples of the parameterized input from within the range of values; training a control policy, via the processing element. The training includes: providing the plurality of samples to the control policy, wherein the control policy is adapted to operate an actuator of the robotic device, and generating, via the processing element, a policy action using the control policy; transmitting the policy action to a robotic model, wherein the robotic model includes a physical model of the robotic device. The method further includes deploying the one or more trained control policies to an on-board controller for the robotic device.