Robotic Control Policy Learning via Simulation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robotic control policies are challenging to design and scale for real-world applications due to the complexity and variability of environments, often requiring extensive hand-coding and limited flexibility, which can lead to inefficiencies and potential damage during training.

Innovation Solution

The use of machine learning techniques to generate and refine robotic control policies through simulated environments, allowing for iterative training and adaptation in both virtual and real-world settings, leveraging reinforcement learning and evolution strategies to improve robustness and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional hand-coding methods are used to design robotic control policies, then the system can operate in predictable environments, but the system lacks flexibility and adaptability when encountering unanticipated situations or complex tasks

Engineering Contradiction:
ImproveflexibilityVSAvoidcomplexity of control policy design
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates virtual copies of the real-world environment through high-fidelity simulation. The simulation replicates physical properties, sensor behaviors, and task conditions, allowing the robotic system to train extensively in a virtual副本 without real-world risks. This copying approach enables the system to learn complex adaptive behaviors that would be difficult to hand-code directly.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary training actions in the virtual environment before deploying to the real world. By pre-training the robotic system extensively in simulation across diverse scenarios and edge cases, the system acquires adaptive capabilities in advance, reducing the need for complex hand-coded control policies when deployed.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If extensive training is conducted in real-world environments to improve robotic performance, then the system achieves higher accuracy and robustness, but the risk of damage and training time increases significantly

Engineering Contradiction:
ImproverobustnessVSAvoidpotential damage during training
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent uses virtual copying to replicate the real-world environment in simulation, allowing all extensive training to occur in the virtual副本. This eliminates physical damage risks while maintaining training effectiveness through high-fidelity environmental replication that preserves the physics and sensor characteristics of the real world.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system cushions against potential damage by conducting all preliminary training in a risk-free virtual environment. The simulation acts as a protective layer, allowing the system to fail and learn from mistakes without real-world consequences, before any real-world deployment occurs.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

3Adaptability or versatility

If traditional control policies are designed to handle diverse real-world variations, then the system can operate in variable environments, but the design and scaling become extremely challenging and time-consuming

Engineering Contradiction:
Improveability to handle environment variabilityVSAvoiddesign and scaling time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent creates a virtual副本 of the real-world environment that captures environmental variability and complexity. By training in this replicated environment, the system learns to handle diverse conditions without requiring manual design for each scenario, dramatically reducing design and scaling time while maintaining adaptability.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system uses self-service learning through autonomous training in simulation. Instead of requiring engineers to hand-code control policies for each environmental variation, the robotic system independently learns adaptive behaviors through self-directed exploration and reinforcement learning in the virtual environment, eliminating manual design bottlenecks.

Inventive Principle:
Principle #25Self-service

4Adaptability or versatility

If machine learning techniques are used to generate control policies through simulation, then the system achieves improved flexibility and accuracy, but the computational resources and training complexity increase

Engineering Contradiction:
ImproveflexibilityVSAvoidtraining system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses virtual copying to create a simulation environment that replicates real-world physics and sensor behavior. This approach enables machine learning training with high flexibility and accuracy while managing computational complexity by performing all intensive training in the virtual副本, avoiding the need for complex real-world training infrastructure.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10792810B1Artificial intelligence system for learning robotic control policies
Publication Date: 2020.10.06 AMAZON TECH INC
  • US10792810B1 patent drawing
  • US10792810B1 patent drawing
  • US10792810B1 patent drawing

AI summary

A machine learning system builds and uses computer models for controlling robotic performance of a task. Such computer models may be first trained using feedback on computer simulations of the robot performing the task, and then refined using feedback on real-world trials of the robot performing the task. Some examples of the computer models can be trained to automatically evaluate robotic task performance and provide the feedback. This feedback can be used by a machine learning system, for example an evolution strategies system or reinforcement learning system, to generate and refine the controller.