Robot Policy Switching Using Pose Uncertainty for Sample-Efficient Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current robotic training methods require significant memory, time, and computing resources, often necessitating large amounts of costly or unavailable training data, leading to brittle and unstable systems that struggle to converge on task solutions effectively.

Innovation Solution

A robotic control algorithm that combines model-based and model-free methods, leveraging the efficiency of model-based movement in free spaces while using reinforcement learning to adapt and learn from environment interactions, with a perception system predicting pose uncertainties to fuse policies and overcome inaccuracies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If model-free reinforcement learning is used for robotic training, then adaptability to environment interactions is improved, but sample efficiency deteriorates requiring significant memory, time, and computing resources

Engineering Contradiction:
Improveadaptability to environment interactionsVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent combines model-free reinforcement learning with model-based prediction methods, creating a hybrid system where the model-based component provides efficient predictions for common scenarios while the model-free component handles novel situations, thereby reducing overall training time while maintaining adaptability

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent pre-trains a model-based predictor on simulation data before deploying to real-world environments. This preliminary action allows the system to start with pre-learned patterns, reducing the amount of additional training needed in the real environment and thus reducing total training time

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If model-free reinforcement learning is used for robotic training, then adaptability to environment interactions is improved, but computing resources deteriorate requiring extreme amounts of training data

Engineering Contradiction:
Improveadaptability to environment interactionsVSAvoidtraining data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent merges model-based prediction with model-free reinforcement learning, allowing the model-based component to handle predictable scenarios with minimal data while the model-free component learns from actual interactions, reducing the total training data requirement while maintaining adaptability

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses simulation environments to generate synthetic training data that copies real-world physics and dynamics. This simulated data can be generated in large quantities without physical resources, reducing the need for extensive real-world training data collection

Inventive Principle:
Principle #26Copying

3Productivity

If traditional robotic training methods are used, then task learning is achieved, but system stability deteriorates resulting in excessively brittle systems that do not reliably converge

Engineering Contradiction:
Improvetask learning capabilityVSAvoidsystem stability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where the model-based predictor is continuously refined based on the difference between predicted and actual outcomes. This feedback loop stabilizes the system by correcting prediction errors over time while maintaining the ability to learn new tasks

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent employs dynamic policy adjustment where the balance between model-based and model-free components is adaptively modified during training. This dynamic approach allows the system to stabilize by relying more on the reliable model-based component while still maintaining task learning capability through the model-free component

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12109701B2Guided uncertainty-aware policy optimization: combining model-free and model-based strategies for sample-efficient learning
Publication Date: 2024.10.08 NVIDIA CORP
  • US12109701B2 patent drawing
  • US12109701B2 patent drawing
  • US12109701B2 patent drawing

AI summary

A robot is controlled using a combination of model-based and model-free control methods. In some examples, the model-based method uses a physical model of the environment around the robot to guide the robot. The physical model is oriented using a perception system such as a camera. Characteristics of the perception system may be are used to determine an uncertainty for the model. Based at least in part on this uncertainty, the system transitions from the model-based method to a model-free method where, in some embodiments, information provided directly from the perception system is used to direct the robot without reliance on the physical model.