Robot Action Correction Learning From Human Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Robots often perform actions incorrectly due to inaccurate models and dynamic environments, and may not recognize these errors, leading to incorrect parameter determination and subsequent actions.

Innovation Solution

A method to generate correction instances using human input, which include sensor data and incorrect parameter information, to train and update neural network models, allowing robots to adapt and improve their performance based on human corrections, even across disparate geographic locations and environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If robots use current models to determine action parameters, then they can perform actions autonomously, but the actions may be incorrect due to model inaccuracy and dynamic environments

Engineering Contradiction:
Improveautonomous action performanceVSAvoidaction correctness
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The system implements feedback by collecting human corrections when robots perform incorrect actions. These corrections are fed back into the training process to update and improve the models, creating a closed-loop system that continuously learns from errors and improves action correctness while maintaining autonomous operation

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system enables self-service by allowing robots to autonomously collect their own incorrect action data and human corrections, automatically generate training examples, and participate in model retraining without requiring manual intervention for each correction, thus improving reliability while maintaining automation

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If robots collect correction instances from multiple sources, then model training data increases, but system complexity increases due to disparate robots and environments

Engineering Contradiction:
Improvetraining data volumeVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system achieves universality by creating a standardized correction instance format and training example structure that can be universally applied across multiple disparate robots operating in different environments. This allows training data from diverse sources to be integrated without proportionally increasing system complexity, as the same processing pipeline handles all correction instances

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system manages complexity through parameter changes by abstracting the essential elements of corrections into standardized parameters and features. This allows the system to handle varying data from different robots and environments by transforming them into a common parameter space, increasing training data utility without linearly increasing system complexity

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If robots immediately apply human corrections to local performance, then action accuracy improves, but the ability to learn from aggregated corrections across multiple robots is reduced

Engineering Contradiction:
Improveaction accuracyVSAvoidlearning capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system applies preliminary action by immediately applying human corrections to local robot performance to ensure accurate action execution in the short term, while simultaneously preparing and submitting correction instances for aggregated model retraining. This dual approach maintains action accuracy while preserving the ability to learn from aggregated corrections across multiple robots

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements dynamics by allowing robot performance characteristics to change over time - initially applying corrections locally for immediate accuracy, then progressively incorporating aggregated learnings from multiple robots through periodic model retraining. This dynamic approach balances immediate action accuracy with long-term learning capability

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12064876B2Determining and utilizing corrections to robot actions
Publication Date: 2024.08.20 GDM HOLDING LLC
  • US12064876B2 patent drawing
  • US12064876B2 patent drawing
  • US12064876B2 patent drawing

AI summary

Methods, apparatus, and computer-readable media for determining and utilizing human corrections to robot actions. In some implementations, in response to determining a human correction of a robot action, a correction instance is generated that includes sensor data, captured by one or more sensors of the robot, that is relevant to the corrected action. The correction instance can further include determined incorrect parameter(s) utilized in performing the robot action and/or correction information that is based on the human correction. The correction instance can be utilized to generate training example(s) for training one or model(s), such as neural network model(s), corresponding to those used in determining the incorrect parameter(s).