Robot Action Correction via Human Feedback Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Robots may perform actions incorrectly due to inaccurate models and dynamic environments, often failing to recognize these errors themselves.

Innovation Solution

A method to generate correction instances based on human input, which are used to train neural network models across multiple robots, allowing for model revision and real-time adjustment of robot actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If robot models are used to determine actions in dynamic environments, then automation and productivity are improved, but measurement precision and reliability deteriorate due to model inaccuracies and environmental variability

Engineering Contradiction:
Improverobot action performanceVSAvoidaction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements feedback by collecting human corrections of robot actions and using these corrections to retrain neural network models. The feedback loop continuously improves model accuracy by comparing robot predictions with human expert judgments, thereby resolving the contradiction between automation and precision.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system changes the parameters of neural network models through iterative retraining with correction instances. By updating model parameters based on human feedback from multiple robots operating in diverse environments, the system adapts to environmental variability while maintaining high productivity.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If robot models are trained on limited data, then device complexity is reduced, but adaptability to diverse environments and robustness deteriorate

Engineering Contradiction:
Improvemodel complexityVSAvoidenvironmental adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The system creates universal neural network models that can operate across multiple diverse environments by collecting correction instances from multiple robots in different geographic locations and environments. The aggregated training data makes the models universally applicable rather than environment-specific.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary action by proactively collecting and aggregating correction instances from multiple robots before deploying updated models. This advance data collection and model retraining ensures robots are prepared for diverse environments before encountering new situations.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If human corrections are collected from multiple robots in disparate locations, then model robustness and adaptability are improved, but loss of time and communication overhead increase

Engineering Contradiction:
Improvemodel robustnessVSAvoidmodel update time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system merges correction instances from multiple distributed robots into a centralized training process. By combining data from various sources and aggregating correction instances, the system achieves improved model robustness while managing the time cost through efficient data consolidation and batch processing.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP4219088B1Determining and utilizing corrections to robot actions
Publication Date: 2026.03.18 GDM HOLDING LLC
  • EP4219088B1 patent drawingFigure 1
  • EP4219088B1 patent drawingFigure 2A
  • EP4219088B1 patent drawingFigure 2B

AI summary

Methods, apparatus, and computer-readable media for determining and utilizing human corrections to robot actions. In some implementations, in response to determining a human correction of a robot action, a correction instance is generated that includes sensor data, captured by one or more sensors of the robot, that is relevant to the corrected action. The correction instance can further include determined incorrect parameter(s) utilized in performing the robot action and/or correction information that is based on the human correction. The correction instance can be utilized to generate training example(s) for training one or model(s), such as neural network model(s), corresponding to those used in determining the incorrect parameter(s). In various implementations, the training is based on correction instances from multiple robots. After a revised version of a model is generated, the revised version can thereafter be utilized by one or more of the multiple robots.