Robot Gripper Learning From Human Demonstration With Reinforcement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional robots require complex manual programming and expensive hardware to perform human-like tasks, making them costly and time-consuming, and conventional deep learning methods face challenges with large training data sets and computational costs.

Innovation Solution

The method combines machine vision and reinforcement learning to enable robots to observe human demonstrations, identify objects, and recreate performances by iteratively improving a cumulative award system, reducing the need for extensive programming and training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional manual programming and hardware are used to enable robots to perform human-like tasks, then task performance capability is improved, but device complexity and cost increase

Engineering Contradiction:
Improvetask performance capabilityVSAvoidprogramming complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The robot performs self-learning by observing human demonstrations through vision systems and automatically generating control policies through reinforcement learning, eliminating the need for manual programming by humans. The system learns tasks autonomously by iterating through demonstration observation, policy generation, and performance evaluation cycles.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual programming mechanisms with automated learning mechanisms. Instead of humans writing code to control robots, the system uses vision-based observation and reinforcement learning algorithms to automatically generate control policies, substituting mechanical programming processes with intelligent learning processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If conventional deep learning methods are used for robot learning, then task automation is improved, but training data requirements and computational costs increase

Engineering Contradiction:
Improvetask automationVSAvoidtraining data volume
Core Design Contradiction:
Extent of automationVSQuantity of substance

Solution Approach 1:

The system performs preliminary action by having a human demonstrate the task once, and the robot captures this demonstration through vision systems. This single demonstration serves as pre-collected training data that the robot then processes through reinforcement learning to learn the task, eliminating the need for extensive training data collection.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The robot uses feedback from performing the learned task and comparing its performance against the observed demonstration to iteratively improve its control policy. This feedback-driven reinforcement learning process enables the robot to learn from limited initial data by continuously refining its understanding through performance evaluation and policy adjustment.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11292129B2Performance recreation system
Publication Date: 2022.04.05 AIVOT LLC
  • US11292129B2 patent drawing
  • US11292129B2 patent drawing
  • US11292129B2 patent drawing

AI summary

The present disclosure generally relates to performance recreation, and in particular, the recreation of observed human performance using reinforcement learning. In this regard, a first object is identified from a plurality of objects. The manipulation of the first object is tracked from a first position to a second position. A characterization of the manipulation is generated. A policy that controls a mechanical gripper to recreate the manipulation is generated based on an iteratively increasing cumulative award. The mechanical gripper iteratively recreates the manipulation to increase a cumulative award with each recreation.