Robot Performance Recreation Using Vision-Guided Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional robots require complex manual programming and extensive training data for human-like tasks, making them costly and time-consuming, and conventional deep learning methods struggle with large environmental state-action pairs.

Innovation Solution

A method combining machine vision and reinforcement learning enables robots to observe and recreate human performances by identifying objects and tracking human appendages, generating a policy to maximize a cumulative award through iterative learning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional manual programming is used for robotic mimicking, then task-specific performance can be achieved, but device complexity and cost increase significantly

Engineering Contradiction:
Improvetask-specific performanceVSAvoidprogramming complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The robot performs self-learning through reinforcement learning algorithms, automatically acquiring task-specific performance without requiring complex manual programming. The system learns optimal policies through iterative trial-and-error in the environment, eliminating the need for extensive human programming effort while maintaining reliable task execution

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces traditional mechanical programming approaches with intelligent software-based reinforcement learning. Instead of manually configuring robot behavior through complex programming, the system uses AI algorithms that automatically learn task performance through environmental interaction and reward-based feedback

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If conventional deep learning methods are used for environmental perception, then object recognition capability is improved, but computational complexity increases due to large state-action pairs

Engineering Contradiction:
Improveobject recognition accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex environmental state-space into manageable components through hierarchical reinforcement learning. The system divides large state-action pairs into smaller sub-tasks and sub-states, making computational processing more efficient while maintaining accurate object recognition and environmental perception capabilities

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts perception parameters and state representations based on task requirements and environmental context. By changing the granularity and focus of environmental representation, the system maintains high object recognition accuracy while reducing the computational burden of processing all possible state-action pairs

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12591229B2Performance recreation system
Publication Date: 2026.03.31 AIVOT LLC
  • US12591229B2 patent drawing
  • US12591229B2 patent drawing
  • US12591229B2 patent drawing

AI summary

The present disclosure generally relates to performance recreation, and in particular, the recreation of observed human performance using reinforcement learning. In this regard, a first object is identified from a plurality of objects. The manipulation of the first object is tracked from a first position to a second position. A characterization of the manipulation is generated. A policy that controls a mechanical gripper to recreate the manipulation is generated based on an iteratively increasing cumulative award. The mechanical gripper iteratively recreates the manipulation to increase a cumulative award with each recreation.