Robot Action Imitation Using RGB Keypoint Pose Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing action imitation methods for humanoid robots are either costly due to the need for high-precision depth cameras or cumbersome wearable devices, limiting their application and popularity.

Innovation Solution

A computer-implemented method using a pre-trained convolutional neural network to process RGB images from an ordinary camera, calculating key point positions and rotational angles of linkages, allowing robots to imitate human actions without high-precision depth cameras, thereby reducing costs and expanding application scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If wearable control devices are used to collect motion information, then measurement precision is improved, but device complexity and cost increase

Engineering Contradiction:
Improvemotion information accuracyVSAvoidwearable equipment complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses video images as a copy of human motion to extract motion information, replacing the need for wearable control devices. The video processing module captures and processes visual copies of human actions to derive joint angle information, achieving measurement precision without physical wearables.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical wearable control device system with a video-based optical system. Instead of using physical sensors attached to the body, the system uses video cameras and image processing algorithms to extract motion data, substituting mechanical measurement with optical-field measurement.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If high-precision depth cameras are used for vision-based action imitation, then measurement precision is improved, but cost increases

Engineering Contradiction:
Improvedepth measurement accuracyVSAvoidproduction cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent uses ordinary video cameras instead of expensive high-precision depth cameras. The system processes standard video images through specialized algorithms to extract depth and motion information, replacing costly specialized hardware with cheaper, widely available components.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Solution Approach 2:

The patent replaces the optical depth-sensing system with a video-based system that uses 2D image processing to infer 3D motion. Instead of using depth cameras that directly measure distance, the system uses video frames and computational algorithms to reconstruct motion information, substituting direct optical measurement with computational inference.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If wearable control devices are used, then action imitation accuracy is improved, but ease of operation deteriorates due to cumbersome wearing and assembling processes

Engineering Contradiction:
Improveaction imitation accuracyVSAvoiduser experience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent enables the system to automatically capture and process video information without requiring user intervention for device setup. The video processing module automatically tracks human motion from video frames, extracting joint angle information without manual calibration or device assembly, making the system self-sufficient and user-friendly.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses video copies of human motion to directly drive robot imitation without requiring physical interaction. The system captures video of human actions and processes these visual copies to generate corresponding robot motions, eliminating the need for users to wear or assemble any equipment.

Inventive Principle:
Principle #26Copying

4Measurement precision

If wearable control devices or high-precision depth cameras are used, then measurement precision is improved, but adaptability deteriorates due to limited application scenarios

Engineering Contradiction:
Improvemotion data accuracyVSAvoidapplication scenario range
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal action imitation system that can handle various types of human motions using a single video-based approach. The system can process different actions (walking, gesturing, manipulating objects) through the same video processing pipeline, making it adaptable to diverse application scenarios without requiring specialized equipment for each scenario.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces specialized measurement systems with a general-purpose video-based system. By using standard video cameras and computational algorithms, the system achieves broad adaptability across different environments and applications, from industrial settings to domestic environments, without being constrained by specialized hardware requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11850747B2Action imitation method and robot and computer readable medium using the same
Publication Date: 2023.12.26 UBTECH ROBOTICS CORP LTD
  • US11850747B2 patent drawing
  • US11850747B2 patent drawing
  • US11850747B2 patent drawing

AI summary

The present disclosure provides an action imitation method as well as a robot and a computer readable storage medium using the same. The method includes: collecting a plurality of action images of a to-be-imitated object; processing the action images through a pre-trained convolutional neural network to obtain a position coordinate set of position coordinates of a plurality of key points of each of the action images; calculating a rotational angle of each of the linkages of the to-be-imitated object based on the position coordinate sets of the action images; and controlling a robot to move according to the rotational angle of each of the linkages of the to-be-imitated object. In the above-mentioned manner, the rotational angle of each linkage of the to-be-imitated object can be obtained by just analyzing and processing the images collected by an ordinary camera without the help of high-precision depth camera.