Robot Action Imitation Using RGB Keypoint Pose Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing action imitation methods for humanoid robots are either costly due to the need for high-precision depth cameras or cumbersome wearable devices, limiting their application and popularity.
Innovation Solution
A computer-implemented method using a pre-trained convolutional neural network to process RGB images from an ordinary camera, calculating key point positions and rotational angles of linkages, allowing robots to imitate human actions without high-precision depth cameras, thereby reducing costs and expanding application scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If wearable control devices are used to collect motion information, then measurement precision is improved, but device complexity and cost increase
Solution Approach 1:
The patent uses video images as a copy of human motion to extract motion information, replacing the need for wearable control devices. The video processing module captures and processes visual copies of human actions to derive joint angle information, achieving measurement precision without physical wearables.
Solution Approach 2:
The patent replaces the mechanical wearable control device system with a video-based optical system. Instead of using physical sensors attached to the body, the system uses video cameras and image processing algorithms to extract motion data, substituting mechanical measurement with optical-field measurement.
2Measurement precision
If high-precision depth cameras are used for vision-based action imitation, then measurement precision is improved, but cost increases
Solution Approach 1:
The patent uses ordinary video cameras instead of expensive high-precision depth cameras. The system processes standard video images through specialized algorithms to extract depth and motion information, replacing costly specialized hardware with cheaper, widely available components.
Solution Approach 2:
The patent replaces the optical depth-sensing system with a video-based system that uses 2D image processing to infer 3D motion. Instead of using depth cameras that directly measure distance, the system uses video frames and computational algorithms to reconstruct motion information, substituting direct optical measurement with computational inference.
3Reliability
If wearable control devices are used, then action imitation accuracy is improved, but ease of operation deteriorates due to cumbersome wearing and assembling processes
Solution Approach 1:
The patent enables the system to automatically capture and process video information without requiring user intervention for device setup. The video processing module automatically tracks human motion from video frames, extracting joint angle information without manual calibration or device assembly, making the system self-sufficient and user-friendly.
Solution Approach 2:
The patent uses video copies of human motion to directly drive robot imitation without requiring physical interaction. The system captures video of human actions and processes these visual copies to generate corresponding robot motions, eliminating the need for users to wear or assemble any equipment.
4Measurement precision
If wearable control devices or high-precision depth cameras are used, then measurement precision is improved, but adaptability deteriorates due to limited application scenarios
Solution Approach 1:
The patent creates a universal action imitation system that can handle various types of human motions using a single video-based approach. The system can process different actions (walking, gesturing, manipulating objects) through the same video processing pipeline, making it adaptable to diverse application scenarios without requiring specialized equipment for each scenario.
Solution Approach 2:
The patent replaces specialized measurement systems with a general-purpose video-based system. By using standard video cameras and computational algorithms, the system achieves broad adaptability across different environments and applications, from industrial settings to domestic environments, without being constrained by specialized hardware requirements.
Data Source
AI summary
The present disclosure provides an action imitation method as well as a robot and a computer readable storage medium using the same. The method includes: collecting a plurality of action images of a to-be-imitated object; processing the action images through a pre-trained convolutional neural network to obtain a position coordinate set of position coordinates of a plurality of key points of each of the action images; calculating a rotational angle of each of the linkages of the to-be-imitated object based on the position coordinate sets of the action images; and controlling a robot to move according to the rotational angle of each of the linkages of the to-be-imitated object. In the above-mentioned manner, the rotational angle of each linkage of the to-be-imitated object can be obtained by just analyzing and processing the images collected by an ordinary camera without the help of high-precision depth camera.


