Robot Action Imitation Using Monocular 3D Pose Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing humanoid robots require a to-be-imitated object to wear light-sensitive elements and are restricted by their working environment, limiting their ability to perform action imitation tasks, especially in outdoor settings.
Innovation Solution
A method using monocular cameras to collect two-dimensional images from different view angles, predicting three-dimensional coordinates of key points on the to-be-imitated object, and generating action control instructions for a robot to imitate actions without the need for light-sensitive devices, enabling action imitation in various environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple cameras and light-sensitive elements are used to accurately locate joints, then measurement precision is improved, but device complexity increases and ease of operation deteriorates
Solution Approach 1:
The patent extracts and removes the light-sensitive elements from the system, replacing them with a monocular camera that captures images of the to-be-imitated object directly. This eliminates the need for additional wearable devices while maintaining the capability to track and locate key points through image processing algorithms.
Solution Approach 2:
The patent uses a monocular camera to capture visual images of the to-be-imitated object, creating a visual copy that can be processed to extract key point coordinates. This visual copying approach replaces the direct electromagnetic signal capture method, simplifying the system while preserving measurement functionality.
2Measurement precision
If multiple cameras and light-sensitive elements are required, then measurement precision is improved, but ease of operation worsens
Solution Approach 1:
The patent removes the requirement for light-sensitive elements to be worn by the to-be-imitated object. Instead, it uses a monocular camera to capture images and process them to obtain key point coordinates, significantly improving ease of operation as no special equipment needs to be worn.
Solution Approach 2:
The system allows the to-be-imitated object to be captured using ordinary clothing or no special equipment at all. The monocular camera and processing algorithms serve the function of tracking key points without requiring the object to provide any special service or wear any devices.
3Measurement precision
If light-sensitive elements are worn by the to-be-imitated object, then measurement precision is improved, but adaptability deteriorates due to environmental restrictions
Solution Approach 1:
The patent extracts and eliminates the light-sensitive elements from the system, replacing them with a monocular camera that captures visual images. This allows the system to operate in various lighting conditions and environments without requiring special wearable equipment, significantly improving environmental adaptability.
Solution Approach 2:
The monocular camera system serves multiple functions: it captures images for key point detection, works in various lighting conditions, and does not require the to-be-imitated object to wear any special equipment. This universal approach enhances the system's adaptability across different environments.
Data Source
AI summary
The present disclosure provides action imitation method as well as a robot and a computer readable storage medium using the same. The method includes: collecting at least a two-dimensional image of a to-be-imitated object; obtaining two-dimensional coordinates of each key point of the to-be-imitated object in the two-dimensional image and a pairing relationship between the key points of the to-be-imitated object; converting the two-dimensional coordinates of the key points of the to-be-imitated object in the two-dimensional image into space three-dimensional coordinates corresponding to the key points of the to-be-imitated object through a pre-trained first neural network model, and generating an action control instruction of a robot based on the space three-dimensional coordinates corresponding to the key points of the to-be-imitated object and the pairing relationship between the key points, where the action control instruction is for controlling the robot to imitate an action of the to-be-imitated object.


