Robot Action Imitation Using Monocular 3D Pose Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing humanoid robots require a to-be-imitated object to wear light-sensitive elements and are restricted by their working environment, limiting their ability to perform action imitation tasks, especially in outdoor settings.

Innovation Solution

A method using monocular cameras to collect two-dimensional images from different view angles, predicting three-dimensional coordinates of key points on the to-be-imitated object, and generating action control instructions for a robot to imitate actions without the need for light-sensitive devices, enabling action imitation in various environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple cameras and light-sensitive elements are used to accurately locate joints, then measurement precision is improved, but device complexity increases and ease of operation deteriorates

Engineering Contradiction:
Improvejoint location precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the light-sensitive elements from the system, replacing them with a monocular camera that captures images of the to-be-imitated object directly. This eliminates the need for additional wearable devices while maintaining the capability to track and locate key points through image processing algorithms.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses a monocular camera to capture visual images of the to-be-imitated object, creating a visual copy that can be processed to extract key point coordinates. This visual copying approach replaces the direct electromagnetic signal capture method, simplifying the system while preserving measurement functionality.

Inventive Principle:
Principle #26Copying

2Measurement precision

If multiple cameras and light-sensitive elements are required, then measurement precision is improved, but ease of operation worsens

Engineering Contradiction:
Improvejoint location precisionVSAvoidoperation convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent removes the requirement for light-sensitive elements to be worn by the to-be-imitated object. Instead, it uses a monocular camera to capture images and process them to obtain key point coordinates, significantly improving ease of operation as no special equipment needs to be worn.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system allows the to-be-imitated object to be captured using ordinary clothing or no special equipment at all. The monocular camera and processing algorithms serve the function of tracking key points without requiring the object to provide any special service or wear any devices.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If light-sensitive elements are worn by the to-be-imitated object, then measurement precision is improved, but adaptability deteriorates due to environmental restrictions

Engineering Contradiction:
Improveaction recognition precisionVSAvoidenvironmental adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent extracts and eliminates the light-sensitive elements from the system, replacing them with a monocular camera that captures visual images. This allows the system to operate in various lighting conditions and environments without requiring special wearable equipment, significantly improving environmental adaptability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The monocular camera system serves multiple functions: it captures images for key point detection, works in various lighting conditions, and does not require the to-be-imitated object to wear any special equipment. This universal approach enhances the system's adaptability across different environments.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11940774B2Action imitation method and robot and computer readable storage medium using the same
Publication Date: 2024.03.26 UBTECH ROBOTICS CORP LTD
  • US11940774B2 patent drawing
  • US11940774B2 patent drawing
  • US11940774B2 patent drawing

AI summary

The present disclosure provides action imitation method as well as a robot and a computer readable storage medium using the same. The method includes: collecting at least a two-dimensional image of a to-be-imitated object; obtaining two-dimensional coordinates of each key point of the to-be-imitated object in the two-dimensional image and a pairing relationship between the key points of the to-be-imitated object; converting the two-dimensional coordinates of the key points of the to-be-imitated object in the two-dimensional image into space three-dimensional coordinates corresponding to the key points of the to-be-imitated object through a pre-trained first neural network model, and generating an action control instruction of a robot based on the space three-dimensional coordinates corresponding to the key points of the to-be-imitated object and the pairing relationship between the key points, where the action control instruction is for controlling the robot to imitate an action of the to-be-imitated object.