3D Pose Estimation Using Synthetic Depth Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for determining the 3D pose of objects in RGB images require annotated images, which are time-consuming and expensive to create, and often impractical due to the complexity of information in RGB images.

Innovation Solution

A method that maps features from unannotated RGB images to a depth domain using a trained pose estimator network, which is optimized with depth maps generated from RGB-D cameras or CAD models, allowing for 3D pose determination without annotated RGB images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If annotated RGB images are used for training pose estimation systems, then measurement precision of 3D pose is improved, but loss of time and manufacturing cost increase due to manual annotation requirements

Engineering Contradiction:
Improve3D pose determination accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates synthetic depth maps as copies of real-world scenes using CAD models and rendering algorithms. These synthetic depth maps serve as training data substitutes, eliminating the need for manual annotation of real RGB images while providing sufficient geometric information for accurate 3D pose estimation

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces depth maps as an intermediary representation between RGB images and 3D pose parameters. By training on depth maps (which can be synthetically generated) and then applying the trained model to RGB images, the system bridges the gap between easy-to-generate training data and the target application domain

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If annotated RGB images are used for training pose estimation systems, then measurement precision of 3D pose is improved, but loss of money increases due to annotation costs

Engineering Contradiction:
Improve3D pose determination accuracyVSAvoidannotation cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent creates synthetic depth maps as copies of real-world scenes using CAD models and rendering algorithms. These synthetic depth maps serve as training data substitutes, eliminating the need for manual annotation of real RGB images while providing sufficient geometric information for accurate 3D pose estimation

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent uses computationally inexpensive synthetic depth maps generated from CAD models as disposable training data. These synthetic examples can be generated in large quantities at minimal cost, replacing expensive manual annotation processes while providing adequate training signal for the neural network

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Ease of manufacture

If depth maps are used instead of RGB images for training, then ease of annotation is improved, but loss of information occurs due to reduced visual detail

Engineering Contradiction:
Improveannotation easeVSAvoidvisual information
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The patent extracts only the essential geometric information needed for 3D pose estimation from the full RGB image data, representing it in the simplified depth map domain. This extraction removes unnecessary visual details (colors, textures, lighting) while retaining the critical depth and shape information required for accurate pose determination

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11532094B2Systems and methods for three-dimensional pose determination
Publication Date: 2022.12.20 QUALCOMM TECHNOLOGIES INC
  • US11532094B2 patent drawing
  • US11532094B2 patent drawing
  • US11532094B2 patent drawing

AI summary

A method is described. The method includes mapping features extracted from an unannotated red-green-blue (RGB) image of the object to a depth domain. The method further includes determining a three-dimensional (3D) pose of the object by providing the features mapped from the unannotated RGB image of the object to the depth domain to a trained pose estimator network.