Pose Determination via Semantic Segmentation and 3D Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing camera registration methods in augmented reality, autonomous driving, and robotics face challenges in accurately determining the pose of an image capture device under changing conditions, such as illumination and construction, due to the difficulty in matching 2D images with pre-registered 3D scene points, especially in urban environments with untextured 2.5D models.

Innovation Solution

A method that involves semantic segmentation of captured images to generate segmented images, using a 3D tracker to estimate an initial pose, and generating multiple 3D renderings based on this pose to select the alignment that matches the segmented image, thereby correcting tracker drift without requiring additional reference images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If pre-registered image collections are used for pose determination, then the system can provide a reference framework for matching, but the method fails under changing conditions such as illumination, season, or construction work

Engineering Contradiction:
Improvepose determination reliabilityVSAvoidadaptability to changing conditions
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the image matching process into two distinct stages: coarse localization using pre-registered image collections to obtain an initial pose estimate, and fine alignment using semantic segmentation to match 3D rendered images with 2D captured images. This segmentation allows the system to leverage the stability of pre-registered collections while adapting to changing conditions through semantic feature matching that is invariant to illumination and seasonal variations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary pose estimation using pre-registered image collections before conducting the final semantic segmentation-based alignment. This preliminary action provides an initial guess that constrains the search space for the subsequent fine alignment process, improving both efficiency and accuracy while maintaining adaptability to changing conditions.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If feature matching is performed under changing conditions, then the system can handle various environments, but matching accuracy deteriorates due to illumination, season, or construction work

Engineering Contradiction:
Improveenvironmental adaptabilityVSAvoidfeature matching precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces semantic segmentation as an intermediary process between the captured image and the pre-registered 3D model. Instead of directly matching 2D image features with 3D point cloud features, the system first segments the captured image into semantic regions, then renders corresponding 3D images from the model, and finally aligns these segmented representations. This intermediary step eliminates the negative effects of illumination and seasonal changes on direct feature matching.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a synthetic copy of the captured image by rendering a 3D image from the pre-registered model using the estimated pose parameters. This rendered 3D image serves as a reference that can be directly compared with the segmented captured image, providing a consistent reference framework that is invariant to environmental changes while maintaining geometric accuracy.

Inventive Principle:
Principle #26Copying

3Measurement precision

If 2D image points are matched to 3D scene points, then pose can be computed, but the method is challenged by untextured 2.5D models in urban environments

Engineering Contradiction:
Improvepose computation accuracyVSAvoidfeature detectability
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent transitions from 2D image space to 3D rendering space by generating 3D rendered images from the untextured 2.5D model using the estimated pose parameters. This dimensionality change allows the system to leverage the geometric information in the 3D model while comparing it with the 2D captured image in a unified 3D rendering framework, making the matching process invariant to the lack of texture information in urban environments.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10546387B2Pose determination with semantic segmentation
Publication Date: 2020.01.28 QUALCOMM INC
  • US10546387B2 patent drawing
  • US10546387B2 patent drawing
  • US10546387B2 patent drawing

AI summary

A method determines a pose of an image capture device. The method includes accessing an image of a scene captured by the image capture device. A semantic segmentation of the image is performed, to generate a segmented image. An initial pose of the image capture device is generated using a three-dimensional (3D) tracker. A plurality of 3D renderings of the scene are generated, each of the plurality of 3D renderings corresponding to one of a plurality of poses chosen based on the initial pose. A pose is selected from the plurality of poses, such that the 3D rendering corresponding to the selected pose aligns with the segmented image.