Pose Determination via Semantic Segmentation and 3D Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing camera registration methods in augmented reality, autonomous driving, and robotics face challenges in accurately determining the pose of an image capture device under changing conditions, such as illumination and construction, due to the difficulty in matching 2D images with pre-registered 3D scene points, especially in urban environments with untextured 2.5D models.
Innovation Solution
A method that involves semantic segmentation of captured images to generate segmented images, using a 3D tracker to estimate an initial pose, and generating multiple 3D renderings based on this pose to select the alignment that matches the segmented image, thereby correcting tracker drift without requiring additional reference images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If pre-registered image collections are used for pose determination, then the system can provide a reference framework for matching, but the method fails under changing conditions such as illumination, season, or construction work
Solution Approach 1:
The patent segments the image matching process into two distinct stages: coarse localization using pre-registered image collections to obtain an initial pose estimate, and fine alignment using semantic segmentation to match 3D rendered images with 2D captured images. This segmentation allows the system to leverage the stability of pre-registered collections while adapting to changing conditions through semantic feature matching that is invariant to illumination and seasonal variations.
Solution Approach 2:
The patent performs preliminary pose estimation using pre-registered image collections before conducting the final semantic segmentation-based alignment. This preliminary action provides an initial guess that constrains the search space for the subsequent fine alignment process, improving both efficiency and accuracy while maintaining adaptability to changing conditions.
2Adaptability or versatility
If feature matching is performed under changing conditions, then the system can handle various environments, but matching accuracy deteriorates due to illumination, season, or construction work
Solution Approach 1:
The patent introduces semantic segmentation as an intermediary process between the captured image and the pre-registered 3D model. Instead of directly matching 2D image features with 3D point cloud features, the system first segments the captured image into semantic regions, then renders corresponding 3D images from the model, and finally aligns these segmented representations. This intermediary step eliminates the negative effects of illumination and seasonal changes on direct feature matching.
Solution Approach 2:
The patent creates a synthetic copy of the captured image by rendering a 3D image from the pre-registered model using the estimated pose parameters. This rendered 3D image serves as a reference that can be directly compared with the segmented captured image, providing a consistent reference framework that is invariant to environmental changes while maintaining geometric accuracy.
3Measurement precision
If 2D image points are matched to 3D scene points, then pose can be computed, but the method is challenged by untextured 2.5D models in urban environments
Solution Approach 1:
The patent transitions from 2D image space to 3D rendering space by generating 3D rendered images from the untextured 2.5D model using the estimated pose parameters. This dimensionality change allows the system to leverage the geometric information in the 3D model while comparing it with the 2D captured image in a unified 3D rendering framework, making the matching process invariant to the lack of texture information in urban environments.
Data Source
AI summary
A method determines a pose of an image capture device. The method includes accessing an image of a scene captured by the image capture device. A semantic segmentation of the image is performed, to generate a segmented image. An initial pose of the image capture device is generated using a three-dimensional (3D) tracker. A plurality of 3D renderings of the scene are generated, each of the plurality of 3D renderings corresponding to one of a plurality of poses chosen based on the initial pose. A pose is selected from the plurality of poses, such that the 3D rendering corresponding to the selected pose aligns with the segmented image.


