Semantic 3D Reconstruction Using Category Shape Priors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional multiview stereo (MVS) reconstruction methods are limited by the lack of texture, specularities, and wide baselines, leading to sparse and noisy outputs with holes or artifacts, especially in scenarios with few images.

Innovation Solution

A method that incorporates semantic information using learned category-level shape priors and object detection, modeling the object shape as a warped version of a category mean with instance-specific details, and refining the shape using anchor points and photoconsistency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional multiview stereo (MVS) reconstruction methods are used, then the system can reconstruct 3D shapes from multiple images, but the reconstruction quality deteriorates in scenarios with few images, lack of texture, specularities, or wide baselines, leading to sparse and noisy outputs with holes or artifacts

Engineering Contradiction:
Improvereconstruction qualityVSAvoidreconstruction reliability under challenging conditions
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The system performs preliminary actions by learning category-level shape priors and object detection models from extensive training data before actual reconstruction. This pre-learning phase enables the system to have prior knowledge about object shapes and appearances, which is then applied during reconstruction to compensate for challenges like few images, lack of texture, or wide baselines, thereby maintaining reconstruction quality and reliability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces semantic information as an intermediary between the input images and the 3D reconstruction process. By incorporating learned category-level shape priors and object detection results as intermediate representations, the system bridges the gap between challenging imaging conditions and reliable reconstruction, enabling accurate results even when traditional MVS would fail

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If dense reconstruction is performed without semantic information, then the process is simpler, but the reconstruction accuracy deteriorates on textured surfaces with few views and on difficult surfaces without texture

Engineering Contradiction:
Improvesystem complexityVSAvoidreconstruction accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The system performs preliminary learning of category-level shape priors and object detection models from training data before reconstruction. This pre-computed semantic information is then integrated into the reconstruction pipeline, enabling accurate dense reconstruction on textured surfaces with few views and on difficult surfaces without texture, while keeping the actual reconstruction process computationally efficient

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameter space by incorporating semantic information dimensions (category-level shape priors, object detection features) into the reconstruction process. By augmenting the traditional geometric parameters with semantic parameters, the system achieves higher reconstruction accuracy without proportionally increasing computational complexity

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If traditional MVS systems are used, then the reconstruction process is straightforward, but imaging costs increase due to the requirement of many views to achieve acceptable reconstruction quality

Engineering Contradiction:
Improvereconstruction process simplicityVSAvoidnumber of images required
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The system performs preliminary learning of category-level shape priors from extensive training data before actual reconstruction. This pre-acquired knowledge enables the system to achieve acceptable reconstruction quality with fewer views, thereby reducing imaging costs while maintaining process simplicity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By introducing learned category-level shape priors as an intermediary, the system reduces the number of views needed for reconstruction. The semantic priors compensate for the limited geometric information from fewer images, enabling cost-effective imaging while maintaining reconstruction quality

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9489768B2Semantic dense 3D reconstruction
Publication Date: 2016.11.08 NEC CORP
  • US9489768B2 patent drawing
  • US9489768B2 patent drawing
  • US9489768B2 patent drawing

AI summary

A method to reconstruct 3D model of an object includes receiving with a processor a set of training data including images of the object from various viewpoints; learning a prior comprised of a mean shape describing a commonality of shapes across a category and a set of weighted anchor points encoding similarities between instances in appearance and spatial consistency; matching anchor points across instances to enable learning a mean shape for the category; and modeling the shape of an object instance as a warped version of a category mean, along with instance-specific details.