Multi-View Reconstruction Using Foreground Segmentation for Occlusion Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-view reconstruction systems fail to provide photo-realistic outputs due to inability to differentiate between foreground and background objects, leading to occlusion issues and inconsistent color rendering, which affects the accuracy and realism of rendered images.

Innovation Solution

A system comprising multiple cameras, a CEM module for environment modeling, an FES module for segmenting foreground from background, and a configuration and rendering engine that allows user-selectable novel views, ensuring less than 10% discrepancy in pixel raster values between output images and camera frames, effectively addressing occlusion and color inconsistencies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If multi-view reconstruction is performed using known systems, then 3D data representation can be created, but photo-realistic rendering quality is lost due to occlusion issues

Engineering Contradiction:
Improverendering accuracyVSAvoidocclusion handling
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The system segments the 3D scene into multiple layers including foreground objects, background environment, and occluding objects. This segmentation allows independent processing of each layer to properly handle occlusion relationships, where foreground objects can occlude background objects and occluding objects can occlude other objects. The layered approach enables photo-realistic rendering by correctly establishing the occlusion hierarchy that known systems fail to maintain.

Inventive Principle:
Principle #1Segmentation

2Productivity

If virtual rendering camera is projected through objects, then rendering can be performed, but occlusion accuracy is destroyed

Engineering Contradiction:
Improverendering efficiencyVSAvoidocclusion accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system performs preliminary actions by pre-processing the 3D scene data to identify and categorize all objects into foreground, background, and occluding layers before rendering. Occlusion relationships are pre-calculated and stored as part of the scene representation. When the virtual rendering camera projects through the scene, the pre-established occlusion hierarchy allows efficient ray-tracing that automatically respects occlusion boundaries, maintaining both rendering efficiency and occlusion accuracy.

Inventive Principle:
Principle #10Preliminary action

3Stability of the object's composition

If color consistency is maintained across views, then flat rendering is achieved, but photo-realistic quality is lost due to abrupt color variations

Engineering Contradiction:
Improvecolor consistencyVSAvoidphoto-realistic quality
Core Design Contradiction:
Stability of the object's compositionVSManufacturing precision

Solution Approach 1:

The system applies local quality by allowing different color rendering behaviors in different spatial regions. In regions where occlusion occurs, the system permits abrupt color variations to accurately represent the transition from one object to another. In non-occluded regions, color consistency is maintained. This localized approach to color rendering enables photo-realistic quality by preserving the natural color discontinuities that occur at object boundaries while maintaining smooth colors within uniform regions.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11463678B2System for and method of social interaction using user-selectable novel views
Publication Date: 2022.10.04 INTEL CORP
  • US11463678B2 patent drawing
  • US11463678B2 patent drawing
  • US11463678B2 patent drawing

AI summary

A system for social interaction using a photo-realistic novel view of an event includes a multi-view reconstruction system for developing transmission data of the event a plurality of client-side rendering devices, each rendering device receiving the transmission data from the multi-view reconstruction system and rendering the transmission data as the photo-realistic novel view. A method of social interaction using a photo-realistic novel view of an event includes transmitting by a server side transmission data of the event; receiving by a first user on a first rendering device the data transmission; selecting by the first user a path for rendering on the first rendering device at least on novel view; rendering by the first rendering device the at least one novel view; and saving by the user on the first rendering device novel view date for the at least one novel view.