AR/VR 3D Projection via Neural Network Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating three-dimensional (3D) image data in augmented and virtual reality systems are inefficient due to high processing requirements and the need for dense data such as depth maps or optical flow maps, which limits their processing speed and transfer rates.

Innovation Solution

A method using a single lens camera to capture a sequence of images, semantically segmenting the object, stabilizing the images, and computing on-the-fly interpolation parameters to generate stereoscopic pairs for real-time 3D projections in AR/VR environments without polygon generation or texture mapping.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If dense depth maps or optical flow maps are used to describe scene structure, then 3D model accuracy is improved, but processing time and data transfer rates deteriorate

Engineering Contradiction:
Improve3D model accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent extracts only the essential geometric information needed for 3D reconstruction from the image sequence, rather than using dense depth maps or optical flow maps. By selecting and processing only key feature points and their trajectories, the system achieves 3D model accuracy while significantly reducing processing time and data transfer requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the scene into discrete feature points and trajectories, processing only the necessary geometric elements rather than entire dense maps. This segmentation approach maintains 3D reconstruction accuracy by focusing on key structural information while reducing overall data volume and processing requirements.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If computer generation of polygons or texture mapping is used to produce 3D models, then 3D model quality is improved, but processing resources and time deteriorate

Engineering Contradiction:
Improve3D model qualityVSAvoidprocessing resources
Core Design Contradiction:
Manufacturing precisionVSPower

Solution Approach 1:

The patent creates a lightweight 3D model representation by copying and transforming only essential geometric data (feature points and trajectories) rather than generating complete polygon meshes with texture mapping. This approach maintains adequate 3D model quality for AR/VR applications while dramatically reducing processing resource requirements.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If multiple images are combined to produce panoramas or 3D images, then immersive experience is improved, but data processing complexity deteriorates

Engineering Contradiction:
Improveimmersive experienceVSAvoiddata processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary processing of the image sequence by identifying and tracking feature points before 3D reconstruction. By pre-processing the images to extract only essential geometric information (feature points and their trajectories), the system reduces subsequent processing complexity while maintaining the ability to generate immersive 3D and panoramic views.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10719939B2Real-time mobile device capture and generation of AR/VR content
Publication Date: 2020.07.21 FUSION INC
  • US10719939B2 patent drawing
  • US10719939B2 patent drawing
  • US10719939B2 patent drawing

AI summary

Various embodiments describe systems and processes for generating AR/VR content. In one aspect, a method for generating a three-dimensional (3D) projection of an object is provided. A sequence of images along a camera translation may be obtained using a single lens camera. Each image contains at least a portion of overlapping subject matter, which includes the object. The object is semantically segmented from the sequence of images using a trained neural network to form a sequence of segmented object images, which are then refined using fine-grained segmentation. On-the-fly interpolation parameters are computed and stereoscopic pairs are generated for points along the camera translation from the refined sequence of segmented object images for displaying the object as a 3D projection in a virtual reality or augmented reality environment. Segmented image indices are then mapped to a rotation range for display in the virtual reality or augmented reality environment.