Large-Scale Scene Light Field Reconstruction With Sparse 3D Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for real-time reconstruction and intelligent understanding of large-scale scene light fields face challenges in achieving high precision and speed due to limitations in 2D and 3D convolutional neural networks, which fail to capture global 3D information and are computationally inefficient, leading to low accuracy and slow processing.

Innovation Solution

A real-time light field reconstruction network model using sparse convolutional networks and online segmentation modules is trained with the ScanNet dataset to extract features from 3D voxels and voxel color information, enabling high-precision semantic and instance segmentation by constraining temporal consistency and clustering instance embeddings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If 2D and 3D convolutional neural networks are used for scene light field reconstruction, then the system can process image data, but the models fail to capture global 3D information and are computationally inefficient

Engineering Contradiction:
Improvesegmentation accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent divides the 3D scene into volumetric voxels and processes them through sparse convolutional operations, segmenting the global 3D space into manageable local regions while maintaining global context through hierarchical processing structures

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from traditional 2D image processing to 3D volumetric voxel processing, enabling the system to capture global 3D information by operating on three-dimensional data structures rather than two-dimensional images

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If traditional convolutional neural networks are used for real-time reconstruction, then the system can generate 3D models, but the processing is computationally inefficient and slow

Engineering Contradiction:
Improvereconstruction qualityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies sparse convolutional operations that process only relevant local voxel regions rather than the entire 3D volume, improving computational efficiency by focusing calculations on areas with significant visual information while maintaining reconstruction quality

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts processing parameters and computational resources based on scene complexity and real-time requirements, enabling adaptive performance optimization between reconstruction quality and processing speed

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If offline 3D segmentation methods are used, then high precision can be achieved, but the processing cannot be performed in real-time

Engineering Contradiction:
Improvesegmentation accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent pre-processes and organizes 3D voxel data into structured representations before segmentation, preparing the scene graph and voxel mappings in advance to enable rapid real-time segmentation without sacrificing accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system maintains continuous processing of 3D scene data through real-time voxel updates and incremental segmentation, ensuring consistent and accurate segmentation results while operating continuously at frame rates suitable for real-time applications

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12380579B2Intelligent understanding system for real-time reconstruction of large-scale scene light field
Publication Date: 2025.08.05 GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
  • US12380579B2 patent drawing
  • US12380579B2 patent drawing
  • US12380579B2 patent drawing

AI summary

An intelligent understanding apparatus for real-time reconstruction of a large-scale scene light field includes the following. A data obtaining module obtains a 3D instance depth map, and obtain 3D voxels and voxel color information through simultaneous positioning and map generation. The model constructing module constructs and trains a real-time light field reconstruction network model using a ScanNet dataset. The real-time light field reconstruction network model extracts features of the 3D voxels and voxel color information, and obtain a semantic segmentation result and an instance segmentation result. The semantic segmentation module inputs the 3D voxel and voxel color information corresponding to the 3D instance depth map into the trained real-time light field reconstruction network model, and determine an output as a semantic segmentation result and an instance segmentation result corresponding to the 3D instance depth map.