3D Scene Reconstruction via Depth Clustering and Diffusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional three-dimensional scene reconstruction methods consume excessive computational power and fail to ensure accuracy in the reconstruction process.

Innovation Solution

Acquire a color image and a depth image simultaneously after calibration and alignment, determine prior depths of object pixel points through depth clustering, and perform depth diffusion based on color information to generate a minimum bounding box of the object, approximating complex objects with simple geometric structures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional SFM and MVS reconstruction methods are used, then three-dimensional scene reconstruction can be achieved, but excessive computational power is consumed

Engineering Contradiction:
Improvereconstruction accuracyVSAvoidcomputational power consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent segments the reconstruction process into two distinct stages: first reconstructing the background scene using SFM-MVS methods, then reconstructing foreground objects using depth clustering and diffusion on the residual differences. This segmentation allows computationally intensive operations to be applied only where necessary, reducing overall computational power consumption while maintaining reconstruction accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts foreground objects from the background scene by computing depth residuals (differences between original depth maps and background-reconstructed depth maps). By separating foreground extraction into a distinct step, the system avoids applying full SFM-MVS reconstruction to all scene elements, thereby reducing computational overhead while preserving reconstruction reliability for important foreground objects.

Inventive Principle:
Principle #2Taking out (Extraction)

2Productivity

If traditional SFM and MVS reconstruction methods are used, then three-dimensional scene reconstruction can be achieved, but the reconstruction accuracy cannot be ensured

Engineering Contradiction:
Improvereconstruction efficiencyVSAvoidreconstruction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces depth clustering and depth diffusion as intermediary processes between raw depth map acquisition and final reconstruction. These intermediary steps refine depth information by propagating depth values across similar regions and resolving ambiguities, thereby improving reconstruction accuracy without requiring the full computational expense of traditional SFM-MVS methods applied to all scene elements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary background reconstruction and foreground extraction before applying detailed object reconstruction. By pre-processing the scene to separate background and foreground components, the system can apply more accurate but computationally intensive reconstruction techniques only to foreground objects where accuracy is most critical, thereby improving overall reconstruction accuracy while maintaining efficiency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250209732A1Three-dimentional scene reconstruction method and apparatus, device, and storage medium
Publication Date: 2025.06.26 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250209732A1 patent drawing
  • US20250209732A1 patent drawing
  • US20250209732A1 patent drawing

AI summary

The embodiments of the present application provide a three-dimensional scene reconstruction method and apparatus, a device and a storage medium. The method comprises: acquiring a color image and a depth image of a three-dimensional scene at the same time; determining, based on the depth image, prior depths of object pixel points obtained after depth clustering in a detection frame of each object in the color image; performing depth diffusion on the pixel points in the detection frame based on color information of each pixel point in the detection frame and the prior depths of the object pixel points to obtain an actual depth of each pixel point in the detection frame, so as to generate a minimum bounding box of the object.