3D Video Model Reconstruction via Depth Stream Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for real-time holographic communication face challenges in universality and data transmission quality due to the need for high-speed networks and compression of 3D video, which results in information loss.

Innovation Solution

The method involves receiving depth video streams from multiple camera perspectives, determining a 3D video model, performing light field rendering based on interaction parameters, and sending target light field rendering views to construct a 3D image, reducing the need for high-speed networks and enhancing universality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If 3D video is compressed for transmission, then data transmission volume is reduced, but information is lost and view quality deteriorates

Engineering Contradiction:
Improvedata transmission volumeVSAvoidview quality
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent extracts only the essential depth information from the 3D video data, separating it from the full-color video data. By transmitting only depth video streams instead of complete 3D video, the system reduces data transmission volume while preserving the critical information needed for 3D reconstruction at the display end.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a lightweight copy of the 3D scene in the form of depth video streams, which can be transmitted efficiently. The full 3D reconstruction is then copied back at the display end through light field rendering, avoiding the need to transmit the complete high-volume 3D video data while maintaining view quality.

Inventive Principle:
Principle #26Copying

2Speed

If high-speed networks like 5G are used for transmission, then data transmission speed is improved, but universality deteriorates due to poor compatibility with lower-speed networks

Engineering Contradiction:
Improvedata transmission speedVSAvoidnetwork compatibility
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent changes the data format parameter from full 3D video to depth video streams, which have significantly reduced data volume. This parameter change enables the system to achieve acceptable transmission speeds on lower-speed networks like 4G, thereby improving universality and compatibility across different network conditions without requiring high-speed 5G infrastructure.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If multiple cameras are used to capture 3D video, then view quality is improved, but device complexity and cost increase

Engineering Contradiction:
Improveview qualityVSAvoidcamera system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts depth information as a separate data stream from the camera system. By using depth sensors or depth estimation algorithms to capture only the essential 3D structural information, the system can achieve good 3D reconstruction quality with fewer cameras compared to traditional multi-camera 3D video systems that capture full color and depth data from multiple perspectives.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20240296626A1Method, apparatus, electronic device and storage medium for reconstructing 3D images
Publication Date: 2024.09.05 DOUYIN VISION CO LTD
  • US20240296626A1 patent drawing
  • US20240296626A1 patent drawing
  • US20240296626A1 patent drawing

AI summary

This disclosure discloses a method, an apparatus, an electronic device, and a storage medium for reconstructing a 3D image. The method of reconstructing a 3D image includes: receiving depth video streams of at least two camera perspectives of a same scene; determining a 3D video model corresponding to the depth video streams of the at least two camera perspectives; performing a light field rendering on the 3D video model based on an obtained interaction parameter to obtain a plurality of target light field rendering views; and sending the plurality of target light field rendering views to a display end to construct a 3D image corresponding to the depth video streams at the display end.