3D Video Reconstruction Using Pose Transformation Matrices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Three-dimensional video processing faces challenges due to its unique data structure, particularly in aligning and reconstructing depth video streams from multiple camera perspectives, which affects immersion and accuracy in augmented and virtual reality applications.

Innovation Solution

A method and device for processing 3D videos by acquiring depth video streams from multiple cameras, registering them using pose transformation matrices, and performing three-dimensional reconstruction to align and fuse the data, enabling accurate 3D model creation and improved user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If depth video streams from multiple cameras are acquired and registered using pose transformation matrices, then the accuracy and immersion of 3D video reconstruction is improved, but the device complexity and processing difficulty increase

Engineering Contradiction:
Improve3D video reconstruction accuracyVSAvoidvideo processing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-calculating and storing pose transformation matrices for each camera perspective before the actual 3D reconstruction process. These matrices are computed in advance based on camera positions and orientations, allowing the registration module to quickly align depth video streams without performing complex real-time calculations during reconstruction, thus improving accuracy while managing system complexity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary registration module that acts as a mediator between the multiple camera input streams and the final 3D reconstruction process. This module uses pose transformation matrices to align and coordinate the depth video streams from different perspectives, simplifying the overall system architecture by centralizing the complex alignment operations in a dedicated intermediate processing stage

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If depth video streams are registered and three-dimensional reconstruction is performed, then user immersion in AR and VR applications is improved, but the processing time and computational resources increase

Engineering Contradiction:
Improveuser immersion qualityVSAvoidvideo processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies segmentation by dividing the 3D video processing into distinct modular stages: an acquisition module that collects depth video streams from multiple cameras, a registration module that aligns streams using pose transformation matrices, and a reconstruction module that performs 3D reconstruction. This segmentation allows each module to be optimized independently and enables parallel processing where possible, reducing overall processing time while maintaining immersion quality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses preliminary action by pre-computing pose transformation matrices and preparing registration data before the actual reconstruction process. This advance preparation reduces the computational burden during real-time reconstruction, allowing faster processing while maintaining high user immersion quality in AR and VR applications

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240314289A1Method and device for processing three-dimensional video, and storage medium
Publication Date: 2024.09.19 DOUYIN VISION CO LTD
  • US20240314289A1 patent drawing
  • US20240314289A1 patent drawing

AI summary

Provided are a method and a device for processing a three-dimensional video and a storage medium. The method includes the steps described below. Depth video streams from perspectives of at least two cameras of the same scene are acquired. According to preset registration information, the depth video streams from the perspectives of the at least two cameras are registered. According to the registered depth video streams from the perspectives of the at least two cameras, three-dimensional reconstruction is performed to obtain a 3D video.