Volumetric Video Transcoding via Multiplane Image Projection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in efficiently transmitting and processing volumetric videos (VVs) due to high bandwidth requirements, resource-intensive processing, and lack of a computationally and memory-efficient canonical representation, especially when adapting to arbitrary topologies and adaptive-bitrate streaming systems.
Innovation Solution
The method involves identifying a portion of a volumetric video, obtaining viewpoint information, predicting a user's viewpoint, and generating a multiplane image (MPI) representation compatible with two-dimensional hardware, reducing data transmission and processing loads by focusing on relevant data within the predicted viewpoint.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If volumetric videos are transmitted using conventional compression technology, then the videos can be delivered to communication devices, but the bandwidth requirements become excessively high and network resources are heavily consumed
Solution Approach 1:
The volumetric video is segmented into multiple 2D video representations corresponding to different viewpoints. Instead of transmitting the entire 3D volumetric data, only the necessary 2D projections are generated and transmitted based on predicted user viewing angles, significantly reducing bandwidth consumption while maintaining video delivery capability
Solution Approach 2:
The patent transforms 3D volumetric video data into 2D representations that can be efficiently compressed and transmitted using conventional 2D video codecs. By projecting volumetric data onto 2D planes from predicted viewpoints, the system reduces data dimensionality and enables efficient bandwidth utilization without sacrificing essential visual information
2Reliability
If volumetric videos are decoded using software processing, then the videos can be rendered on communication devices, but the processing load and computational overhead become significant
Solution Approach 1:
The system performs viewpoint prediction and selects appropriate 2D video representations in advance, before the actual playback. By pre-processing the volumetric data into multiple viewpoint projections and predicting which views will be needed, the system reduces real-time computational requirements during playback, allowing efficient rendering on resource-constrained devices
Solution Approach 2:
Instead of performing complex 3D volumetric decoding on the communication device, the system generates and transmits pre-processed 2D video copies from predicted viewpoints. The communication device simply needs to decode and display these 2D representations, avoiding the heavy computational burden of real-time 3D reconstruction and rendering
3Productivity
If a canonical representation of volumetric videos is created to improve efficiency, then bandwidth and processing loads are reduced, but the ability to provide high resolution across arbitrary topologies is compromised
Solution Approach 1:
The system dynamically adapts the viewpoint selection and 2D projection generation based on predicted user behavior and viewing context. By using machine learning models to forecast which viewpoints will be most relevant, the system can allocate computational resources to generate high-resolution 2D representations for those specific views, maintaining quality where it matters most while reducing overall data transmission requirements
Data Source
AI summary
Aspects of the subject disclosure may include, for example, transmitting viewpoint information associated with a first portion of a three-dimensional (3D)/volumetric video to a device, wherein the viewpoint information comprises a first coordinate in 3D space associated with a first viewing direction in a playback of the first portion and a first timestamp associated with the first portion, receiving, from the device, a multiplane image (MPI) representation of a second portion of the 3D video responsive to the transmitting of the viewpoint information, and providing an image of the MPI representation to a display device. Other embodiments are disclosed.


