Volumetric Video Patch Packing for Bandwidth Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Volumetric video data, captured by multiple 3D cameras, requires significant bandwidth for storage and transmission due to its high volume, and existing compression methods are inefficient for dynamic 3D scenes, especially in virtual reality applications where 6 degrees of freedom are needed, as they fail to effectively compress and render 3D models with changing geometry and attributes.
Innovation Solution
The method involves decomposing volumetric video frames into patches, projecting these patches onto 2D planes, and using standard 2D video compression techniques, along with signaling and encoding the bitstream to indicate the presence of multiple video data components, allowing for efficient temporal compression and rendering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If volumetric video data is captured using multiple 3D cameras, then the coverage and quality of the 3D scene are improved, but the data volume and bandwidth requirements increase significantly
Solution Approach 1:
The volumetric video data is divided into multiple patches, each representing a specific region of the 3D scene. These patches are then packed into a 2D video frame structure, allowing selective transmission and rendering of only the relevant patches needed for the current viewpoint, thereby reducing the effective data volume while maintaining complete scene coverage capability
Solution Approach 2:
The patent transforms 3D volumetric data into a 2D packed video frame representation by projecting and packing multiple 3D patches onto 2D planes. This dimensionality reduction allows the use of efficient 2D video compression standards while preserving the ability to reconstruct 3D views through selective unpacking and rendering of relevant patches
2Productivity
If standard 2D video compression techniques are used for volumetric data, then compression efficiency is improved, but the ability to represent dynamic 3D geometry and attributes is lost
Solution Approach 1:
The volumetric data is segmented into multiple patches that are independently processed and packed into 2D frames. Each patch can be compressed using standard 2D video codecs while retaining its 3D spatial information through metadata markers, enabling both efficient compression and 3D geometry preservation
Solution Approach 2:
The patent introduces metadata markers as an intermediary layer between the 2D compressed video data and the original 3D volumetric structure. These markers contain information about patch locations, dimensions, and spatial relationships, allowing the decoder to reconstruct 3D geometry from 2D compressed data without losing versatility
3Quantity of substance
If multiple video data components are packed into one video frame, then bandwidth for transmission is reduced, but the complexity of encoding and decoding increases
Solution Approach 1:
Different volumetric video data components (such as geometry, texture, and attribute data) are segmented into separate patches and then packed into different regions of the same 2D video frame. Each component can be independently encoded using standard video codecs, reducing transmission bandwidth while maintaining manageable encoding complexity through modular processing
Solution Approach 2:
Multiple video data components are merged into a single packed video frame structure, where each component occupies specific regions marked by metadata. This merging reduces the number of separate transmission streams needed while the modular structure with clear delimiters keeps encoding and decoding complexity manageable
Data Source
AI summary
The embodiments relate to a method comprising receiving as an input a volumetric video frame comprising volumetric video data (910); decomposing the volumetric video frame into one or more patches, wherein a patch comprises a volumetric video data component (920); packing several patches, where at least two patches of the several patches comprise a different volumetric video data component with respect to each other, into one video frame (930); generating a bitstream comprising an encoded video frame (940); signaling, in or along the bitstream, existence of encoded video frame containing patches of more than one different volumetric video data component (950); and transmitting the encoded bitstream to a storage for rendering (960). The embodiments also relate to a technical equipment for implementing the method.


