Volumetric Video Encoding and Decoding Through 2D Projection Views
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing volumetric video coding technologies suffer from poor spatial and temporal compression performance, particularly in dynamic 3D scenes where geometry and attributes change, leading to inefficient compression.
Innovation Solution
A system for capturing, encoding, decoding, and reconstructing three-dimensional scenes using multiple cameras and microphones to create a scene model, projecting 3D data onto 2D planes, and employing standard 2D video coding tools for efficient temporal compression, with projection surfaces optimized for individual objects to improve coverage and utilize standard video encoding hardware.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If volumetric video data is represented using traditional formats (triangle meshes, point clouds, or voxel arrays), then the data can be stored and processed, but the spatial and temporal coding performance deteriorates
Solution Approach 1:
The patent transforms 3D volumetric video data into 2D projection views, projecting three-dimensional scene geometry and attributes onto two-dimensional planes. This dimensionality reduction enables the use of efficient 2D video coding standards while preserving the ability to reconstruct 3D scenes from multiple viewpoints, thereby improving coding performance without excessive complexity
Solution Approach 2:
The patent divides the volumetric scene into multiple discrete projection views or layers, each representing a different perspective or depth plane. By segmenting the 3D data into manageable 2D components, the system can apply standard video coding techniques to each segment independently, improving overall compression efficiency while maintaining spatial and temporal coherence
2Productivity
If standard 2D video coding tools are used for volumetric data, then coding efficiency improves, but the 3D spatial information may be lost or degraded
Solution Approach 1:
The patent applies different coding strategies to different regions or layers of the volumetric data. By tailoring the projection and coding approach to local spatial characteristics, the system maintains high spatial accuracy in critical regions while achieving efficient compression overall, preventing uniform degradation of 3D spatial information
Solution Approach 2:
The patent combines multiple 2D projection views or layers into a composite volumetric representation. By synthesizing information from multiple coded 2D perspectives, the system reconstructs accurate 3D spatial structure, preventing loss of spatial information while benefiting from efficient 2D coding applied to each component
Data Source
Figure 1
Figure 2a~2b
Figure 3a~3b
AI summary
There are provided methods, apparatuses, systems and computer program products for coding volumetric video, where a first texture picture is coded, the first texture picture comprising a first projection of texture data of a first source volume of a digital scene model, the scene model comprising a number of further source volumes, the first projection being from the first source volume to a first projection surface, a first geometry picture is coded, the first geometry picture representing a mapping of the first projection surface to the first source volume, and first projection geometry information of the first projection is coded, the first projection geometry information comprising information of position of the first projection surface in the scene model.