Volumetric Video Encoding Using Auxiliary Patch Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current volumetric video coding methods fail to effectively compress and reconstruct 3D scenes, particularly in dense point clouds or voxel arrays, as they lack information about the nature of the 3D objects, leading to inefficient compression and limited 6DOF capabilities in applications like medical imaging and immersive media.
Innovation Solution
The method involves projecting a 3D representation onto 2D patches, generating geometry, texture, and occupancy maps, along with auxiliary patch information that includes metadata and indicators for reconstructing the 3D representation, and encoding these in a bitstream using standard 2D video compression techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If 3D representation is projected onto 2D planes using standard video compression, then compression efficiency is improved, but information about the nature of 3D objects (solidity, cavities) is lost
Solution Approach 1:
The patent segments the 3D scene into multiple 2D patches projected from different viewpoints, allowing standard 2D compression techniques to be applied to each patch independently. This segmentation enables efficient compression while preserving 3D object information through the collection of multiple patches, each carrying spatial and depth data that reconstructs the original 3D structure with solidity and cavity information.
Solution Approach 2:
The patent transforms 3D volumetric data into 2D patch representations for compression, then reconstructs the 3D scene by projecting these patches back. This dimensional transformation enables the use of efficient 2D compression algorithms while maintaining 3D object nature information through the depth maps and spatial coordinates encoded in the patches.
2Measurement precision
If dense point clouds or voxel arrays are used to represent 3D scenes, then 3D scene representation accuracy is improved, but device complexity and processing requirements increase
Solution Approach 1:
The patent extracts essential 3D information (depth maps, surface normals, texture coordinates) from dense point clouds or voxel arrays and encodes only these extracted features in the bitstream. This extraction approach maintains high 3D scene representation accuracy while significantly reducing the data volume and processing complexity compared to transmitting complete dense 3D models.
Solution Approach 2:
The patent creates 2D patch copies from 3D data that preserve the essential geometric and textural information needed for accurate 3D scene representation. These 2D patches serve as compressed copies that can be efficiently processed and reconstructed, reducing the computational burden while maintaining measurement precision.
3Device complexity
If only hollow 3D surface information is reconstructed, then compression is simplified, but applications requiring solid object information (medical imaging, autonomous navigation) cannot benefit
Solution Approach 1:
The patent encodes different types of information (surface geometry, depth, texture, and solidity indicators) in different patches or layers of the bitstream. This local quality differentiation allows the reconstruction system to provide hollow surface information where simple compression is sufficient, while preserving solid object information and cavity data where applications like medical imaging and autonomous navigation require it, thereby enhancing adaptability without uniformly increasing complexity.
Data Source
AI summary
There are disclosed a method and an apparatus for video encoding. A method can include projecting a 3D representation of at least one object onto at least one 2D patch (400); generating a geometry image, a texture image, an occupancy map and auxiliary patch information from the 2D patch, wherein the auxiliary patch information comprises metadata relating to properties of the patch and one or more indicators for configuring reconstruction of the 3D representation of the at least one object (402); and encoding the geometry image, the texture image, the occupancy map and the auxiliary patch information in or along a bitstream (404).


