Dynamic Mesh Wavelet Decoding via Base Mesh Segmentation for Lower Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing volumetric video encoding and decoding technologies face challenges in achieving high compression efficiency and hardware compatibility, particularly for dynamic mesh content, leading to increased computational complexity and latency in immersive applications.
Innovation Solution
A wavelet-based encoding approach is used to decode a base mesh of three-dimensional object data, incorporating geometry and occupancy components, along with a displacement field comprising wavelet-encoded and quantized position displacements, within a video component framework, enabling efficient decoding and reducing processing complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If full mesh reconstruction is performed for three-dimensional object data, then mesh accuracy and completeness are improved, but computational complexity and processing time increase
Solution Approach 1:
The mesh is segmented into a base mesh and a displacement field. The base mesh contains the majority of geometric information and is decoded first, while the displacement field contains only the necessary correction data. This segmentation allows the system to achieve high mesh accuracy without processing the entire mesh at full resolution, thereby reducing computational complexity.
Solution Approach 2:
The invention extracts only the essential geometric information into the base mesh and separates the dynamic deformation information into the displacement field. By taking out only the necessary base mesh data for reconstruction and using wavelet encoding to represent only the critical displacement information, the system achieves accurate mesh representation with reduced processing requirements.
2Adaptability or versatility
If traditional video decoding is used for mesh data, then hardware compatibility is improved, but compression efficiency decreases
Solution Approach 1:
The invention makes the video decoding framework universal by enabling it to handle both traditional video data and three-dimensional mesh data through the same decoding pipeline. The video component can decode both 2D video frames and 3D mesh representations, including the base mesh and displacement field, thereby achieving hardware compatibility without sacrificing compression efficiency for mesh-specific data.
Solution Approach 2:
The system changes the data representation parameters by encoding mesh geometry and displacement information in a format compatible with video decoding standards. By transforming mesh data into a video-compatible bitstream format with appropriate parameter settings, the system achieves efficient compression while maintaining hardware compatibility through existing video decoding capabilities.
3Manufacturing precision
If detailed mesh data is transmitted for immersive applications, then visual quality is improved, but data size and transmission time increase
Solution Approach 1:
The mesh data is segmented into a compact base mesh representation and a displacement field. The base mesh transmits only the essential geometric structure, while the displacement field uses wavelet encoding to represent only the necessary deformation information. This segmentation significantly reduces the total data size while maintaining high visual quality for immersive applications.
Solution Approach 2:
The system changes the data encoding parameters by applying wavelet transform and quantization to the displacement field, creating a compressed representation that maintains visual quality while reducing data size. The parameter settings for wavelet decomposition levels and quantization precision are optimized to achieve the best compression ratio for immersive application requirements.
4Manufacturing precision
If complex mesh processing is performed in real-time, then processing accuracy is improved, but latency increases
Solution Approach 1:
The mesh processing is segmented into two independent stages: base mesh decoding and displacement field decoding. The base mesh is decoded first to provide the geometric framework, followed by the displacement field decoding that adds only the necessary deformation corrections. This segmentation enables real-time processing with minimal latency while maintaining processing accuracy, as each stage can be processed independently and efficiently.
Data Source
AI summary
An apparatus including at least one processor; and at least one memory including computer program code. The at least one memory and the computer program code are configured to, with the at least one processor, cause the apparatus at least to: decode, for a frame of three-dimensional object data, a base mesh of the three-dimensional object data, wherein the base mesh has been generated with a geometry component; and decode, for the frame, a displacement field, where the displacement field comprises wavelet encoded and quantized position displacements.


