Volumetric Video Coding Using Region Segmentation and Offset Signaling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Volumetric video content, captured by multiple 3D cameras, poses a significant challenge due to its high data volume, requiring substantial bandwidth for storage and transmission, and existing compression methods are inefficient, especially for dynamic 3D scenes, which limits 6DOF capabilities in VR and AR applications.
Innovation Solution
The method involves segmenting a 3D object into regions and inserting signals indicating intra-frame and inter-frame offsets, as well as depth smoothness constraints into a bitstream for efficient coding, allowing for improved compression by projecting volumetric models onto 2D planes and using standard 2D video coding tools, enhancing coding efficiency and 6DOF capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If volumetric video is captured using multiple 3D cameras, then the 6DOF capabilities and 3D reconstruction quality are improved, but the data volume and bandwidth requirements increase significantly
Solution Approach 1:
The patent segments the 3D object into multiple regions and processes each region separately. By dividing the volumetric data into smaller manageable parts, the system can apply region-specific prediction and coding strategies, reducing the overall data volume while preserving 3D reconstruction quality.
Solution Approach 2:
The patent introduces parameters such as intra-frame offset, inter-frame offset, and depth smoothness constraint to efficiently represent 3D geometry. By changing the parameter representation from raw coordinates to offset values relative to reference frames and regions, the data volume is significantly reduced while maintaining measurement precision.
2Device complexity
If standard 2D video coding tools are used for volumetric video, then the device complexity is reduced, but the coding efficiency is insufficient
Solution Approach 1:
The patent makes standard 2D video coding tools multi-functional by applying them to both 2D projections and 3D volumetric data. The same coding tools are used for different purposes (2D video and 3D volumetric video), reducing device complexity while improving coding efficiency through specialized adaptations for volumetric data.
Solution Approach 2:
The patent introduces intermediate representations (geometry images, depth maps, and offset signals) that bridge between 3D volumetric data and 2D coding tools. These intermediaries enable standard 2D coders to efficiently process volumetric video without requiring complex 3D-specific coding machinery.
3Quantity of substance
If region segmentation and offset signaling are applied, then the bit rate is reduced, but the device complexity increases
Solution Approach 1:
The patent segments the 3D object into regions and applies offset signaling within each region. This segmentation allows the system to reduce bit rate by exploiting local correlations within regions while keeping the complexity manageable through systematic processing of divided data.
Solution Approach 2:
The patent changes the parameter representation to use offsets relative to reference regions and frames. This parameter transformation reduces the bit rate by encoding only the differences rather than absolute values, while the complexity is managed through consistent application of offset calculation and signaling mechanisms.
Data Source
AI summary
The embodiments relate to a method comprising receiving (1311) a volumetric video comprising a three-dimensional object; segmenting (1312) the three-dimensional object into a plurality of regions; for one or more regions of a three-dimensional object (1313): inserting into a bitstream or signaling along a bitstream a signal indicating one or more of the following: intra frame offset relating to three-dimensional geometry value (Z) between two regions within a frame; inter frame offset relating to three-dimensional geometry value (Z) between two regions in different frames; depth smoothness constraint relating to three-dimensional geometry value (Z) and transmitting (1314) the bitstream to a decoder. The embodiments relate to a method for receiving and decoding the bitstream, as well as to technical equipment for implementing any of the methods.


