Volumetric Video Stream Encoding for 3DoF and Volumetric Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack a stream format and associated methods that can efficiently encode and decode volumetric video data, allowing for both 3DoF and volumetric rendering while minimizing data requirements.
Innovation Solution
A method and device for encoding a 3D scene into a stream, where the stream is structured in elements of syntax, including generating color and depth data for visible and non-visible points of the scene, and encoding these data in separate elements of syntax for efficient decoding and rendering in either 3DoF or volumetric modes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If 6DoF volumetric video is encoded using multiple views and depth information, then immersion and depth perception are improved, but data volume and encoding complexity increase significantly
Solution Approach 1:
The patent segments the volumetric video data into multiple view components and depth components separately. Each view captures a portion of the 360-degree scene, and depth information is extracted and encoded independently. This segmentation allows the system to provide 6DoF immersion when needed while enabling fallback to lower-data modes (like 3DoF or single-view) when data volume becomes excessive, thus resolving the contradiction between immersion quality and data volume.
Solution Approach 2:
The patent implements a flexible encoding scheme where full 6DoF volumetric data (multiple views + depth) is encoded only when necessary for maximum immersion. The system can selectively encode partial data sets (e.g., only color information, or limited views) depending on bandwidth and storage constraints. This partial action approach maintains the capability for high immersion when resources allow, while reducing data volume when constraints exist.
2Ease of manufacture
If legacy standard encoding methods are used for multiview plus depth video, then encoding simplicity is maintained, but data requirements become unmanageably large for broadcasting or streaming
Solution Approach 1:
The patent merges multiple view data and depth data into a unified volumetric representation that shares common spatial and temporal structures. By combining these data types in a coordinated manner with joint encoding techniques, the system achieves better compression efficiency than encoding each view and depth map separately using legacy standards. This merging reduces overall data volume while maintaining encoding feasibility.
Solution Approach 2:
The patent creates a universal encoding framework that can handle multiple data configurations (full 6DoF, partial 6DoF, 3DoF, single-view) within a single standardized structure. This multi-functional encoding system adapts to different data volume requirements while maintaining a consistent encoding approach, eliminating the need for separate legacy encoding processes and enabling efficient broadcasting and streaming across various bandwidth conditions.
3Productivity
If streams are encoded for one specific rendering type (3DoF or volumetric), then encoding efficiency is improved, but adaptability to different rendering modes is lost
Solution Approach 1:
The patent implements a dynamic encoding structure where the stream can adapt its composition based on the intended rendering mode. The encoding process incorporates flexible data organization that allows efficient extraction for 3DoF rendering (using only necessary view portions) or full volumetric rendering (utilizing all views and depth information). This dynamic structure maintains encoding efficiency for each specific mode while enabling versatility across multiple rendering types through a single unified stream format.
Data Source
AI summary
A sequence of three-dimension scenes is encoded as a video by an encoder and transmitted to a decoder which retrieves the sequence of 3D scenes. Points of a 3D scene visible from a determined point of view are encoded as a color image in a first track of the stream in order to be decodable independently from other tracks of the stream. The color image is compatible with a three degrees of freedom rendering. Depth information and depth and color of residual points of the scene are encoded in separate tracks of the stream and are decoded only in case the decoder is configured to decode the scene for a volumetric rendering.


