Volumetric Video View-Driven Specularity Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in encoding and decoding volumetric video content with consistent rendering of light effects, particularly in 3DoF+ scenarios, where visual artifacts occur due to the lack of sufficient information from Multi-View+Depth (MVD) frames to recover physically true illumination and material properties.
Innovation Solution
A method for encoding a 3D scene as a multiviews-plus-depth (MVD) frame involves selecting a reference view based on field of view coverage, generating an atlas image packing patches with 3D scene information, and encoding metadata that includes acquisition parameters and an identifier for the reference view, ensuring consistent light effects during decoding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If MVD frames are used to encode volumetric video content, then 3DoF+ viewing experience with parallax is enabled, but visual artifacts occur due to insufficient information to recover physically true illumination and material properties
Solution Approach 1:
The patent applies preliminary action by pre-selecting a reference view from the MVD frames that provides the most reliable lighting information, and pre-processing the atlas images to separate albedo and lighting components. This preparation is done before decoding, so that when viewport images are generated, the consistent lighting information is already available, preventing visual artifacts without requiring complete physical illumination recovery from all MVD frames.
Solution Approach 2:
The patent introduces an intermediary approach by using a reference view as a mediator to provide consistent lighting information across different viewport images. Instead of attempting to recover complete physical illumination from all MVD frames, the reference view serves as an intermediary source of lighting data that is applied consistently during viewport synthesis, resolving the contradiction between enabling 3DoF+ and maintaining rendering consistency.
2Measurement precision
If complete physical illumination recovery is attempted from MVD frames, then physically true material properties can be recovered, but the complexity of encoding and decoding increases significantly
Solution Approach 1:
The patent applies the taking out principle by extracting only the essential lighting information from a selected reference view, rather than attempting to recover complete physical illumination from all MVD frames. The atlas images are processed to separate albedo and lighting components, and only the necessary lighting data from the reference view is preserved and applied during decoding. This extraction approach reduces encoding and decoding complexity while maintaining sufficient visual quality.
Solution Approach 2:
The patent applies local quality by treating different parts of the volumetric video data with different processing approaches. The reference view receives special processing to extract lighting information, while other views are used primarily for geometric and textural data. This localized processing strategy reduces overall complexity by focusing computational effort only where needed for lighting consistency.
3Reliability
If atlas images are generated from multiple views to improve lighting consistency, then visual artifacts are reduced, but the data size and processing requirements increase
Solution Approach 1:
The patent applies preliminary action by pre-selecting a single reference view from the MVD frames that provides the most reliable lighting information. This reference view is processed to separate albedo and lighting components before encoding. By performing this selection and processing in advance, the patent avoids the need to store and process multiple atlas images during decoding, thereby reducing data size while maintaining lighting consistency.
Solution Approach 2:
The patent applies the taking out principle by extracting lighting information from only the selected reference view, rather than combining data from multiple views. This extraction approach reduces the quantity of data that needs to be encoded and decoded, as only the essential lighting parameters from the reference view are preserved, not complete atlas images from all views.
Data Source
AI summary
Methods and devices are provided for encoding, transmitting and decoding 3DoF+ volumetric video. At the encoding stage one input view (among all the input ones) is selected to convey the viewport dependent light effect and its id is transmitted to the decoder as an extra metadata. On the decoder side, when patches coming from this selected view are available for the rendering of the viewport, they are preferentially used regarding the other candidates whatever the view to synthesize position.


