Decoder Metadata for Virtual View Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In virtual reality applications, especially those using free navigation, the encoder and decoder lack a priori knowledge of the final user's viewpoint, leading to suboptimal synthesis of virtual views due to the lack of correlation between decoding and synthesis processes, resulting in medium-quality virtual views when the number of captured and decoded views is insufficient.
Innovation Solution
A method where the decoder provides metadata associated with reconstructed views to an image processing module, facilitating the synthesis of virtual views by reducing operational complexity and enabling more powerful image processing algorithms, thereby improving view synthesis quality and reducing the number of cameras required.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If the decoder provides only reconstructed views without additional metadata, then the decoding process remains simple, but the image processing module cannot efficiently synthesize virtual views due to lack of correlation with the decoding process
Solution Approach 1:
The decoder performs preliminary actions by extracting and providing metadata during the decoding process itself, before the image processing module needs to synthesize virtual views. This preliminary extraction of depth information, motion vectors, and other relevant data establishes the correlation between decoding and synthesis, enabling higher quality virtual view synthesis without requiring the image processing module to recalculate this information from scratch
Solution Approach 2:
Metadata acts as an intermediary between the decoder and the image processing module. This intermediary carries essential information (depth maps, motion vectors, occlusion maps) that bridges the gap between the decoding process and the virtual view synthesis process, enabling the image processing module to efficiently generate high-quality virtual views while maintaining a relatively simple decoder architecture
2Manufacturing precision
If more captured views are used to improve virtual view synthesis quality, then the synthesis quality improves, but the number of cameras and data transmission requirements increase
Solution Approach 1:
The decoding process itself provides the necessary metadata (depth information, motion vectors, occlusion maps) that the image processing module needs for virtual view synthesis. This self-service approach means the system doesn't need additional cameras or external data sources - the existing multi-view video data and its decoded metadata are sufficient to enable high-quality virtual view synthesis with fewer captured views
Data Source
AI summary
A method and a device for decoding a data stream representative of a multi-view video. Syntax elements are obtained from at least one part of the stream data, and used to reconstruct at least one image of a view of the video. Then, at least one item of metadata in a predetermined form is obtained from at least one obtained syntax element, and provided to an image processing module. Also, a method and a device for processing images configured to read the at least one item of metadata in the predetermined form and use it to generate at least one image of a virtual view from a reconstructed view of the multi-view video.


