Multi-View Video Synthesis Depth Data Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current immersive video processing schemes face challenges in efficiently transmitting and synthesizing intermediate views, particularly due to high bit rates and poor quality depth information, which can lead to motion sickness and increased computational load on terminals like smartphones.
Innovation Solution
A method for processing multi-view video data that allows for flexible selection of modes for obtaining synthesis data, either by decoding it from the encoded data stream or estimating it on the client side, optimizing encoding cost and synthesis quality based on available tools and content characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If depth maps are calculated and transmitted for all views, then the quality of intermediate view synthesis is improved, but the bit rate and transmission cost increase significantly
Solution Approach 1:
The patent extracts and transmits only the necessary depth information for intermediate view synthesis rather than complete depth maps for all views. The depth estimator identifies and transmits only the depth data that will be useful for synthesizing intermediate views, eliminating redundant depth information transmission.
Solution Approach 2:
The depth estimator on the server side performs preliminary analysis to determine which depth information will be needed for intermediate view synthesis before encoding. This allows the system to prepare and transmit only the relevant depth data in advance, avoiding the need to transmit all depth maps.
2Manufacturing precision
If complete depth maps are generated and transmitted, then synthesis quality is improved, but redundant information increases transmission cost
Solution Approach 1:
The system extracts only the essential depth information needed for intermediate view synthesis from the complete depth maps. The depth estimator identifies regions and depth data that are actually required for synthesis, separating useful information from redundant data before transmission.
Solution Approach 2:
The patent applies different quality levels to different regions of depth maps based on their importance for intermediate view synthesis. Critical regions are transmitted with high quality, while less important regions are either transmitted with lower quality or omitted entirely, optimizing the balance between synthesis quality and transmission cost.
3Manufacturing precision
If depth information is captured with dedicated sensors, then depth quality is improved, but device complexity and cost increase
Solution Approach 1:
The patent replaces dedicated depth sensing hardware with computational depth estimation methods. Instead of using complex sensor systems to capture depth information, the system uses software-based depth estimation algorithms that analyze texture images to generate depth maps, significantly reducing device complexity while maintaining acceptable depth quality.
Solution Approach 2:
The depth estimator acts as an intermediary that converts readily available texture image data into depth information. This intermediary process allows the system to obtain depth data without direct sensor measurement, using computational methods to bridge the gap between 2D images and 3D depth information.
4Adaptability or versatility
If multiple captured views are transmitted, then intermediate view synthesis capability is improved, but the number of views and data amount increase
Solution Approach 1:
The system extracts and transmits only the essential view data and associated depth information needed for intermediate view synthesis. Rather than transmitting all captured views, the depth estimator identifies the minimum set of view data required to reconstruct intermediate views, reducing the overall data amount while maintaining synthesis capability.
Solution Approach 2:
The depth estimator performs preliminary analysis to determine which captured views contain the most valuable information for intermediate view synthesis. This allows the system to select and transmit only the most useful views in advance, reducing data transmission requirements while preserving synthesis flexibility.
Data Source
AI summary
A method for processing multi-view video data including, for at least one block of an image of a view encoded in an encoded data stream representing the multi-view video, obtaining at least one information item, which specifies a mode for obtaining a synthesis data item, from among first and second obtaining modes. The synthesis data item is used to synthesize at least one image of an intermediate view of the multi-view video, the intermediate view not being encoded in the encoded data stream. The first obtaining mode involves decoding an information item representing the synthesis data item from the encoded data stream, the second obtaining mode involves obtaining the synthesis data item from at least the reconstructed encoded image. At least one part of an image of the intermediate view is synthesised from at least the reconstructed encoded image and the synthesis data item obtained by the specified mode.


