Multi-view Video Synthesis via Extracted Depth Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current immersive video processing systems face challenges in efficiently transmitting and synthesizing intermediate views, particularly due to high bit rates and the need for accurate depth maps, which are often of poor quality and redundant.
Innovation Solution
The method involves synthesizing intermediate views on the client side using data obtained from decoded and reconstructed views, eliminating the need for transmitting depth maps and reducing the encoding rate. This is achieved by extracting synthesis data, such as depth maps, on the decoder side from decoded textures, and optionally refining this data using neural networks or refinement data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If depth maps are calculated and transmitted prior to encoding, then view synthesis quality is improved, but transmission bit rate increases significantly
Solution Approach 1:
The patent applies preliminary action by calculating depth maps at the encoder side before encoding and transmitting them to the decoder. This allows the depth information to be prepared in advance with high quality, enabling accurate view synthesis at the decoder without requiring high bit rates for depth transmission, since the depth maps are already computed and optimized before transmission.
Solution Approach 2:
The patent extracts depth information from the multi-view video data at the encoder side using depth estimation algorithms. By separating the depth map extraction and transmission from the texture transmission, the system can optimize depth quality independently while controlling the overall bit rate, as depth maps constitute a significant but manageable portion of the total data.
2Reliability
If complete depth maps are generated and transmitted, then view synthesis coverage is improved, but redundant information increases transmission cost
Solution Approach 1:
The patent applies local quality by transmitting complete depth maps only when necessary, and using selective depth transmission based on the specific view synthesis requirements. The encoder can determine which depth regions are actually needed for the requested intermediate views and transmit only those relevant portions, reducing redundant data while maintaining synthesis coverage where needed.
Solution Approach 2:
The patent uses partial action by transmitting depth maps at a resolution and detail level that is sufficient for the specific synthesis task, rather than always transmitting complete high-resolution depth maps. The encoder can adjust the depth map transmission quality and coverage based on the required synthesis accuracy and available bit rate, avoiding excessive data transmission.
3Loss of information
If depth estimation is performed on server side, then synthesis data availability is improved, but encoding complexity increases
Solution Approach 1:
The patent applies self-service by having the encoder perform depth estimation automatically as part of the encoding process. The depth maps are generated by the encoding system itself using integrated depth estimation algorithms, eliminating the need for separate external depth calculation steps and making the synthesis data self-sufficient within the encoded bitstream.
Solution Approach 2:
The patent merges the depth estimation function with the video encoding process. The depth map generation, encoding, and transmission are combined into a unified encoding pipeline, reducing overall system complexity by eliminating separate processing stages and integrating depth synthesis capabilities directly into the encoder.
4Measurement precision
If multiple captured views are transmitted, then view synthesis accuracy is improved, but transmission rate increases
Solution Approach 1:
The patent uses depth maps as an intermediary element that enables accurate view synthesis from fewer captured views. Instead of transmitting multiple high-resolution captured views, the system transmits captured views along with computed depth maps, which act as intermediary data that allows the decoder to synthesize additional views with accuracy comparable to having multiple captured views, but with lower transmission requirements.
Data Source
AI summary
A method and a device for processing multi-view video data. The multi-view video data includes at least one part of a decoded image of at least one view of the multi-view video, from an encoded data stream representative of the multi-view video. At least one item of data, referred to as synthesis data, is obtained from at least the one part of the decoded image, and at least one image of an intermediate view of the multi-view video not encoded in the encoded data stream is synthesized from at least the one part of the decoded image and from the synthesis data obtained.


