Multiview Video Encoding Using Disparity-Compensated Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multiview video encoding and decoding methods lack efficient methods to account for temporal and spatial correlations between views, leading to increased network bandwidth and storage requirements due to independent encoding of videos from multiple camera viewpoints.
Innovation Solution
A method and system that decompose multiview videos using a combination of motion-compensated temporal filtering (MCTF) and disparity-compensated inter-view filtering (DCVF), allowing for adaptive selection of prediction modes including temporal, spatial, and view synthesis modes, to generate a synthesized view of the scene, which correlates frames across different camera views.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If independent encoding of multiview videos is used, then encoding simplicity is maintained, but compression efficiency deteriorates
Solution Approach 1:
The multiview video encoding is segmented into reference view encoding and non-reference view encoding. Reference views are encoded independently using conventional techniques, while non-reference views are encoded using inter-view prediction based on reconstructed reference views. This segmentation allows the system to balance encoding simplicity with improved compression efficiency by applying complex prediction only where beneficial.
Solution Approach 2:
The patent merges temporal prediction and inter-view prediction into a unified disparity compensated prediction framework. By combining these prediction methods and using reconstructed reference views from decoders as prediction bases, the system achieves better compression efficiency while maintaining a manageable encoding structure that builds upon conventional video encoding techniques.
2Productivity
If disparity compensated prediction is used, then compression efficiency is improved, but device complexity increases
Solution Approach 1:
The system performs preliminary encoding of reference views using conventional video encoders before applying disparity compensated prediction to non-reference views. By preparing reconstructed reference views in advance and making them available from decoders, the system reduces the complexity of the prediction process while maintaining compression efficiency benefits.
Solution Approach 2:
Reconstructed reference views serve as intermediaries between the original multiview videos and the final encoded output. These intermediaries facilitate the prediction process by providing a stable basis for disparity compensated prediction, reducing the direct complexity between multiple video streams while maintaining compression efficiency.
3Adaptability or versatility
If conventional video encoding is used, then compatibility is maintained, but network bandwidth increases
Solution Approach 1:
The system creates copied reconstructed views from reference videos using disparity compensated prediction. These copied views are then used as reference material for encoding other views, reducing the amount of original video data that needs to be transmitted while maintaining compatibility with conventional encoding standards through the use of standard video encoders for reference views.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Multiview videos are acquired by overlapping cameras. Side information is used to synthesize multiview videos. A reference picture list is maintained for current frames of the multiview videos, the reference picture indexes temporal reference pictures and spatial reference pictures of the acquired multiview videos and the synthesized reference pictures of the synthesized multiview video. Each current frame of the multiview videos is predicted according to reference pictures indexed by the associated reference picture list with a skip mode and a direct mode, whereby the side information is inferred from the synthesized reference picture. Alternatively, the depth images corresponding to the multiview videos of the input data, and this data are encoded as part of the bitstream depending on a SKIP type.