Multiview Video Random Access via V-Frame Insertion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multiview video encoding and decoding methods fail to efficiently utilize temporal and spatial correlations between views, leading to increased network bandwidth and storage requirements due to independent encoding of each camera view, and lack a method to account for correlations in time.
Innovation Solution
A method that decomposes a multiview bit stream into low and high band frames using motion-compensated temporal filtering (MCTF) and disparity-compensated inter-view filtering (DCVF), with a reference picture list managing temporal, spatial, and synthesized reference pictures for prediction, and inserting V-frames for random temporal access.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If independent encoding is used for each camera view, then encoding simplicity is maintained, but compression efficiency decreases and network bandwidth and storage increase
Solution Approach 1:
The patent merges the encoding of multiple camera views into a single integrated encoding process that exploits temporal and spatial correlations between views. Instead of independently encoding each view, the system uses prediction models that reference other views and time points, thereby reducing the total data量 required for transmission and storage while maintaining encoding feasibility through standardized prediction techniques.
2Ease of manufacture
If independent encoding is used for each camera view, then encoding simplicity is maintained, but compression efficiency decreases and storage requirements increase
Solution Approach 1:
The patent merges the encoding of multiple camera views into a single integrated encoding process that exploits temporal and spatial correlations between views. Instead of independently encoding each view, the system uses prediction models that reference other views and time points, thereby reducing the total data量 required for transmission and storage while maintaining encoding feasibility through standardized prediction techniques.
3Adaptability or versatility
If conventional prediction methods are used, then encoding compatibility is maintained, but temporal and spatial correlations between views are not utilized
Solution Approach 1:
The patent implements a universal prediction framework that can operate within existing video coding standards while simultaneously exploiting temporal and spatial correlations. The prediction model is designed to be compatible with conventional encoders but enhances compression efficiency by utilizing multiple reference sources (temporal references from previous frames and spatial references from other camera views) in a multi-functional prediction process.
4Ease of operation
If no prediction dependency information is provided, then decoding simplicity is maintained, but random access efficiency decreases
Solution Approach 1:
The patent performs preliminary action by pre-calculating and storing prediction dependency information during the encoding process. This metadata about which frames and views are needed for prediction is prepared in advance and transmitted with the encoded data, enabling the decoder to efficiently determine the exact sequence of decoding operations required for random access without needing to analyze the entire bitstream in real-time.
Data Source
AI summary
A method randomly accesses multiview videos. Multiview videos are acquired of a scene with corresponding cameras arranged at poses, such that there is view overlap between any pair of cameras. V-frames are generated from the multiview videos. The V-frames are encoded using only spatial prediction. Then, the V-frames are inserted periodically in an encoded bit stream to provide random temporal access to the multiview videos. Additional view dependency information enables the decoding of a reduced number of frames prior to accessing randomly a target frame for a specified view and time, and decoding the target frame.


