Multi-view Video Prediction Encoding Using Inter-view Residuals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video codecs face challenges in efficiently encoding and decoding high-resolution or high-quality video content, particularly in multi-view video systems where inter prediction and motion compensation methods are limited, leading to suboptimal data compression and increased bit rates.
Innovation Solution
A multi-view video prediction method and apparatus that generate a base layer image stream through inter prediction between base view images and an enhancement layer image stream through inter-view prediction, using residual values to encode additional view images, allowing for efficient prediction and decoding of additional view images by performing disparity and motion compensation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional video encoding methods based on macroblocks are used, then encoding simplicity is maintained, but coding efficiency for high-resolution video is insufficient
Solution Approach 1:
The video sequence is divided into multiple views (base view and additional views), with each view further segmented into picture groups and individual pictures. This hierarchical segmentation enables efficient processing of high-resolution video by breaking down the large data volume into manageable units that can be encoded and transmitted separately while maintaining overall coding efficiency.
Solution Approach 2:
The patent introduces a multi-view dimension to traditional single-view video coding. By encoding video from multiple viewpoints simultaneously and establishing prediction relationships across views, the system achieves better compression for high-resolution content by exploiting redundancy in the additional dimension of view information.
2Quantity of substance
If inter prediction is performed between base view and additional view images, then bit rate is reduced, but decoding complexity increases
Solution Approach 1:
The base view images are encoded and transmitted first, establishing a foundation for subsequent additional view decoding. Prediction parameters and reference image information are prepared in advance during base view encoding, allowing the decoder to efficiently reconstruct additional views without performing complex real-time calculations, thus reducing overall decoding complexity despite the multi-view nature.
Solution Approach 2:
The base view images serve as an intermediary reference for decoding additional view images. By using the base view as a mediator, the system reduces the bit rate for additional views through prediction while managing decoding complexity through a structured two-stage process where the base view facilitates efficient reconstruction of additional views.
3Manufacturing precision
If multiple additional view images are encoded with full redundancy, then image quality is maintained, but data compression efficiency decreases
Solution Approach 1:
Different quality levels are applied to different views based on their importance and usage. The base view is encoded with higher quality to serve as a reliable reference, while additional views use prediction-based encoding with acceptable quality for their specific purposes. This local quality differentiation maintains overall image quality while significantly improving data compression efficiency.
Solution Approach 2:
The encoding parameters are changed across different views and picture types. Base views use standard encoding parameters, while additional views utilize prediction parameters derived from base views and temporal information. This parameter variation allows the system to maintain adequate image quality for all views while achieving superior compression efficiency compared to uniform encoding.
Data Source
AI summary
A multi-view video prediction method and a multi-view video prediction restoring method. The multi-view video prediction method includes generating a base layer image stream including residual values of I-picture base view key pictures and base view images of a base view by performing inter prediction between the base view images; and generating an enhancement layer image stream comprising residual values of additional view images of an additional view by performing inter-view prediction for predicting the additional view images with reference to the base view images, performing inter prediction for predicting a different additional view key picture with reference to an additional view key picture from among the additional view images, and performing inter prediction for predicting an additional view image other than the additional view key picture with reference to the additional view images.


