Multilayer video signal encoding/decoding method and device
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video compression techniques struggle to efficiently encode and decode high-resolution, high-quality multi-layer video signals, particularly in determining reference layers and blocks for inter-layer prediction, which is crucial for effective texture information derivation and upsampling in multi-layer video processing.
Innovation Solution
A method and apparatus that determine a candidate reference picture for a current picture by using a maximum temporal indicator or temporal ID, derive the number of active references, acquire a reference layer ID, and perform inter-layer prediction using the active reference picture, incorporating spatial offset information to determine a reference block for accurate prediction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If high-resolution, high-quality video data is transmitted or stored using existing media and compression techniques, then video quality is maintained, but transmission cost and storage cost increase
Solution Approach 1:
The patent applies parameter changes by utilizing inter-layer prediction techniques that reference multiple temporal layers (different time points) of video data. By predicting current picture blocks from reference blocks at different temporal positions, the method reduces the number of bits required to encode high-resolution video, thereby lowering transmission and storage costs while maintaining video quality
Solution Approach 2:
The patent employs preliminary action through the use of reference pictures from previous temporal layers that are prepared and stored in advance. These reference pictures are used to predict current picture content, allowing the encoder to transmit fewer bits for the current high-resolution video while still achieving high reconstruction quality
2Productivity
If inter-layer prediction is performed using candidate reference pictures from reference layers, then texture information derivation efficiency is improved, but complexity in determining reference layers and blocks increases
Solution Approach 1:
The patent applies segmentation by dividing the reference picture selection process into distinct stages: first identifying candidate reference pictures from multiple temporal layers, then selecting the optimal reference picture based on prediction accuracy. This segmented approach systematically manages the complexity of inter-layer prediction while improving texture derivation efficiency
Solution Approach 2:
The patent employs dynamics by making the reference layer and reference block selection adaptive based on picture content and temporal characteristics. The system dynamically adjusts which temporal layer is used as reference and which blocks are selected, optimizing prediction accuracy for different video content types and motion patterns
3Measurement precision
If multiple candidate reference pictures are evaluated for inter-layer prediction, then prediction accuracy is improved, but processing time and computational complexity increase
Solution Approach 1:
The patent applies partial action by evaluating only a selected subset of candidate reference pictures from available temporal layers rather than all possible candidates. The method strategically chooses reference pictures from specific temporal positions that are most likely to provide accurate prediction, reducing processing time while maintaining high prediction accuracy
Solution Approach 2:
The patent employs feedback mechanisms by using prediction error metrics to evaluate the quality of candidate reference pictures and adjust subsequent reference picture selection. This feedback-driven approach ensures high prediction accuracy while minimizing the number of candidates that need to be evaluated, thus reducing processing time
Data Source
AI summary
A method for decoding a multilayer video signal, according to the present invention, is characterized by: selecting candidate reference pictures for a current picture from among corresponding pictures in one or more reference layers by using either a maximum temporal level indicator for a current layer or a temporal level identifier for the current picture belonging to the current layer; inducing the number of active reference layer pictures for the current picture on the basis of the number of the candidate reference pictures; obtaining a reference layer identifier on the basis of the number of the active reference layer pictures; determining an active reference picture for use in the interlayer prediction of the current picture by using the reference layer identifier; and performing interlayer prediction on the current picture by using the active reference picture.


