Scalable Video Decoding Using Inter-Layer Reference Pictures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for high-resolution and high-quality images leads to increased data amounts, resulting in higher transmission and storage costs, and existing image compression techniques are inadequate for efficiently handling stereographic image content.
Innovation Solution
A method and device for using a picture from a lower layer as an inter-layer reference picture in encoding/decoding a scalable video signal, allowing for inter-layer prediction and upsampling to effectively manage memory and derive texture information of the upper layer.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If high-resolution and high-quality image data is transmitted and stored using existing methods, then image quality is improved, but transmission and storage costs increase
Solution Approach 1:
The image data is divided into multiple layers with different resolutions and qualities. The decoded picture buffer stores multiple layer versions (base layer and enhancement layers), allowing selective transmission and storage of different quality levels. This segmentation enables efficient resource allocation where only necessary layer combinations are transmitted, reducing overall data volume while maintaining quality options.
Solution Approach 2:
The patent introduces a temporal dimension to image storage by maintaining multiple temporal versions of pictures at different resolutions in the decoded picture buffer. Instead of storing only current high-resolution frames, the system stores historical frames at multiple resolution levels across different time points, enabling inter-layer and inter-temporal prediction that reduces the data amount needed for high-quality reconstruction.
2Measurement precision
If multiple reference pictures are stored in memory for high-quality decoding, then decoding accuracy is improved, but memory usage increases
Solution Approach 1:
Different regions of the decoded picture buffer are allocated with different quality levels and storage durations. Base layer pictures are stored with higher fidelity and longer retention, while enhancement layer pictures use lower fidelity storage. The reference picture lists are constructed selectively, using high-quality base layer references for critical prediction areas and lower-quality enhancement references for less critical regions, optimizing memory usage while maintaining decoding accuracy where needed.
Solution Approach 2:
The patent implements a nested structure where enhancement layer pictures are stored within the same buffer framework as base layer pictures, but with different quality characteristics. The reference picture lists are nested hierarchically, with base layer references forming the core and enhancement layer references providing additional detail. This nested organization allows efficient memory sharing between layers, reducing total memory requirements while maintaining access to multiple reference pictures for accurate decoding.
3Productivity
If inter-layer prediction is performed using lower layer pictures, then coding efficiency is improved, but complexity of reference picture management increases
Solution Approach 1:
The patent establishes reference picture lists in advance before the actual prediction process. The decoded picture buffer pre-loads and organizes multiple layer pictures at different temporal identifiers, creating ready-to-use reference lists that simplify the prediction operation. This preliminary organization of reference pictures across layers and time points reduces the computational complexity during actual encoding/decoding, as the prediction engine can directly access pre-arranged reference lists without complex search or selection operations.
Solution Approach 2:
The reference picture lists are dynamically constructed and updated based on the specific encoding/decoding context. The system adaptively selects which lower layer pictures to include in reference lists based on temporal identifiers, picture types, and prediction requirements. This dynamic management allows the system to optimize reference picture selection for each specific case, improving coding efficiency while keeping the management complexity manageable through context-adaptive rules rather than rigid fixed structures.
Data Source
AI summary
A scalable video signal decoding method according to the present invention determines, based on the temporal identifier of a lower layer, whether a corresponding picture of the lower layer is used as an inter-layer reference picture for the current picture of an upper layer, creates a list of reference pictures for the current picture based on the determination, and performs inter-layer prediction on the current block in the current picture based on the created list of reference pictures.


