Low-Delay Hierarchical B GOP Structure for Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding systems face challenges in achieving high coding efficiency and low delay while managing computational complexity, especially in handling moving objects and panned/zoomed scenes, which require higher bitrates due to substantial prediction residues.
Innovation Solution
A low-delay hierarchical B Group of Pictures (GOP) structure is introduced, where pictures are divided into temporal layers with I-pictures only in the lowest layer and low-delay B-pictures used throughout, allowing pictures in lower layers to reference only prior pictures, and using smaller quantization parameters for lower temporal layers to enhance system performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If IBBP GOP structure with bi-directional prediction is used, then coding efficiency is improved, but computational complexity increases due to bi-directional motion estimation
Solution Approach 1:
The video sequence is segmented into multiple temporal layers (Layer 0, Layer 1, Layer 2, etc.), where each layer contains specific picture types. Layer 0 contains only I-pictures, Layer 1 contains low-delay B-pictures referencing Layer 0, Layer 2 contains P-pictures referencing Layer 1, and higher layers follow similar patterns. This segmentation allows different layers to use different prediction strategies, reducing overall computational complexity while maintaining coding efficiency.
Solution Approach 2:
The patent introduces dynamic adaptation of prediction modes across temporal layers. Lower layers use simpler prediction (I-pictures only), while higher layers use more complex prediction (low-delay B-pictures with bidirectional reference to past and future layers). This dynamic structure allows the system to adapt computational complexity to the specific needs of different temporal resolutions, achieving better rate-distortion performance without uniformly high complexity throughout.
2Loss of time
If IPPP GOP structure with forward prediction is used, then processing delay is reduced, but coding efficiency decreases compared to bi-directional prediction
Solution Approach 1:
The video sequence is segmented into multiple temporal layers (Layer 0, Layer 1, Layer 2, etc.), where each layer contains specific picture types. Layer 0 contains only I-pictures, Layer 1 contains low-delay B-pictures referencing Layer 0, Layer 2 contains P-pictures referencing Layer 1, and higher layers follow similar patterns. This segmentation allows different layers to use different prediction strategies, reducing overall computational complexity while maintaining coding efficiency.
Solution Approach 2:
The patent introduces dynamic adaptation of prediction modes across temporal layers. Lower layers use simpler prediction (I-pictures only), while higher layers use more complex prediction (low-delay B-pictures with bidirectional reference to past and future layers). This dynamic structure allows the system to adapt computational complexity to the specific needs of different temporal resolutions, achieving better rate-distortion performance without uniformly high complexity throughout.
3Device complexity
If I-picture only processing is used, then computational complexity is low, but coding efficiency is poor
Solution Approach 1:
The video sequence is segmented into multiple temporal layers (Layer 0, Layer 1, Layer 2, etc.), where each layer contains specific picture types. Layer 0 contains only I-pictures, Layer 1 contains low-delay B-pictures referencing Layer 0, Layer 2 contains P-pictures referencing Layer 1, and higher layers follow similar patterns. This segmentation allows different layers to use different prediction strategies, reducing overall computational complexity while maintaining coding efficiency.
Solution Approach 2:
The patent introduces dynamic adaptation of prediction modes across temporal layers. Lower layers use simpler prediction (I-pictures only), while higher layers use more complex prediction (low-delay B-pictures with bidirectional reference to past and future layers). This dynamic structure allows the system to adapt computational complexity to the specific needs of different temporal resolutions, achieving better rate-distortion performance without uniformly high complexity throughout.
Data Source
AI summary
A method and apparatus for encoding a video sequence comprising a plurality of pictures are disclosed. In video coding systems, the temporal redundancy is exploited using motion compensated prediction. The video sequence is often organized into multiple GOP (group of pictures) where different types of GOP may be used. In conventional coding systems, IPPP and IBBP GOP structure is often used. In H.264/AVC and the emerging High Efficiency Video Coding (HEVC), hierarchical GOP structure, including hierarchical P GOP structure and hierarchical B GOP structure, has been introduced to allow temporal scalability. Furthermore, low-delay IBBB GOP structure has been also introduced, for low-delay application. In the present invention, a low-delay hierarchical B GOP structure is disclosed. The new structure uses low-delay B-pictures only so as to minimize the processing delay while the hierarchical structure provides the temporal scalability. The low-delay hierarchical B GOP structure has been shown to result in substantial improvement in coding efficiency.


