Cross-Layer Alignment Signaling in Multi-Layer Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding standards do not efficiently indicate cross-layer alignment of picture types, leading to potential loss of coding efficiency and increased random access delays, especially in scenarios requiring frequent random access across layers.
Innovation Solution
The method involves signaling a syntax element, such as the vps_cross_layer irap_align flag, to indicate cross-layer alignment within the Video Parameter Set (VPS), allowing network abstraction layer unit types to be set uniformly across all pictures in an access unit without requiring entropy decoding, thereby facilitating efficient random access and layer switching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If cross-layer alignment is not indicated efficiently in current video coding standards, then coding efficiency is maintained through standard processing, but random access delays increase and coding efficiency is lost in multi-layer scenarios
Solution Approach 1:
The patent applies preliminary action by signaling the cross-layer alignment indication in the VPS (Video Parameter Set) before actual decoding occurs. This allows receiving devices to pre-determine whether cross-layer alignment is applied to IRAP pictures in advance, enabling efficient random access without requiring full decompression or entropy decoding during the access operation.
Solution Approach 2:
The patent extracts the cross-layer alignment indication from the entropy-coded bitstream and places it in the VPS syntax structure. This extraction allows the indication to be accessible without entropy decoding, separating the alignment information from the main video data decoding process and enabling faster access.
2Loss of information
If cross-layer alignment indication is embedded in entropy-coded bitstream, then comprehensive information is available, but decoding complexity and processing time increase
Solution Approach 1:
The patent extracts the cross-layer alignment indication from the entropy-coded bitstream and embeds it directly in the VPS syntax. This allows the indication to be read without performing entropy decoding operations, significantly reducing decoding complexity while maintaining full information availability about cross-layer alignment.
Solution Approach 2:
The indication is prepared and placed in the VPS during the encoding phase, before the actual video decoding process. This preliminary action ensures that receiving devices can access the alignment information immediately upon receiving the bitstream without requiring complex entropy decoding operations.
3Device complexity
If cross-layer alignment is applied without explicit indication, then bitstream simplicity is maintained, but random access performance degrades in multi-layer scenarios
Solution Approach 1:
The patent uses the VPS structure, which is already a universal component in multi-layer video coding, to carry the cross-layer alignment indication. This multi-functional approach allows the VPS to serve both its traditional parameter setting role and the new function of indicating cross-layer alignment, without adding separate complex signaling structures.
Solution Approach 2:
The patent introduces a new syntax parameter (cross-layer alignment indication) in the VPS that changes the behavior of picture type handling. When this parameter indicates cross-layer alignment is applied, all IRAP pictures in the access unit are treated as having the same NAL unit type, enabling efficient layer switching without modifying the fundamental bitstream structure.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
In one example, the disclosure is directed to techniques that include receiving a bitstream comprising at least a syntax element, a first network abstraction layer unit type, and a coded access unit comprising a plurality of pictures. The techniques further include determining a value of the syntax element which indicates whether the access unit was coded using cross-layer alignment. The techniques further include determining the first network abstraction layer unit type for a picture in the access unit and determining whether the first network abstraction layer unit type equals a value in a range of type values. The techniques further include setting a network abstraction layer unit type for all other pictures in the coded access unit to equal the value of the first network abstraction layer unit type if the first network abstraction layer unit type is equal to a value in the range of type values.