Step-wise Temporal Sub-layer Access in Scalable Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding standards, such as SHVC and MV-HEVC, face limitations in scalability designs, including restrictions on prediction hierarchies across layers and the inability to switch temporal levels at the lowest level, which hinder rate-distortion performance and flexible sub-layer access.
Innovation Solution
The introduction of step-wise temporal sub-layer access pictures allows for layer-wise initialization and bitrate adaptation by selectively decoding pictures on the lowest temporal sub-layer, enabling flexible access and decoding of temporal sub-layers within a bitstream.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If pictures of an access unit are required to have the same temporal level, then decoding structure is simplified, but prediction hierarchies cannot be determined differently across layers and sub-layer up-switch points are limited
Solution Approach 1:
The patent divides the bitstream into multiple temporal sub-layers (first temporal sub-layer and second temporal sub-layer) with different temporal levels. This segmentation allows independent control of prediction hierarchies in each sub-layer, enabling frequent sub-layer up-switch points while maintaining simplified decoding within each segment.
Solution Approach 2:
The patent introduces a new dimension by allowing temporal level switching within the same access unit across different temporal sub-layers. This enables the system to vary temporal levels dynamically without requiring all pictures in an access unit to have the same temporal level, thus improving adaptability while maintaining structural organization.
2Stability of the object's composition
If temporal level switch pictures are not allowed at the lowest temporal level, then decoding stability is maintained, but access points for partial temporal level decoding cannot be indicated
Solution Approach 1:
The patent introduces a step-wise temporal sub-layer access picture type that can be indicated at the lowest temporal level. This preliminary action enables the decoder to prepare for and initiate decoding at specific temporal levels without requiring temporal level switch pictures, thus maintaining stability while enabling flexible access points.
Solution Approach 2:
The step-wise temporal sub-layer access picture acts as an intermediary mechanism that bridges the gap between maintaining decoding stability and enabling flexible access points. It allows the system to indicate access points for partial temporal level decoding without introducing temporal level switching at the lowest level, thus resolving the contradiction through a mediating structure.
3Manufacturing precision
If frequent sub-layer up-switch points are used, then rate-distortion performance improves, but decoding complexity increases
Solution Approach 1:
By segmenting the bitstream into multiple temporal sub-layers with clear boundaries and access points, the patent enables frequent sub-layer up-switch points to be implemented with manageable decoding complexity. Each sub-layer can be decoded independently to a certain extent, reducing the overall complexity burden.
Solution Approach 2:
The patent introduces dynamic temporal sub-layer switching capabilities that allow the system to adaptively choose when to switch between sub-layers based on rate-distortion optimization. This dynamic approach enables frequent switching points when beneficial while maintaining the ability to simplify decoding when appropriate, thus balancing performance and complexity.
Data Source
AI summary
A method comprising: encoding a first picture on a first scalability layer and on a lowest temporal sub-layer; encoding a second picture on a second scalability layer and on the lowest temporal sub-layer, wherein the first picture and the second picture represent the same time instant, encoding one or more first syntax elements, associated with the first picture, with a value indicating that a picture type of the first picture is other than a step-wise temporal sub-layer access (STSA) picture; encoding one or more second syntax elements, associated with the second picture, with a value indicating that a picture type of the second picture is a step-wise temporal sub-layer access picture; and encoding at least a third picture on a second scalability layer and on a temporal sub-layer higher than the lowest temporal sub-layer.


