Temporal Sublayers for Spatial Scalability in Video Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video encoding and decoding technologies, such as HEVC and VVC, face challenges in efficiently supporting spatial scalability without relying on scalability layers, which are not supported in all profiles.
Innovation Solution
The proposed method utilizes temporal layers and reference picture resampling (RPR) to enable spatial scalability in profiles that do not support scalability layers. This involves assigning pictures that would be in different scalability layers to different temporal layers, ensuring that all pictures of one temporal layer have the same spatial resolution, and using RPR for inter prediction across temporal layers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If scalability layers are used to support spatial scalability, then spatial resolution flexibility is improved, but profile compatibility and device complexity are worsened due to lack of support in all profiles
Solution Approach 1:
The patent segments spatial scalability support into multiple temporal layers, where each temporal layer contains pictures at a specific spatial resolution. This allows the decoder to selectively output pictures from different temporal layers based on capability, achieving spatial scalability without requiring scalability layer infrastructure that is not supported in all profiles.
Solution Approach 2:
The patent transitions from using scalability layers (one dimension of organization) to using temporal layers (another dimension) to achieve spatial scalability. By organizing pictures with different spatial resolutions into different temporal layers and using RPR for inter-layer prediction, the patent achieves spatial flexibility through temporal layer structure.
2Adaptability or versatility
If multiple spatial resolutions are decoded simultaneously, then spatial scalability is improved, but decoding complexity and resource usage are worsened
Solution Approach 1:
The patent introduces dynamic control mechanisms including output suppression flags and temporal layer identifiers that allow the decoder to adaptively select which temporal layers to output based on device capabilities and runtime conditions. This dynamic control enables spatial scalability while managing decoding complexity through selective processing.
Solution Approach 2:
The patent changes the operational parameters of the decoder by introducing temporal layer identifiers and output suppression flags that control which pictures are decoded and output. By parameterizing the decoding process to selectively handle different temporal layers, the patent achieves spatial scalability while controlling resource usage.
3Reliability
If all decoded pictures are output, then completeness of output is improved, but resource efficiency and bandwidth usage are worsened
Solution Approach 1:
The patent extracts and separates pictures into different temporal layers based on their spatial resolution characteristics. By taking out low-resolution pictures into dedicated temporal layers, the system enables selective suppression of these layers when high-resolution output is required, improving resource efficiency while maintaining the option for complete output when needed.
Solution Approach 2:
The patent implements a mechanism where low-resolution pictures in lower temporal layers can be discarded (suppressed from output) when high-resolution pictures from higher temporal layers are available. This discarding and recovering approach optimizes resource efficiency by avoiding redundant low-resolution output while preserving the capability to output all pictures when completeness is required.
Data Source
AI summary
A method and apparatus for decoding and outputting one or more pictures from a bitstream is provided. The method includes obtaining an indication that specifies that the decoder should not output pictures belonging to a temporal layer. The method includes decoding at least one picture from a bitstream wherein picture(s) belong to the temporal layer. The method includes decoding at least one picture from a bitstream wherein picture(s) belong to one temporal layer not equal to. The method includes responsive to receiving the indication, suppressing output of the at least one picture. The method includes outputting the at least one picture.


