Level Information Signaling in Video Coding Bitstreams
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding technologies face challenges in signaling level information for intra random access point (TRAP) only and intra-only representations, inter-layer prediction sublayers, and virtual boundaries, which affects decoder capabilities and bitstream conformance in versatile video coding (VVC) standards.
Innovation Solution
The proposed solution involves signaling level information for TRAP-only and intra-only representations through various syntax structures such as the profile, tier, and level (PTL) syntax, SEI messages, and HRD parameters, and modifying constraints for inter-layer prediction and virtual boundary signaling to ensure decoder compatibility and bitstream conformance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If level information is signaled for TRAP-only and intra-only representations using existing PTL syntax, then decoder capability determination is enabled, but bitstream complexity and signaling overhead increase
Solution Approach 1:
The patent extracts level information signaling from the general PTL syntax and creates dedicated syntax structures specifically for TRAP-only and intra-only representations. This separation allows level information to be signaled only when needed for these specific representations, reducing overall bitstream complexity while ensuring decoder capability determination for trick mode playback.
Solution Approach 2:
The patent introduces new intermediate syntax structures (e.g., trick_mode_level_info syntax) that act as mediators between the video content and decoder capabilities. These intermediate structures provide a standardized interface for signaling level information specific to TRAP-only and intra-only representations without affecting the overall PTL syntax complexity.
2Productivity
If constraints are modified for inter-layer prediction information signaling, then signaling efficiency is improved, but compatibility with existing decoders may be reduced
Solution Approach 1:
The patent makes the inter-layer prediction information signaling dynamic by introducing conditional syntax elements that are only present when inter-layer prediction is actually used. The syntax structure adapts based on the presence of reference pictures from other layers, allowing efficient signaling when needed while maintaining compatibility when not needed.
Solution Approach 2:
The patent changes the parameter structure for inter-layer prediction by replacing fixed two-dimensional syntax elements with more flexible signaling mechanisms. The max_tid_il_ref_pics_plus1 parameter is signaled conditionally and with variable precision depending on the actual usage, improving signaling efficiency while maintaining backward compatibility through default values.
3Loss of substance
If virtual boundary signaling is modified to reduce syntax elements, then bitstream size is reduced, but precision in defining virtual boundaries may be lost
Solution Approach 1:
The patent segments the virtual boundary definition into multiple components: a primary virtual boundary position syntax element for the main boundary location, and optional additional syntax elements for refined positioning. This segmentation allows the bitstream to contain only the essential boundary information while providing mechanisms for higher precision when needed, reducing overall bitstream size while maintaining precision where required.
Solution Approach 2:
The patent applies partial action by signaling virtual boundary information at different levels of detail depending on the specific case. For many common cases, a simplified virtual boundary syntax is used that reduces bitstream size, while more complex cases can utilize additional precision syntax elements, achieving an optimal balance between compression and precision.
Data Source
AI summary
Methods and apparatus for video processing are described. The processing may include video encoding, video decoding, or video transcoding. An example video processing method includes performing a conversion between a video and a bitstream of the video including one or more output layer sets according to a format rule. At least one of the one or more output layer sets consists of a trick mode access representation including only intra random access points pictures or only intra-coded pictures. The format rule specifies whether or how level information for the trick mode representation is indicated in the bitstream.


