VVC Decoder Configuration Records for Complete Sublayer and PTL Signaling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding standards, such as VVC, face challenges in efficiently signaling decoder configuration information and random access recovery points, leading to incomplete or unspecified parameters that hinder effective content selection and decoding processes.

Innovation Solution

The VVC decoder configuration record is modified to include additional parameters like decoded picture buffer size, maximum picture output reordering, and latency flags before or after PTL records, ensuring byte-alignment and reserved bits for proper signaling, and the 'roll' sample group semantics are clarified to accurately describe layer correlations and access points.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If the VVC decoder configuration record includes additional parameters like decoded picture buffer size, maximum picture output reordering, and latency flags, then the completeness of decoder configuration information is improved, but the complexity of the configuration record structure increases

Engineering Contradiction:
Improvecompleteness of decoder configuration informationVSAvoidcomplexity of configuration record structure
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The decoder configuration information is segmented into multiple distinct syntax elements organized in a hierarchical structure. Each parameter (decoded picture buffer size, maximum picture output reordering, latency flags) is separated into its own syntax element, allowing independent encoding and decoding. This segmentation enables the configuration record to accommodate comprehensive parameters while maintaining structured organization and manageable complexity.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If the VVC decoder configuration record specifies detailed parameters for sublayers and PTL records, then the precision of decoding control is improved, but the amount of signaling data increases

Engineering Contradiction:
Improveprecision of decoding controlVSAvoidamount of signaling data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The configuration record implements local quality by providing specific syntax elements for different sublayer levels. Each sublayer can have its own PTL (Profile Tier Level) records with precisely tailored parameters. This allows the decoding control precision to be optimized locally for each sublayer's specific requirements rather than using a uniform configuration for all sublayers, thereby reducing overall signaling data while maintaining high precision where needed.

Inventive Principle:
Principle #3Local quality

3Reliability

If the configuration record includes byte-alignment requirements and reserved bits, then the reliability of data parsing is improved, but the overhead of the configuration structure increases

Engineering Contradiction:
Improvereliability of data parsingVSAvoidoverhead of configuration structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The configuration record structure incorporates preliminary actions by pre-defining byte-alignment requirements and reserved bits positions before actual data parsing occurs. The syntax elements are arranged with explicit alignment rules (e.g., 32-bit alignment) and reserved bits are allocated in advance at specific positions. This preliminary structuring ensures reliable parsing by eliminating ambiguity during decoding, while the systematic placement of alignment and reserved elements minimizes overall structural overhead.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12375696B2Decoder configuration information in VVC video coding
Publication Date: 2025.07.29 BYTEDANCE INC
  • US12375696B2 patent drawing
  • US12375696B2 patent drawing
  • US12375696B2 patent drawing

AI summary

A mechanism for processing video data is disclosed. A conversion is performed between a visual media data and a visual media data file. The visual media data file includes a versatile video coding (VVC) decoder configuration record and a plurality of pictures coded into one or more sublayers, wherein the VVC decoder configuration record comprises a number of the one or more sublayers and one or more VVC profile tier level (PTL) records for the one or more sublayers based on the number of the one or more sublayers.