Random Access Point Access Units for Multi-Layer Video Decoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding standards struggle to efficiently support scalable video coding, particularly in Versatile Video Coding (VVC), where multi-layer bitstreams require complex decoding capabilities and frequent random access points, leading to inefficiencies in bandwidth usage and decoding complexity.

Innovation Solution

The implementation of random access point access units (RAP AUs) in scalable video coding, which include specific syntax elements and format rules to manage access units, allowing for efficient decoding and reduced complexity in multi-layer bitstreams, especially in VVC, by ensuring each RAP AU contains a picture for each layer and using side information to indicate start access units.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If frequent random access points are implemented in multi-layer bitstreams, then decoding flexibility and reliability are improved, but device complexity and decoding complexity increase

Engineering Contradiction:
Improvedecoding flexibilityVSAvoiddecoding complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The bitstream is divided into access units (AUs) that can be independently decoded, with each AU containing complete picture data for all layers. This segmentation enables frequent random access points while managing complexity through standardized AU boundaries and self-contained structures.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Access units are designed to serve multiple functions: they can start new coded video sequences, support random access, and maintain compatibility across different decoder configurations. The same AU structure works for both single-layer and multi-layer decoders, reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multi-layer bitstreams are used for scalable video coding, then adaptability and video quality are improved, but device complexity and bandwidth requirements increase

Engineering Contradiction:
Improvevideo quality scalabilityVSAvoiddecoder complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The video stream is segmented into layers and access units, allowing decoders to process only the necessary layers based on capabilities. Base layer decoders handle lower layers while enhancement layer decoders add quality, enabling scalability without requiring all decoders to process all layers simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Decoders can selectively decode only the base layer or specific enhancement layers based on device capabilities and bandwidth constraints. This partial action approach allows adaptability across different devices without requiring full multi-layer processing capability in all decoders.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If access units include pictures for all video layers, then decoding reliability is improved, but bandwidth consumption and data volume increase

Engineering Contradiction:
Improvedecoding reliabilityVSAvoidbandwidth consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Access units are segmented to contain complete picture data for random access reliability, but the segmentation allows selective transmission of only necessary layers. Enhanced AUs can include all layers for reliability, while base AUs may transmit only essential layers, optimizing bandwidth usage.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12432386B2Random access point access unit in scalable video coding
Publication Date: 2025.09.30 BYTEDANCE INC
  • US12432386B2 patent drawing
  • US12432386B2 patent drawing
  • US12432386B2 patent drawing

AI summary

Methods, devices, and systems for configuring different access units in scalable video coding are described. In one example aspect, a method of video processing include performing a conversion between a video having one or more pictures in one or more video layers and a bitstream of a video. The bitstream includes a coded video sequence that has one or more access units. The bitstream further includes a first syntax element indicating whether an access unit includes a picture for each video layer making up the coded video sequence.