Context Modeling for LFNST Signaling in Video Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video encoding and decoding technologies face inefficiencies in reducing the signaling overhead of Low-Frequency Non-separable Transformation (LFNST) indices/flags, which affects the coding efficiency of context-adaptive binary arithmetic coding (CABAC) in video compression standards.

Innovation Solution

The introduction of new context models that allow for selective context-based coding of LFNST indices/flags based on parameters such as color component, block size, and number of non-zero transform coefficients, enabling more efficient signaling and coding of video data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a single context model is used for CABAC coding of LFNST indices/flags, then the device complexity is reduced, but the coding efficiency and bitrate optimization are insufficient

Engineering Contradiction:
Improvecoding efficiencyVSAvoidcontext modeling complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies local quality by creating separate context models for different color components (luma and chroma). Instead of using a single context model for all LFNST indices/flags, the invention divides the context modeling into component-specific models that capture the unique statistical characteristics of each color component, thereby improving coding efficiency without requiring excessive complexity

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If multiple context models are used for different color components, then the bitrate is reduced and coding efficiency is improved, but the device complexity increases

Engineering Contradiction:
ImprovebitrateVSAvoidcontext modeling complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the context modeling into distinct models for different color components (luma and chroma). This segmentation allows each context model to be optimized for its specific component's statistical properties, enabling better compression and lower bitrate while keeping the complexity manageable through structured organization of the context models

Inventive Principle:
Principle #1Segmentation

3Productivity

If LFNST indices/flags are coded without context adaptation, then the ease of operation is improved, but the signaling overhead is reduced and coding efficiency is improved

Engineering Contradiction:
Improvecoding efficiencyVSAvoidsignaling overhead
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent implements dynamic context adaptation where the context model is updated based on the actual data characteristics. The system dynamically adjusts the context probabilities during encoding/decoding operations, allowing efficient signaling of LFNST indices/flags while adapting to the specific patterns in the video data, thus reducing signaling overhead without losing information

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11949870B2Context modeling for low-frequency non-separable transformation signaling for video coding
Publication Date: 2024.04.02 QUALCOMM INC
  • US11949870B2 patent drawing
  • US11949870B2 patent drawing
  • US11949870B2 patent drawing

AI summary

An example method includes determining a color component of a unit of video data; determining, based at least on the color component, a context for context-adaptive binary arithmetic coding (CABAC) a syntax element that specifies a value of a low-frequency non-separable transform (LFNST) index for the unit of video data; CABAC decoding, based on the determined context and via a syntax structure for the unit of video data, the syntax element that specifies the value of the LFNST index for the unit of video data; and inverse-transforming, based on a transform indicated by the value of the LFNST index, transform coefficients of the unit of video data.