Adaptive KLT Matrix Signaling for Efficient Video Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding standards, such as VVC, HEVC, and AVS, do not achieve optimal coding efficiency due to limitations in transform coding methods, particularly in adapting to diverse video content characteristics and complexity in training and signaling of Karhunen-Loève Transform (KLT) matrices.

Innovation Solution

Implementing adaptive Karhunen-Loève Transform (KLT) matrices derived from video sequences, enabling their signaling in sequence, picture, or slice headers, and applying Low-Frequency Non-Separable Transforms (LFNST) with reduced complexity, along with enhanced Multiple Transform Selection (MTS) and Subblock Transforms (SBT) to improve coding efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If adaptive KLT matrices are derived and signaled for each picture or slice, then coding efficiency is improved, but device complexity and signaling overhead increase

Engineering Contradiction:
Improvecoding efficiencyVSAvoidtransform training complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the video sequence into different hierarchical levels (sequence-level, picture-level, slice-level) for KLT matrix derivation and signaling. This allows adaptive transform coding to be applied at appropriate granularities, balancing coding efficiency gains against signaling overhead and complexity at each level.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic selection of KLT matrices based on video content characteristics. The encoder derives and selects appropriate KLT matrices adaptively for different pictures or slices, allowing the transform to dynamically adapt to local content properties rather than using a fixed transform throughout the sequence.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If KLT matrices are derived from each picture, then adaptability to video content is improved, but computational complexity increases

Engineering Contradiction:
Improveadaptability to video contentVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies different KLT matrices to different local regions (pictures or slices) based on their specific content characteristics. This local adaptation allows each region to be transformed using a matrix optimized for its particular statistical properties, improving coding efficiency without requiring a single complex matrix for the entire sequence.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent applies KLT derivation and signaling selectively at picture or slice levels rather than at every block level, providing sufficient adaptability to capture local content variations while avoiding excessive computational complexity that would result from more granular application.

Inventive Principle:
Principle #16Partial or excessive action

3Quantity of substance

If transform coding is applied to compress video data, then bit rate is reduced, but coding efficiency is limited by existing transform methods

Engineering Contradiction:
Improvebit rateVSAvoidcoding efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent changes the transform parameters by deriving and applying Karhunen-Loève Transform matrices that are adapted to the specific statistical properties of the video content. This parameter adaptation allows for more efficient energy compaction in the transform domain, achieving better coding efficiency at the same bit rate or lower bit rate for the same quality.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250211789A1Methods and apparatus for transform training and coding
Publication Date: 2025.06.26 BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
  • US20250211789A1 patent drawing
  • US20250211789A1 patent drawing
  • US20250211789A1 patent drawing

AI summary

The disclosure generally includes a device and methods for video decoding and encoding. The methods include deriving, a karhunen-loève transform (KLT) matrix from a video sequence, each picture in the video sequence, or each slice in the video sequence, generating an adaptive KLT matrix signal, and signaling the adaptive KLT matrix signal in a sequence parameter set (SPS) header, a picture parameter set (PPS) header, and/or slice header.