Adaptive KLT Matrix Signaling for Efficient Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding standards, such as VVC, HEVC, and AVS, do not achieve optimal coding efficiency due to limitations in transform coding methods, particularly in adapting to diverse video content characteristics and complexity in training and signaling of Karhunen-Loève Transform (KLT) matrices.
Innovation Solution
Implementing adaptive Karhunen-Loève Transform (KLT) matrices derived from video sequences, enabling their signaling in sequence, picture, or slice headers, and applying Low-Frequency Non-Separable Transforms (LFNST) with reduced complexity, along with enhanced Multiple Transform Selection (MTS) and Subblock Transforms (SBT) to improve coding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If adaptive KLT matrices are derived and signaled for each picture or slice, then coding efficiency is improved, but device complexity and signaling overhead increase
Solution Approach 1:
The patent segments the video sequence into different hierarchical levels (sequence-level, picture-level, slice-level) for KLT matrix derivation and signaling. This allows adaptive transform coding to be applied at appropriate granularities, balancing coding efficiency gains against signaling overhead and complexity at each level.
Solution Approach 2:
The patent implements dynamic selection of KLT matrices based on video content characteristics. The encoder derives and selects appropriate KLT matrices adaptively for different pictures or slices, allowing the transform to dynamically adapt to local content properties rather than using a fixed transform throughout the sequence.
2Adaptability or versatility
If KLT matrices are derived from each picture, then adaptability to video content is improved, but computational complexity increases
Solution Approach 1:
The patent applies different KLT matrices to different local regions (pictures or slices) based on their specific content characteristics. This local adaptation allows each region to be transformed using a matrix optimized for its particular statistical properties, improving coding efficiency without requiring a single complex matrix for the entire sequence.
Solution Approach 2:
The patent applies KLT derivation and signaling selectively at picture or slice levels rather than at every block level, providing sufficient adaptability to capture local content variations while avoiding excessive computational complexity that would result from more granular application.
3Quantity of substance
If transform coding is applied to compress video data, then bit rate is reduced, but coding efficiency is limited by existing transform methods
Solution Approach 1:
The patent changes the transform parameters by deriving and applying Karhunen-Loève Transform matrices that are adapted to the specific statistical properties of the video content. This parameter adaptation allows for more efficient energy compaction in the transform domain, achieving better coding efficiency at the same bit rate or lower bit rate for the same quality.
Data Source
AI summary
The disclosure generally includes a device and methods for video decoding and encoding. The methods include deriving, a karhunen-loève transform (KLT) matrix from a video sequence, each picture in the video sequence, or each slice in the video sequence, generating an adaptive KLT matrix signal, and signaling the adaptive KLT matrix signal in a sequence parameter set (SPS) header, a picture parameter set (PPS) header, and/or slice header.


