Implicit Transform Selection in Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding standards face challenges in efficiently managing multiple transforms, leading to increased computational complexity and overhead bits, particularly in the decision-making processes for implicit multiple transform sets and subblock transforms, which affect coding efficiency and throughput.
Innovation Solution
The proposed solution involves making decisions on transform applications based on decoded coefficients and characteristics of video blocks, allowing for implicit selection of transforms and enabling multiple transform sets within specific coding modes, regardless of sequence or picture-level settings, to optimize conversions between video blocks and bitstream representations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If explicit multiple transform set (MTS) process is used with transform indices, then transform selection flexibility is improved, but computational complexity and overhead bits increase
Solution Approach 1:
The patent extracts the transform selection decision from the explicit signaling process and embeds it within the implicit MTS decision-making framework. By deriving transform application decisions from decoded coefficients and block characteristics rather than separate transform indices, the system maintains adaptability while reducing computational overhead and complexity.
Solution Approach 2:
The patent merges the transform selection function with the implicit MTS decision process. Instead of treating transform index selection as a separate explicit process, it combines transform decisions with the existing implicit MTS logic that operates on decoded coefficients and block characteristics, thereby reducing overall system complexity while preserving flexibility.
2Productivity
If implicit multiple transform set (MTS) process is applied regardless of enablement flags, then coding efficiency is improved, but control overhead increases
Solution Approach 1:
The patent enables the implicit MTS process to self-determine its application based on intrinsic block characteristics and decoded coefficients. By making the transform decision autonomous and context-dependent rather than flag-controlled, the system improves coding efficiency through adaptive transform selection while minimizing control overhead by eliminating redundant enablement signaling.
3Adaptability or versatility
If multiple transform sets are enabled at sequence or picture level, then transform adaptability is improved, but bitstream overhead increases
Solution Approach 1:
The patent applies transform adaptability locally at the block level rather than globally at sequence or picture level. By making transform decisions based on local decoded coefficients and block characteristics, the system achieves necessary adaptability for each block while avoiding the bitstream overhead associated with global transform enablement flags and extensive transform index signaling.
4Productivity
If transform decisions are based on decoded coefficients and block characteristics, then coding efficiency is improved, but decision-making complexity increases
Solution Approach 1:
The patent performs preliminary analysis of decoded coefficients and block characteristics during the decoding process to pre-determine transform application decisions. By preparing and evaluating transform suitability metrics early in the decoding pipeline based on already-decoded data, the system improves coding efficiency while managing decision-making complexity through structured preliminary assessment rather than complex real-time decisions.
Data Source
AI summary
Devices, systems and methods for digital video coding, which includes using multiple transforms, are described. In a representative aspect, a method for video processing includes making a decision, based on one or more decoded coefficients and in an absence of one or more transform indices, regarding an application of a transform to a current block of a video, and performing, based on the decision, a conversion between the current block and a bitstream representation of the video.


