Selective Transform Disable for Inter-Predicted Video Blocks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding systems face increased encoding time and signaling overhead without significant coding gains due to certain techniques used for certain types of coding units, particularly with inter-prediction methods like affine motion compensation, combined inter and intra prediction, triangular partition, and geometric merge, which do not benefit from transform coding operations such as multiple transform selection (MTS) and transform skip (TrSkip).
Innovation Solution
A video coding apparatus that determines whether specific inter-prediction techniques are used for coding blocks and disables operations like MTS and TrSkip for these techniques, skipping the rate distortion search and using separable transforms like DCT2 for coding blocks predicted using affine motion compensation, combined inter and intra prediction, triangular partition, and geometric merge, thereby reducing processing complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If transform coding operations (MTS, TrSkip) are applied to all coding blocks, then coding gain is improved, but encoding time and signaling overhead increase
Solution Approach 1:
The patent applies transform coding operations selectively based on the prediction mode of each coding block. Specifically, MTS and TrSkip are disabled for certain inter-prediction techniques (affine motion compensation, combined inter and intra prediction, triangular partition, geometric merge) while remaining enabled for other modes. This local differentiation ensures coding gain is achieved where beneficial without incurring unnecessary encoding overhead.
2Manufacturing precision
If transform coding operations (MTS, TrSkip) are applied to all coding blocks, then coding gain is improved, but signaling overhead increases
Solution Approach 1:
The patent implements local quality by conditionally enabling transform coding operations based on prediction mode. For coding blocks using affine motion compensation, combined inter and intra prediction, triangular partition, or geometric merge, the MTS and TrSkip operations are disabled, eliminating the need to signal MTS indices or transform skip flags for these blocks. This reduces signaling overhead while maintaining coding efficiency for modes where transform operations provide minimal benefit.
3Measurement precision
If rate distortion search is performed for multiple candidate transforms, then transform selection accuracy is improved, but processing complexity increases
Solution Approach 1:
The patent extracts and removes the rate distortion search process for multiple candidate transforms from the encoding pipeline for specific prediction modes. By disabling MTS for affine motion compensation, combined inter and intra prediction, triangular partition, and geometric merge modes, the patent eliminates the computationally intensive RD search and transform evaluation steps for these modes, reducing processing complexity while accepting that these specific modes do not benefit significantly from transform selection.
Data Source
AI summary
Some operations associated with transform decoding may provide coding gains for intra-predicted coding blocks but not for coding blocks predicted using certain inter-prediction tools or techniques. These operations may include, for example, multiple transform selection (MTS) and/or transform skip, and the inter-prediction tools or techniques may include one or more of affine motion compensation, combined inter and intra prediction (CIIP), a triangular partition mode (TPM), or a geometric merge mode (GEO). Thus, systems, methods, and instrumentalities associated with versatile video coding may be configured such that the aforementioned operations associated with transform decoding may be disabled for coding blocks that are predicted using one or more the inter-prediction tools or techniques described herein. Many benefits may be derived from disabling these operations including, for example, reduction of encoding time and/or signaling overhead.


