Variable Transform Dimensions in Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video stream encoding techniques are inefficient in reducing data size due to fixed transform dimensions for motion compensated prediction residuals, which increases computational complexity and does not optimize rate distortion cost effectively.
Innovation Solution
The method involves encoding video blocks with variable-sized transforms based on motion information, initially encoding sub-blocks with a larger size, then switching to smaller-sized transforms if a lower rate distortion value is achieved, and combining sub-block residuals using motion information to form composite blocks for further encoding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If fixed transform dimensions are used for motion compensated prediction residuals, then encoding simplicity is maintained, but data size reduction efficiency is insufficient and computational complexity increases
Solution Approach 1:
The patent applies dynamics by making transform dimensions variable rather than fixed. The transform size is dynamically adjusted based on motion information characteristics, allowing the encoding system to adapt transform dimensions to the actual content being encoded. This resolves the contradiction by enabling efficient data reduction when conditions warrant larger transforms while avoiding unnecessary complexity when smaller transforms suffice.
Solution Approach 2:
The patent changes the parameter of transform dimension from a fixed value to a variable parameter determined by motion information analysis. By selecting transform sizes based on motion characteristics, the system optimizes the balance between compression efficiency and computational complexity, improving data size reduction without uniformly increasing complexity across all encoding scenarios.
2Manufacturing precision
If fixed transform dimensions are used, then encoding process simplicity is maintained, but rate distortion cost optimization is insufficient
Solution Approach 1:
The patent optimizes rate distortion cost by changing transform dimension parameters based on motion information characteristics. Different motion patterns are matched with appropriate transform sizes to minimize distortion for a given bit rate, achieving better compression efficiency without requiring complex optimization algorithms.
Solution Approach 2:
The patent applies local quality by allowing different transform dimensions to be applied to different regions or blocks based on their local motion characteristics. This enables precise optimization of rate distortion cost for each region while maintaining overall encoding process manageability through localized rather than global complexity.
3Productivity
If larger transform size is used initially, then potential compression efficiency is achieved, but computational overhead increases without guaranteed improvement
Solution Approach 1:
The patent uses dynamics to adjust transform size based on motion information characteristics rather than always applying the largest transform. This adaptive approach achieves compression efficiency when larger transforms are beneficial while avoiding unnecessary computational overhead when smaller transforms are sufficient, resolving the time-cost tradeoff.
Solution Approach 2:
The patent applies partial action by selectively using larger transform sizes only when motion information characteristics indicate potential compression benefits, rather than uniformly applying excessive transform sizes to all blocks. This avoids unnecessary computational overhead while capturing compression efficiency where it matters.
Data Source
AI summary
Coding efficiency may be improved by subdividing a block into smaller sub-blocks for prediction. A first rate distortion value of a block optionally partitioned into smaller prediction sub-blocks of a first size is calculated using respective inter prediction modes and transforms of the first size. The residuals are used to encode the block using a transform of a second size smaller than the first size, generating a second rate distortion value. The values are compared to determine whether coding efficiency gains may result from inter predicting the smaller, second size sub-blocks. If so, the block is encoded by generating prediction residuals for the second size sub-blocks, and neighboring sub-blocks are grouped, where possible, based on common motion information. Each resulting composite residual block is transformed by a transform of the same size to generate another rate distortion value. The encoded block with the lowest rate distortion value is used.


