Inter Prediction Using Transformation Motion Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for high-resolution and high-quality images has led to a need for more efficient image compression techniques to reduce transmission and storage costs.
Innovation Solution
A method and apparatus for improving video coding efficiency by using a transformation prediction model-based inter-prediction method, which derives motion vectors for sub-blocks or sample points based on motion vectors of control points, allowing for accurate prediction even when images are rotated, zoomed, or transformed into a parallelogram.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional inter-prediction methods are used for high-resolution images, then transmission and storage costs increase, but prediction accuracy deteriorates when images are rotated, zoomed, or transformed
Solution Approach 1:
The current block is divided into multiple sub-blocks, and motion vectors are derived for each sub-block independently based on control points. This segmentation allows the prediction to adapt to local variations in motion and transformation, improving accuracy for rotated or zoomed regions while maintaining efficient compression through localized processing.
Solution Approach 2:
The patent uses a transformation prediction model that changes the motion vector parameters based on the geometric transformation (rotation, zoom, parallelogram) of the image block. By deriving motion vectors that account for these transformations, the system maintains prediction accuracy even when the image undergoes complex geometric changes, thereby reducing residual data without sacrificing quality.
2Measurement precision
If motion vectors are derived for each sub-block or sample, then prediction accuracy improves, but computational complexity increases
Solution Approach 1:
The patent applies different levels of motion vector derivation to different regions. Control points are strategically positioned to capture key motion characteristics, and sub-block motion vectors are derived based on these control points rather than computing full motion estimation for every sample. This local quality approach maintains high accuracy where needed while reducing overall computational complexity.
Solution Approach 2:
The transformation prediction model is established in advance using control points, and motion vectors for sub-blocks are derived based on this pre-computed model. This preliminary action of setting up the transformation model allows subsequent sub-block predictions to be made more efficiently, reducing the computational burden compared to performing full motion estimation for each sub-block.
3Productivity
If transformation prediction model is used to handle rotated and zoomed images, then inter-prediction efficiency improves, but the amount of data for residual signals increases
Solution Approach 1:
The patent employs a dynamic transformation prediction model that adapts to the specific geometric transformation (rotation angle, zoom factor, parallelogram skew) of each current block. By dynamically adjusting the motion vectors to match the actual transformation, the prediction more closely matches the original block, thereby reducing the residual data amount while maintaining high coding efficiency for transformed content.
Data Source
AI summary
A video decoding method performed by a decoding apparatus comprises deriving control points (CPs) for the current block; obtaining motion vectors for the CPs; deriving a motion vector of a sub-block or a sample unit in the current block on the basis of the obtained motion vectors; deriving a prediction sample for the current block on the basis of the derived motion vector; and generating a reconstruction sample on the basis of the prediction sample. The method enables effective performance of inter prediction through the motion vectors (transformation prediction), not only when an image in the current block is moved in a plane, but also when the image in the current block is rotated, zoomed in, zoomed out, or transformed into a parallelogram. Accordingly, the amount of data for the residual signal for the current block can be eliminated or reduced, and the overall coding efficiency can be improved.


