Affine Motion Fields Using Control Point Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video compression systems, such as HEVC and VVC, face inefficiencies in inter-prediction due to limitations in motion representation, particularly for complex motions like zoom, rotation, and irregular movements, which are not adequately captured by traditional translation-based motion models.
Innovation Solution
The introduction of subblock-based affine motion modes in VVC allows for more accurate representation of motion by defining affine motion fields using control point motion vectors, enabling better prediction and compression efficiency for complex video content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional translation-based motion models are used, then encoder and decoder complexity is kept simple, but motion representation accuracy for complex motions is insufficient
Solution Approach 1:
The patent divides the motion representation into multiple components: translation motion vectors for block-level movement, and affine motion parameters (rotation, scaling, shearing) for geometric transformations. This segmentation allows complex motions to be represented by combining simpler motion models, improving accuracy without requiring a single complex monolithic model.
Solution Approach 2:
The patent extends motion representation from 2D translation vectors to 2D affine transformation space by adding rotation and scaling dimensions. This dimensional expansion enables the model to capture complex motions like zooming and rotating objects, transforming the motion model from simple linear translation to comprehensive geometric transformation.
2Reliability
If subblock-based affine motion modes are introduced, then prediction accuracy for complex motions is improved, but encoder and decoder complexity increases
Solution Approach 1:
The patent applies different motion models to different regions (blocks) of the video data. Subblocks can use affine motion modes when complex motion is detected, while simpler blocks use basic translation models. This local adaptation ensures high prediction accuracy for complex motions while maintaining simplicity for straightforward motion, balancing reliability and complexity.
Solution Approach 2:
The patent implements dynamic motion model selection where the encoder and decoder can adaptively choose between translation-only models and affine motion models based on the actual motion characteristics of each block. This dynamic switching allows the system to improve prediction accuracy for complex motions while avoiding unnecessary complexity when simple motion models suffice.
3Productivity
If affine motion fields with control point motion vectors are used, then compression efficiency is improved, but processing time and computational complexity increase
Solution Approach 1:
The patent applies affine motion modeling selectively only to blocks that exhibit complex motion characteristics, rather than universally to all blocks. This partial application reduces the overall computational burden while maintaining high compression efficiency for the most challenging frames, achieving a balance between productivity improvement and processing time cost.
Data Source
Figure 1~2
Figure 3~4
Figure 5
AI summary
The present application relates to a method of decoding a video picture by means of motion compensated temporal bi-prediction of an inter coded block using two reference video pictures in two separate reference video picture lists, and two affine motion fields defined by at least two control point motion vectors, denoted CPMVs, associated to each reference picture, refined CPMVs being obtained as output of the following first step (171) or, of the second step (172) or of the second step (172) that follows the first step (171), wherein the first or second step is/are bypassed according to binary values.