Affine Motion Fields Using Control Point Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video compression systems, such as HEVC and VVC, face inefficiencies in inter-prediction due to limitations in motion representation, particularly for complex motions like zoom, rotation, and irregular movements, which are not adequately captured by traditional translation-based motion models.

Innovation Solution

The introduction of subblock-based affine motion modes in VVC allows for more accurate representation of motion by defining affine motion fields using control point motion vectors, enabling better prediction and compression efficiency for complex video content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional translation-based motion models are used, then encoder and decoder complexity is kept simple, but motion representation accuracy for complex motions is insufficient

Engineering Contradiction:
Improvemotion representation accuracyVSAvoidmotion model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the motion representation into multiple components: translation motion vectors for block-level movement, and affine motion parameters (rotation, scaling, shearing) for geometric transformations. This segmentation allows complex motions to be represented by combining simpler motion models, improving accuracy without requiring a single complex monolithic model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extends motion representation from 2D translation vectors to 2D affine transformation space by adding rotation and scaling dimensions. This dimensional expansion enables the model to capture complex motions like zooming and rotating objects, transforming the motion model from simple linear translation to comprehensive geometric transformation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If subblock-based affine motion modes are introduced, then prediction accuracy for complex motions is improved, but encoder and decoder complexity increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidencoder and decoder complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies different motion models to different regions (blocks) of the video data. Subblocks can use affine motion modes when complex motion is detected, while simpler blocks use basic translation models. This local adaptation ensures high prediction accuracy for complex motions while maintaining simplicity for straightforward motion, balancing reliability and complexity.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamic motion model selection where the encoder and decoder can adaptively choose between translation-only models and affine motion models based on the actual motion characteristics of each block. This dynamic switching allows the system to improve prediction accuracy for complex motions while avoiding unnecessary complexity when simple motion models suffice.

Inventive Principle:
Principle #15Dynamics

3Productivity

If affine motion fields with control point motion vectors are used, then compression efficiency is improved, but processing time and computational complexity increase

Engineering Contradiction:
Improvecompression efficiencyVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies affine motion modeling selectively only to blocks that exhibit complex motion characteristics, rather than universally to all blocks. This partial application reduces the overall computational burden while maintaining high compression efficiency for the most challenging frames, achieving a balance between productivity improvement and processing time cost.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP4436184A1Decoding video picture data using affine motion fields defined by control point motion vectors
Publication Date: 2024.09.25 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • EP4436184A1 patent drawingFigure 1~2
  • EP4436184A1 patent drawingFigure 3~4
  • EP4436184A1 patent drawingFigure 5

AI summary

The present application relates to a method of decoding a video picture by means of motion compensated temporal bi-prediction of an inter coded block using two reference video pictures in two separate reference video picture lists, and two affine motion fields defined by at least two control point motion vectors, denoted CPMVs, associated to each reference picture, refined CPMVs being obtained as output of the following first step (171) or, of the second step (172) or of the second step (172) that follows the first step (171), wherein the first or second step is/are bypassed according to binary values.