Video Decoding Inter Prediction Modes for Motion Vector Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding technologies face inefficiencies in motion vector prediction, particularly in lossy compression scenarios, leading to suboptimal compression ratios and increased bandwidth requirements due to rounding errors and redundancy in motion vector data.

Innovation Solution

The introduction of advanced inter prediction modes, including MMVD, SbTMVP, CIIP, triangle, affine merge, and AMVP modes, which enhance motion vector prediction by utilizing motion vector differences and spatial-temporal correlations to improve compression efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If motion vector prediction is used to reduce data redundancy, then compression ratio is improved, but rounding errors and redundancy in motion vector data lead to suboptimal compression and increased bandwidth requirements

Engineering Contradiction:
Improvebit usage for motion vector dataVSAvoidcompression efficiency
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments the motion vector prediction process into multiple inter prediction modes (MMVD, SbTMVP, CIIP, triangle, affine merge, AMVP), each handling different motion characteristics. This segmentation allows more precise motion compensation and reduces rounding errors, thereby improving compression efficiency while maintaining reduced bit usage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces multiple inter prediction modes with different parameters and methodologies for motion vector prediction. By changing the prediction parameters and methods based on the specific video content and motion characteristics, the system achieves better compression efficiency without increasing bandwidth requirements, as each mode is optimized for specific scenarios.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If multiple advanced inter prediction modes are introduced to enhance motion vector prediction, then compression efficiency is improved, but device complexity increases

Engineering Contradiction:
Improvebit usage for motion vector dataVSAvoiddecoding complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements a dynamic mode selection mechanism where the decoder chooses among multiple inter prediction modes (MMVD, SbTMVP, CIIP, triangle, affine merge, AMVP) based on the specific video content and motion characteristics. This dynamic approach allows the system to use simpler modes when sufficient and more complex modes only when necessary, balancing compression efficiency with decoding complexity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies the principle of partial action by implementing a hierarchy of inter prediction modes with varying complexity. Not all modes are applied to every block; instead, the system uses the minimum necessary complexity for each block type, applying advanced modes only where they provide significant compression benefits, thus managing device complexity while improving overall compression efficiency.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3906688B1Mathod and apparatus for video decoding.
Publication Date: 2025.12.03 TENCENT AMERICA LLC
  • EP3906688B1 patent drawingFigure 1
  • EP3906688B1 patent drawingFigure 2
  • EP3906688B1 patent drawingFigure 3

AI summary

Aspects of the disclosure provide methods and apparatuses for video encoding/decoding. In some examples, an apparatus for video decoding includes receiving circuitry and processing circuitry. The processing circuitry decodes prediction information of a current block from a coded video bitstream. The prediction information is indicative of a subset of inter prediction modes associated with a merge flag being false. Then, the processing circuitry decodes at least an additional flag that is used for selecting a specific inter prediction mode from the subset of inter prediction modes. Further, the processing circuitry reconstructs samples of the current block according to the specific inter prediction mode.