Bi-Prediction Motion Vector Coding Across Complex Motion Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video compression standards like VVC draft 3 limit the application of MMVD and SMVD motion vector coding tools to translational motion models, lacking flexibility and efficiency in handling more complex motion models and temporal prediction methods.
Innovation Solution
Extend the usage of MMVD and SMVD motion vector coding tools to support all motion models and temporal prediction methods in the VVC draft 3, including affine, ATMVP, planar, regressive, triangle-partition-based, GBI, LIC, and Multi-hypothesis prediction methods, especially in bi-prediction scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If MMVD and SMVD motion vector coding tools are limited to translational motion models only, then the standard maintains simplicity and ease of implementation, but the adaptability and compression efficiency for complex motion models deteriorate
Solution Approach 1:
The patent extends the MMVD and SMVD coding tools to support multiple motion models (translational, affine, ATMVP, planar, regressive, triangle-partition-based, GBI, LIC, and Multi-hypothesis prediction) beyond their original translational-only limitation. This universal application allows a single coding tool to handle diverse motion complexities, improving adaptability without requiring separate dedicated tools for each motion model.
2Productivity
If MMVD and SMVD tools are extended to all motion models, then compression efficiency and coding flexibility improve, but the device complexity and implementation difficulty increase
Solution Approach 1:
The patent introduces dynamic selection mechanisms that allow the encoder to adaptively choose which motion model and corresponding MMVD/SMVD coding approach to use for each video block based on content characteristics. This dynamic adaptation enables the system to achieve high compression efficiency for complex motions when needed while falling back to simpler modes for straightforward cases, balancing performance with implementation complexity.
Solution Approach 2:
The patent modifies the coding parameters and syntax structures to accommodate multiple motion models. By changing the parameter sets and data structures to be model-aware, the system can efficiently encode diverse motion types using extended MMVD/SMVD tools without requiring fundamentally different coding architectures, thus managing complexity through parameterization rather than structural multiplication.
3Measurement precision
If the syntax structure includes information for both first motion mode and second motion mode, then the precision of motion representation improves, but the bitstream size and encoding complexity increase
Solution Approach 1:
The patent employs conditional syntax inclusion where the second motion mode information is only encoded when actually used. The syntax structure includes flags and conditional elements that allow the encoder to omit redundant motion mode information, achieving partial encoding only when necessary for precision. This prevents unnecessary bitstream bloat while maintaining the capability to represent complex motions when needed.
Solution Approach 2:
The patent segments the motion information encoding into distinct layers or stages, where the first motion mode provides baseline information and the second motion mode provides optional refinements. This segmentation allows the decoder to process essential motion data first and only process additional refinement data when present, improving parsing efficiency and allowing selective precision based on content requirements.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The general aspects extend motion modes, such as merge with motion vector difference, and symmetrical motion vector difference, to motion models beyond a simple translational model, for example in combination with merge and alternative temporal motion vector prediction modes. Embodiments extend the use of MMVD and SMVD motion vector coding tools to all the motion model derivation methods and temporal prediction methods that are supported in proposed video standards, so as to increase the overall compression performance. Particular embodiments describe combining MMVD or SMVD with the affine motion model, the ATMVP motion model, the planar motion model, the regressive motion field, the triangle-partition-based motion model, the GBI temporal prediction method, the LIC temporal prediction method and the Multi-hypothesis prediction method.