Inter Prediction Mode for Video Decoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video processing technologies face challenges in efficiently encoding and decoding next-generation video contents with high spatial resolution, high frame rate, and high dimensionality of scene representation, leading to increased memory storage, memory access rate, and processing power requirements.

Innovation Solution

A method is proposed for deriving a temporal motion vector from one reference picture, selecting a reference picture for this derivation using a signaled syntax, and applying Advanced Temporal Motion Vector Prediction (ATMVP) based on spatial candidates, which reduces memory bandwidth and addresses the additional line buffer problem.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional inter prediction modes are used for next-generation video contents, then processing capability is maintained, but memory bandwidth and processing power requirements increase drastically

Engineering Contradiction:
Improvevideo processing efficiencyVSAvoidmemory bandwidth requirement
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent extracts and utilizes only the necessary motion information from reference pictures by deriving temporal motion vectors from a single reference picture rather than processing multiple reference pictures. This extraction approach reduces memory bandwidth requirements while maintaining prediction accuracy for high-resolution video content.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of reference picture selection by using a signaled syntax to select one reference picture for temporal motion vector derivation. This parameter change optimizes memory access patterns and reduces the quantity of data that needs to be processed, thereby improving video processing efficiency without drastically increasing memory bandwidth requirements.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If multiple reference pictures are processed for motion vector prediction, then prediction accuracy is improved, but additional line buffer problems occur

Engineering Contradiction:
Improvemotion vector prediction accuracyVSAvoidline buffer requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts motion information from a single selected reference picture rather than processing multiple reference pictures simultaneously. This extraction strategy maintains prediction accuracy by focusing computational resources on one high-quality reference while avoiding the line buffer complexity associated with managing multiple reference pictures.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of processing multiple reference pictures and then selecting the best one, the patent inverts the approach by using signaled syntax to pre-select one reference picture for temporal motion vector derivation. This inversion simplifies the buffering requirements while maintaining prediction accuracy.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS12284381B2Image processing method based on inter prediction mode, and device therefor
Publication Date: 2025.04.22 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • US12284381B2 patent drawing
  • US12284381B2 patent drawing
  • US12284381B2 patent drawing

AI summary

In the present disclosure, a method of decoding a video signal and a device therefor are disclosed. Specifically, a method of decoding an image based on an inter prediction mode includes deriving a motion vector of an available spatial neighboring block around a current block; deriving a collocated block of the current block based on the motion vector of the spatial neighboring block; deriving a motion vector in a sub-block unit in the current block based on a motion vector of the collocated block; and generating a prediction block of the current block using the motion vector derived in the sub-block unit, wherein the collocated block may be specified by the motion vector of the spatial neighboring block in one pre-defined reference picture.