Variable Coefficient Deep Learning for Video Inter Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video encoding techniques face challenges in achieving high encoding efficiency due to increasing image sizes, resolutions, and frame rates, leading to higher data amounts that require more hardware resources, necessitating improved compression techniques beyond existing methods like H.264/AVC, HEVC, and VVC.

Innovation Solution

An inter prediction method utilizing a variable coefficient deep learning model that adapts to video characteristics, generating a virtual reference frame through an interpolation model, and transmitting parameters from the encoding apparatus to the decoding apparatus for enhanced encoding efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If video data is stored or transmitted without compression processing, then the original video quality is preserved, but a large amount of hardware resources including memory are required

Engineering Contradiction:
Improvevideo qualityVSAvoidhardware resources
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts and transmits only the essential motion information parameters (motion vectors, reference frame indices) from the full video data, rather than transmitting or storing complete uncompressed video frames. This selective extraction achieves compression while preserving the critical information needed for video reconstruction and quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates simplified copies of video information through motion estimation and compensation, where reference frames are copied and shifted according to motion vectors to reconstruct current frames. This copying approach requires significantly less memory and processing resources compared to storing complete uncompressed video data.

Inventive Principle:
Principle #26Copying

2Reliability

If image size, resolution, and frame rate are increased, then video quality and detail are improved, but the amount of data to be encoded increases

Engineering Contradiction:
Improvevideo qualityVSAvoiddata amount
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent replaces traditional mechanical compression approaches with deep learning-based motion estimation and compensation. Neural networks learn optimal motion patterns and predict frame content more accurately, achieving better compression ratios for high-resolution, high-frame-rate video while maintaining quality by intelligently analyzing and representing motion dynamics.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If existing compression techniques (H.264/AVC, HEVC, VVC) are used, then video data is compressed, but encoding efficiency is limited by increasing data amounts

Engineering Contradiction:
Improveencoding efficiencyVSAvoiddata amount
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent changes the fundamental parameters of motion representation by using deep learning models to learn and encode motion patterns in a more efficient parameter space. Instead of traditional block-based motion estimation, the system uses neural networks to predict motion fields and frame content, achieving superior encoding efficiency for modern video formats with higher resolutions and frame rates.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12192445B2Inter prediction method based on variable coefficient deep learning
Publication Date: 2025.01.07 HYUNDAI MOTOR CO LTD
  • US12192445B2 patent drawing
  • US12192445B2 patent drawing
  • US12192445B2 patent drawing

AI summary

An inter prediction method allows a variable coefficient deep learning model to adaptively learn characteristics of a video; transmits a variable coefficient deep learning model parameter generated from the learning from an image encoding device to an image decoding device; and refers to a virtual reference frame generated by the variable coefficient deep learning model.