Neural Network Video Coding Using Hierarchical Virtual Reference Frames

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional video coding standards face challenges in efficiently compressing video data, particularly due to high bitrate requirements and the inability to effectively handle complex motion in dynamic scenes.

Innovation Solution

The method involves generating a virtual reference frame for the current picture based on its hierarchical level and a nearest decoded picture, which is then used for decoding the video data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional video coding standards (H.264, HEVC, VVC) are used, then compression efficiency is improved through handcrafted prediction and transform tools, but bitrate requirements remain high and complex motion in dynamic scenes cannot be effectively handled

Engineering Contradiction:
Improvecompression efficiencyVSAvoidbitrate requirement
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent replaces traditional handcrafted mechanical prediction systems with a neural network-based system. The neural network learns optimal prediction patterns from training data and automatically adapts to different video content, replacing the fixed block-based hybrid prediction/transform framework with a data-driven approach that can handle complex motion more effectively

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the coding system by introducing hierarchical levels (Level 0, Level 1, Level 2) that process video data at different resolutions and detail levels. This hierarchical parameter structure allows the system to achieve better compression efficiency while reducing bitrate requirements by processing information progressively

Inventive Principle:
Principle #35Parameter changes

2Productivity

If traditional block-based hybrid prediction/transform framework is used, then coding tools can be handcrafted to optimize overall efficiency, but the system lacks adaptability to complex motion patterns in dynamic scenes

Engineering Contradiction:
Improvecoding efficiencyVSAvoidhandling of complex motion
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent introduces dynamic adaptability through hierarchical levels that can independently process different aspects of video data. The system dynamically adjusts which levels are activated based on the complexity of motion patterns detected in the video content, allowing it to adapt to varying scene requirements while maintaining overall coding efficiency

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The neural network-based system replaces the static handcrafted block-based prediction system with a learning-based approach that automatically adapts to complex motion patterns. The neural network processes video data through multiple hierarchical levels, enabling it to handle dynamic scenes more effectively while maintaining coding efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If uncompressed video is used, then no compression is applied, but bitrate requirements become excessively high (e.g., 1.5 Gbit/s for 1080p60) and storage space requirements become unmanageable

Engineering Contradiction:
Improvevideo qualityVSAvoidbitrate and storage requirement
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the video data processing into hierarchical levels (Level 0 for base layer, Level 1 for intermediate details, Level 2 for fine details). This segmentation allows the system to compress video data more effectively by processing different detail levels separately, achieving better compression ratios while maintaining video quality and significantly reducing bitrate and storage requirements

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12244834B2Utilizing hierarchical structure for neural network based tools in video coding
Publication Date: 2025.03.04 TENCENT AMERICA LLC
  • US12244834B2 patent drawing
  • US12244834B2 patent drawing
  • US12244834B2 patent drawing

AI summary

A method, computer program, and computer system is provided for video encoding and decoding. Video data including a current picture is received. A virtual reference frame is generated for the current picture based on hierarchical level associated with the current picture and a nearest decoded picture. The video data is decoded based on the generated reference frame.