Neural Network Video Coding Using Hierarchical Virtual Reference Frames
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional video coding standards face challenges in efficiently compressing video data, particularly due to high bitrate requirements and the inability to effectively handle complex motion in dynamic scenes.
Innovation Solution
The method involves generating a virtual reference frame for the current picture based on its hierarchical level and a nearest decoded picture, which is then used for decoding the video data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional video coding standards (H.264, HEVC, VVC) are used, then compression efficiency is improved through handcrafted prediction and transform tools, but bitrate requirements remain high and complex motion in dynamic scenes cannot be effectively handled
Solution Approach 1:
The patent replaces traditional handcrafted mechanical prediction systems with a neural network-based system. The neural network learns optimal prediction patterns from training data and automatically adapts to different video content, replacing the fixed block-based hybrid prediction/transform framework with a data-driven approach that can handle complex motion more effectively
Solution Approach 2:
The patent changes the fundamental parameters of the coding system by introducing hierarchical levels (Level 0, Level 1, Level 2) that process video data at different resolutions and detail levels. This hierarchical parameter structure allows the system to achieve better compression efficiency while reducing bitrate requirements by processing information progressively
2Productivity
If traditional block-based hybrid prediction/transform framework is used, then coding tools can be handcrafted to optimize overall efficiency, but the system lacks adaptability to complex motion patterns in dynamic scenes
Solution Approach 1:
The patent introduces dynamic adaptability through hierarchical levels that can independently process different aspects of video data. The system dynamically adjusts which levels are activated based on the complexity of motion patterns detected in the video content, allowing it to adapt to varying scene requirements while maintaining overall coding efficiency
Solution Approach 2:
The neural network-based system replaces the static handcrafted block-based prediction system with a learning-based approach that automatically adapts to complex motion patterns. The neural network processes video data through multiple hierarchical levels, enabling it to handle dynamic scenes more effectively while maintaining coding efficiency
3Reliability
If uncompressed video is used, then no compression is applied, but bitrate requirements become excessively high (e.g., 1.5 Gbit/s for 1080p60) and storage space requirements become unmanageable
Solution Approach 1:
The patent segments the video data processing into hierarchical levels (Level 0 for base layer, Level 1 for intermediate details, Level 2 for fine details). This segmentation allows the system to compress video data more effectively by processing different detail levels separately, achieving better compression ratios while maintaining video quality and significantly reducing bitrate and storage requirements
Data Source
AI summary
A method, computer program, and computer system is provided for video encoding and decoding. Video data including a current picture is received. A virtual reference frame is generated for the current picture based on hierarchical level associated with the current picture and a nearest decoded picture. The video data is decoded based on the generated reference frame.


