Virtual Reference Frame Generation for Inter-Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image encoding/decoding technologies face challenges in efficiently handling high-resolution and high-definition images, particularly in accurately predicting pixel values for inter prediction, especially with the increasing demand for UHD resolutions.
Innovation Solution
The development of an encoding and decoding method that generates a virtual reference frame using a deep-learning network architecture, specifically a Generative Adversarial Network (GAN) or Adaptive Convolution Network (ACN), for inter prediction, enabling video interpolation and extrapolation to enhance prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional inter-prediction technology is used for high-resolution images, then encoding/decoding can be performed, but prediction accuracy is insufficient for UHD resolutions
Solution Approach 1:
A virtual reference frame is generated as an intermediary between existing reference frames and the target block to be predicted. This virtual reference frame, created through deep learning-based video interpolation and extrapolation, serves as a mediator that provides more accurate prediction data for high-resolution images than traditional reference frames alone could provide.
Solution Approach 2:
The patent creates a virtual copy (virtual reference frame) of actual reference frames using deep learning networks. This copied and enhanced reference data allows for more accurate inter-prediction in high-resolution images without requiring additional actual reference frames, effectively adapting traditional prediction methods to UHD resolutions.
2Measurement precision
If deep learning networks are used to generate virtual reference frames, then prediction accuracy is improved, but computational complexity increases
Solution Approach 1:
The virtual reference frame is generated in advance (preliminarily) before the actual inter-prediction process. By pre-generating the virtual reference frame using deep learning networks, the complex computational work is performed beforehand, allowing the subsequent prediction process to use this pre-processed data more efficiently.
Solution Approach 2:
The deep learning network is trained to automatically learn and perform the video interpolation and extrapolation tasks itself. Once trained, the network serves itself by generating virtual reference frames without requiring manual intervention or complex external processing, reducing overall system complexity despite the inherent complexity of the network architecture.
3Productivity
If virtual reference frames are generated using video interpolation and extrapolation, then inter-prediction efficiency is improved, but processing time increases
Solution Approach 1:
Video interpolation and extrapolation are performed preliminarily to generate the virtual reference frame before the actual inter-prediction encoding/decoding process. This pre-processing approach allows the main prediction process to run more efficiently by using pre-computed virtual reference data, reducing real-time processing requirements.
Data Source
AI summary
An inter-prediction method and apparatus uses a reference frame generated based on deep learning. In the inter-prediction method and apparatus, a reference frame is selected, and a virtual reference frame is generated based on the selected reference frame. A reference picture list is configured to include the generated virtual reference frame, and inter prediction for a target block is performed based on the virtual reference frame. The virtual reference frame may be generated based on a deep-learning network architecture, and may be generated based on video interpolation and/or video extrapolation that use the selected reference frame.


