Selective Learning-Based Model Application in Video Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video encoding and decoding technologies using neural networks face challenges in optimizing the operation of neural network filters, particularly in deciding when to apply learning-based models for encoding and decoding video frames, which affects the quality and efficiency of the compression and reconstruction processes.
Innovation Solution
The proposed solution involves a method and apparatus that selectively apply a learning-based model for video frame encoding and decoding, where the model is fine-tuned based on the activation status of previously-decoded blocks, and the decision to use the model is encoded into a bitstream, allowing for adaptive processing and improved visual quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If a learning-based model is applied for encoding and decoding video blocks, then the visual quality of reconstructed content is improved, but the computational overhead and processing time increase
Solution Approach 1:
The patent applies the learning-based model selectively to only certain video blocks that benefit most from it, rather than applying it to all blocks. The system determines which blocks require the learning-based model based on specific criteria, applying the model partially to achieve quality improvement while avoiding unnecessary computational overhead on blocks where the model would not provide significant benefit.
Solution Approach 2:
The patent implements different processing qualities for different regions of the video content. By analyzing local block characteristics and applying the learning-based model only where needed, the system creates variable quality processing - high quality where the model is applied and standard quality where it is not - thereby optimizing the overall quality-computation trade-off.
2Productivity
If the learning-based model is selectively fine-tuned based on previously-decoded blocks, then the encoding efficiency is improved, but the decision-making complexity increases
Solution Approach 1:
The patent implements a feedback mechanism where the system monitors the performance and activation status of the learning-based model on previously decoded blocks. This feedback information is used to dynamically adjust and fine-tune the model's application strategy, allowing the system to learn from past decisions and improve encoding efficiency over time while adapting to different content characteristics.
Solution Approach 2:
The patent performs preliminary analysis of video blocks to determine in advance whether the learning-based model should be applied. By pre-assessing block characteristics and predicting which blocks would benefit from the model before actual encoding occurs, the system streamlines the decision-making process and improves encoding efficiency by avoiding unnecessary model applications.
3Adaptability or versatility
If the decision on usage of the learning-based model is encoded into a bitstream, then the adaptability of the decoding process is improved, but the bitstream size increases
Solution Approach 1:
The patent extracts only the essential decision information needed for the learning-based model usage and encodes only this critical data into the bitstream, rather than encoding complete model parameters or extensive metadata. This selective extraction approach maintains the adaptability needed for proper decoding while minimizing the additional bitstream overhead.
Solution Approach 2:
The patent applies partial encoding of the model usage decisions, encoding information only for blocks where the learning-based model is applied or where such information is necessary for reconstruction. This partial approach provides sufficient adaptability for the decoding process while avoiding the overhead of encoding complete information for all blocks.
Data Source
AI summary
The embodiments relate to a method for encoding and a decoding, and apparatuses for the same. The method for encoding comprises receiving a block of a video frame for encoding (1510); making a decision on whether or not a learning-based model is to be applied as a processing step for encoding the block (1520); applying the learning-based model for said input block according to the decision, where the learning-based model has been selectively fine-tuned according to information relating to activation of the learning-based model of previously-decoded blocks (1530); encoding a signal corresponding to the decision on usage of the learning-based model into a bitstream (1540); and encoding the block into a bitstream with an information whether the block is to be used for finetuning (1550).


