Video Block Encoding With Neural Filter Combination Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video encoding and decoding technologies using neural networks face challenges in optimizing the performance of neural network filters, leading to increased bitrate overhead due to the need to transmit combination information, particularly for smaller block sizes.
Innovation Solution
A method and apparatus that combines the output of neural network filters with other data sources, such as conventional filters, using linear combinations, and encodes combination parameters into the bitstream to improve filter performance on subsequent blocks, reducing bitrate overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If neural network filters are used for video encoding and decoding, then filtering performance is improved, but bitrate overhead increases due to transmission of combination information
Solution Approach 1:
The patent extracts only the essential combination information (combination type and optional parameters) from the full filter configuration and transmits only this extracted subset through the bitstream. This allows the decoder to reconstruct the filter behavior without receiving complete filter data, thereby reducing bitrate overhead while maintaining filtering performance.
Solution Approach 2:
The patent applies different levels of combination information transmission based on local needs - combination information is transmitted only when neural network filter combination is actually used for a particular block, and optional parameters are transmitted only when needed. This localized approach avoids unnecessary bitrate overhead for blocks that don't require complex filtering.
2Manufacturing precision
If combination information is transmitted in the bitstream, then filter performance on subsequent blocks is improved, but device complexity increases
Solution Approach 1:
The patent segments the filter configuration into two parts: combination information (transmitted in bitstream) and implementation details (processed locally). This segmentation allows the encoder to transmit only essential guidance information while the actual filter combination operations are performed efficiently at the decoder side, reducing overall system complexity.
Solution Approach 2:
The patent performs preliminary classification of blocks to determine which ones require neural network filter combination, and only for those blocks does it generate and transmit combination information. This preliminary action avoids unnecessary processing and transmission overhead for blocks that can be handled by standard filtering methods.
3Ease of operation
If neural network filters are applied to each block independently, then processing simplicity is maintained, but filter performance on subsequent blocks deteriorates
Solution Approach 1:
The patent implements feedback by transmitting combination information from the encoder to the decoder, allowing the decoder to replicate the encoder's filter combination decisions. This feedback loop ensures that the decoder can apply the same filtering operations as the encoder, maintaining performance consistency across blocks while keeping the processing framework relatively simple.
Solution Approach 2:
The patent merges multiple neural network filters and their outputs in combination processes guided by transmitted combination information. This merging allows the system to leverage multiple filters for improved performance on subsequent blocks while maintaining a unified processing framework that doesn't require completely redesigning the decoding architecture.
Data Source
AI summary
The embodiments relate to method for encoding and decoding, wherein the method for encoding comprises receiving an input block of a video frame for encoding; applying at least a learning-based model (702) for said input block as a processing step for encoding the block; combining (703) an output of a learning-based model with one or more data sources (712, 713) by a combination process; encoding block to a bitstream (40); using a result of the combination process as additional input for the learning-based model for encoding a subsequent block; and encoding to a bitstream combination information (720) used in the combination process, said combination information comprising at least one or more combination parameters. The embodiments also relate to technical equipment for implementing the methods.


