Video Processing With Neural Network Filters for Coding Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding technologies, such as MPEG-2, MPEG-4, ITU-T.263, ITU-T.264/MPEG-4 AVC, ITU-T.265 HEVC, and VVC, require improvements in coding efficiency and effectiveness.
Innovation Solution
The application of a neural network filter for video processing, including visual quality improvement and machine vision tasks, enhances coding efficiency and effectiveness by converting video units into bitstreams.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional video coding technologies (MPEG-2, MPEG-4, H.263, AVC, HEVC, VVC) are used, then video encoding/decoding can be performed, but coding efficiency and coding effectiveness need further improvement
Solution Approach 1:
A neural network filter is introduced as an intermediary component between the video unit and the bitstream. The filter processes the video unit through multiple convolution layers and activation functions to generate enhanced color components, which are then used in the bitstream conversion process. This intermediary neural network filter improves both coding efficiency and effectiveness by optimizing the transformation between video units and bitstreams.
Solution Approach 2:
The patent applies parameter changes by transforming the video unit through neural network convolution operations that modify color component parameters. The neural network filter processes the video data through multiple layers with different kernel sizes and activation functions, changing the parameter representation of the video data to achieve better coding efficiency and effectiveness.
2Productivity
If a neural network filter is applied to the current video unit, then coding efficiency and coding effectiveness are improved, but the complexity of the video processing system increases
Solution Approach 1:
The neural network filter is segmented into multiple convolution layers, each performing a specific function. The filter includes an first convolution layer with first kernel size, a second convolution layer with second kernel size, and so on. This segmentation allows the complex neural network to be broken down into manageable modules that can be processed sequentially, reducing the overall system complexity while maintaining coding efficiency.
Solution Approach 2:
The patent applies partial action by selectively applying the neural network filter to specific color components (e.g., only to chroma components or only to luma components) rather than processing all components equally. This partial application reduces the computational burden and system complexity while still achieving significant coding efficiency improvements where most needed.
Data Source
AI summary
Embodiments of the present disclosure provide a solution for video processing. A method for video processing is proposed. In the method, a conversion between a current video unit of a video and a bitstream of the video is performed. A neural network filter is applied to the current video unit. The neural network filter has a target purpose comprising one of: visual quality improvement, an implementation of a machine vision task, or an image or video processing.


