Video Coding Syntax for Neural Post-Filter Frame-Rate Upsampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding standards lack efficient methods for incorporating neural network post-filter techniques to enhance video quality, particularly in the context of future video coding standards like VVC, which have not fully integrated CNN-based post-filtering for artifact removal.
Innovation Solution
The proposed techniques involve signaling neural network post-filter parameter information through syntax elements, specifying the number of input pictures and concatenation manner for a neural network post-filter picture interpolation process, enabling effective integration of CNN-based post-filtering in video coding systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If video coding standards incorporate neural network post-filter techniques, then video quality is improved and artifacts are reduced, but device complexity and implementation difficulty increase
Solution Approach 1:
The patent segments the neural network post-filter implementation into distinct components: syntax element definition, parameter signaling mechanisms, and separate processing stages (interpolation, filtering). This modular approach allows incremental adoption and reduces overall system complexity while maintaining video quality improvements.
Solution Approach 2:
The patent introduces dynamic parameters that can be adjusted based on content characteristics, such as the number of input pictures and concatenation manner. This allows the system to adapt the complexity of neural network processing to match the actual video content requirements, reducing unnecessary computational overhead while maintaining quality where needed.
2Adaptability or versatility
If neural network post-filter parameters are signaled through syntax elements, then integration with existing video coding standards is improved, but bitstream complexity and processing overhead increase
Solution Approach 1:
The patent designs the syntax elements and parameter structures to be compatible with multiple video coding standards (H.264, H.265, VVC, JEM). This universal approach allows the same signaling mechanism to serve different coding standards, reducing the need for separate implementation paths and lowering overall system complexity despite the added functionality.
Solution Approach 2:
The patent introduces intermediate parameter structures that act as mediators between the neural network post-filter requirements and the existing video coding standard frameworks. These intermediate layers translate complex neural network parameters into standardized syntax elements, reducing bitstream complexity while maintaining integration capability.
3Measurement precision
If multiple input pictures are used for neural network post-filter interpolation, then filtering accuracy is improved, but computational load and processing time increase
Solution Approach 1:
The patent makes the number of input pictures a dynamic parameter that can be adjusted based on content characteristics and processing requirements. This allows the system to use more pictures for high-accuracy regions while using fewer pictures for simpler regions, optimizing the trade-off between filtering accuracy and processing time on a per-region basis.
Solution Approach 2:
The patent applies different numbers of input pictures to different spatial regions or picture types based on their specific quality requirements. Complex or important regions receive higher filtering accuracy with more input pictures, while simpler regions use fewer pictures, reducing overall processing time while maintaining necessary quality levels.
Data Source
AI summary
A device may be configured to perform frame rate upsampling based on information included in a neural network post-filter characteristics message. In one example, the neural network post-filter characteristics message includes a syntax element specifying a number of input pictures to be used as input for a neural network post-filter picture interpolation process and a syntax element having a value specifying a manner in which input pictures are concatenated before being input into the neural network post-filter picture interpolation process.


