Neural In-Loop Filtering for Machine Video Coding Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding technologies struggle to effectively balance compression efficiency and subjective quality for both human and machine tasks, particularly in video coding for machines, where machine vision tasks require specialized encoding and decoding processes.
Innovation Solution
Implementing a neural network in-loop filter (NNLF) that can be dynamically enabled or disabled based on output from VCM processing modules, with parameters signaled in the bitstream or derived during decoding, to enhance reconstructed video for machine tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If neural network in-loop filter (NNLF) is applied to enhance video quality for machine tasks, then video coding efficiency and quality for machine tasks improve, but device complexity and computational overhead increase
Solution Approach 1:
The patent implements dynamic control of NNLF application through signaling in the bitstream, allowing the filtering process to be adaptively enabled or disabled based on video content characteristics and machine task requirements. This dynamic approach optimizes the balance between quality enhancement and computational complexity by applying the neural network filter only when beneficial.
Solution Approach 2:
The patent applies NNLF selectively to specific portions of video content rather than uniformly to all frames. By using region-of-interest (RoI) based approaches and selective filtering flags, the system concentrates computational resources on areas that most benefit from neural network filtering, improving machine task performance while reducing overall complexity.
2Measurement precision
If NNLF is applied to all video content, then overall video quality for machine tasks improves, but computational resources and processing time are wasted on regions that do not benefit from filtering
Solution Approach 1:
The patent implements selective NNLF application at multiple granularities including frame-level flags, slice-level control, and region-of-interest based filtering. This allows the system to apply computational resources only where needed, improving processing efficiency while maintaining quality enhancement where beneficial for machine tasks.
Solution Approach 2:
The patent employs a selective filtering approach where NNLF is applied partially to only those video portions that benefit from it, rather than applying full filtering to all content. This partial action approach optimizes processing efficiency by avoiding unnecessary computation on already-suitable video regions.
3Adaptability or versatility
If NNLF parameters are signaled in the bitstream, then filtering accuracy and adaptability improve, but bitstream size and decoding complexity increase
Solution Approach 1:
The patent segments the NNLF parameter signaling into multiple levels: global parameters signaled at sequence or picture level, and local parameters signaled only where needed at slice or block level. This segmentation reduces overall bitstream overhead while maintaining the adaptability to apply different filtering parameters to different video regions.
Solution Approach 2:
The patent signals NNLF parameters partially, only for those regions where filtering is actually applied. By using conditional signaling based on flags and region indicators, the system avoids transmitting unnecessary parameter data, reducing bitstream overhead while preserving adaptability where needed.
Data Source
AI summary
This disclosure relates generally to video coding, and more particularly to in-loop filtering for video coding for machine tasks based on neural networks. For example, utilization of one or more neural network in-loop filter (NNLF) may be determined (either enabled or disabled) at various coding level. Such determination may be based on output of one or more other VCM-codec-related processed. The usages of the NNLF may be signaled in the bitstream or may be implicitly derived from the output of the one or more other VCM-coded-related processes during decoding process in a decoder or during in-loop decoding process of an encoder.


