Learned Image Codec for Machine Vision Region of Interest Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current frame rate coding methods fail to effectively determine regions of interest in video encoding, leading to inefficient compression and decoding processes, especially when dealing with machine consumption where quality metrics differ from human perception.
Innovation Solution
The proposed solution involves using a neural network-based learned image codec to generate high-quality reconstructed pictures and a conventional video codec to generate low-quality pictures, with subsequent quality measurements and bit allocation calculations to categorize blocks as foreground or background regions, enabling efficient encoding and decoding by signaling ROI information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of substance
If a conventional video codec is used for encoding, then compression efficiency is improved, but reconstruction quality deteriorates
Solution Approach 1:
The patent changes the fundamental parameter of the encoding approach by transitioning from conventional video codecs to neural network-based learned image codecs. This parameter change enables simultaneous achievement of high compression efficiency and high reconstruction quality, as the neural network learns optimal encoding parameters and representations that conventional methods cannot achieve.
2Manufacturing precision
If a neural network based learned image codec is used with high quality settings, then reconstruction quality is improved, but compression efficiency deteriorates
Solution Approach 1:
The patent applies local quality by generating multiple reconstructed pictures with different quality levels (first reconstructed picture with high quality, second reconstructed picture with low quality) and using quality measurements to identify regions of interest. This allows the system to allocate bits selectively - high quality encoding for important regions and lower quality for less important regions - thereby improving overall compression efficiency while maintaining necessary reconstruction quality.
Solution Approach 2:
The patent dynamically adjusts encoding parameters based on content analysis. By computing quality measurements between different reconstructed pictures and identifying regions where quality degradation occurs, the system changes encoding parameters locally to maintain quality only where necessary, improving overall compression efficiency.
3Manufacturing precision
If uniform quality encoding is applied to all regions, then reconstruction quality is maintained, but bit allocation efficiency deteriorates
Solution Approach 1:
The patent implements local quality by categorizing blocks into foreground and background regions based on quality measurements. Different encoding strategies are applied to different regions: foreground regions receive higher bit allocation to maintain quality, while background regions receive lower bit allocation. This resolves the contradiction by maintaining reconstruction quality only where necessary, thereby improving bit allocation efficiency.
Solution Approach 2:
The patent segments the image into different regions (foreground and background) based on quality measurements and applies different encoding parameters to each segment. This segmentation allows efficient bit allocation by concentrating bits on important regions while using fewer bits for less important regions, resolving the contradiction between quality maintenance and bit allocation efficiency.
4Measurement precision
If multiple reconstructed pictures are generated and compared, then region of interest detection accuracy is improved, but computational complexity deteriorates
Solution Approach 1:
The patent changes the approach from using complex proxy networks with feature maps to a more direct quality measurement method. By computing quality measurements directly between different reconstructed pictures and comparing them, the system achieves accurate region of interest detection with reduced computational complexity, avoiding the overhead of complex neural network feature extraction.
Data Source
AI summary
Various embodiments describe an apparatus, a method, and a computer program product. An example apparatus includes at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: encoding an input picture by using a first encoder or first encoding parameters; encoding the input picture by using a second encoder or second encoding parameters; generating a first reconstructed picture based on the encoding of the input picture by using the first encoder or the first encoding parameters; and generating a second reconstructed picture based on the encoding of the input picture by using the second encoder or the second encoding parameters.


