Learned Image Codec for Machine Vision Region of Interest Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current frame rate coding methods fail to effectively determine regions of interest in video encoding, leading to inefficient compression and decoding processes, especially when dealing with machine consumption where quality metrics differ from human perception.

Innovation Solution

The proposed solution involves using a neural network-based learned image codec to generate high-quality reconstructed pictures and a conventional video codec to generate low-quality pictures, with subsequent quality measurements and bit allocation calculations to categorize blocks as foreground or background regions, enabling efficient encoding and decoding by signaling ROI information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of substance

If a conventional video codec is used for encoding, then compression efficiency is improved, but reconstruction quality deteriorates

Engineering Contradiction:
Improvecompression efficiencyVSAvoidreconstruction quality
Core Design Contradiction:
Loss of substanceVSManufacturing precision

Solution Approach 1:

The patent changes the fundamental parameter of the encoding approach by transitioning from conventional video codecs to neural network-based learned image codecs. This parameter change enables simultaneous achievement of high compression efficiency and high reconstruction quality, as the neural network learns optimal encoding parameters and representations that conventional methods cannot achieve.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If a neural network based learned image codec is used with high quality settings, then reconstruction quality is improved, but compression efficiency deteriorates

Engineering Contradiction:
Improvereconstruction qualityVSAvoidcompression efficiency
Core Design Contradiction:
Manufacturing precisionVSLoss of substance

Solution Approach 1:

The patent applies local quality by generating multiple reconstructed pictures with different quality levels (first reconstructed picture with high quality, second reconstructed picture with low quality) and using quality measurements to identify regions of interest. This allows the system to allocate bits selectively - high quality encoding for important regions and lower quality for less important regions - thereby improving overall compression efficiency while maintaining necessary reconstruction quality.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically adjusts encoding parameters based on content analysis. By computing quality measurements between different reconstructed pictures and identifying regions where quality degradation occurs, the system changes encoding parameters locally to maintain quality only where necessary, improving overall compression efficiency.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If uniform quality encoding is applied to all regions, then reconstruction quality is maintained, but bit allocation efficiency deteriorates

Engineering Contradiction:
Improvereconstruction qualityVSAvoidbit allocation efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent implements local quality by categorizing blocks into foreground and background regions based on quality measurements. Different encoding strategies are applied to different regions: foreground regions receive higher bit allocation to maintain quality, while background regions receive lower bit allocation. This resolves the contradiction by maintaining reconstruction quality only where necessary, thereby improving bit allocation efficiency.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the image into different regions (foreground and background) based on quality measurements and applies different encoding parameters to each segment. This segmentation allows efficient bit allocation by concentrating bits on important regions while using fewer bits for less important regions, resolving the contradiction between quality maintenance and bit allocation efficiency.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If multiple reconstructed pictures are generated and compared, then region of interest detection accuracy is improved, but computational complexity deteriorates

Engineering Contradiction:
Improveregion of interest detection accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the approach from using complex proxy networks with feature maps to a more direct quality measurement method. By computing quality measurements directly between different reconstructed pictures and comparing them, the system achieves accurate region of interest detection with reduced computational complexity, avoiding the overhead of complex neural network feature extraction.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240357104A1Determining regions of interest using learned image codec for machines
Publication Date: 2024.10.24 NOKIA TECHNOLOGIES OY
  • US20240357104A1 patent drawing
  • US20240357104A1 patent drawing
  • US20240357104A1 patent drawing

AI summary

Various embodiments describe an apparatus, a method, and a computer program product. An example apparatus includes at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: encoding an input picture by using a first encoder or first encoding parameters; encoding the input picture by using a second encoder or second encoding parameters; generating a first reconstructed picture based on the encoding of the input picture by using the first encoder or the first encoding parameters; and generating a second reconstructed picture based on the encoding of the input picture by using the second encoder or the second encoding parameters.