Hierarchical Auto-regressive Image Compression with Saliency Masks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image compression methods often result in increased storage and bandwidth requirements due to redundancies in coded bit strings and fail to optimally allocate bits to salient regions of interest, leading to suboptimal image quality.
Innovation Solution
A multi-stage encoder/decoder system using hierarchical auto-regressive models and saliency-based masks to identify and remove redundancies, allocating more bits to salient regions and weighting distortions differently across the image, thereby improving perceived clarity without increasing storage and bandwidth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional image compression methods are used, then file size is reduced, but image quality deteriorates due to loss of important visual information
Solution Approach 1:
The patent applies local quality by differentiating between salient and non-salient regions in the image. Salient regions (containing important visual information) are encoded with higher quality and more bits, while non-salient regions are compressed more aggressively. This is achieved through saliency detection mechanisms that identify important regions and allocate compression resources accordingly, resolving the contradiction by making compression quality spatially variable rather than uniform.
Solution Approach 2:
The patent segments the image into multiple regions based on saliency maps that identify salient versus non-salient areas. Different compression strategies are applied to different segments: high-quality encoding for salient regions and aggressive compression for non-salient regions. This segmentation approach allows the system to maintain overall image quality while achieving better compression ratios.
2Ease of operation
If uniform compression is applied across the entire image, then encoding simplicity is maintained, but visual perception quality deteriorates because non-salient regions waste bandwidth while salient regions lack sufficient detail
Solution Approach 1:
The patent implements local quality by applying different compression parameters to different image regions. Salient regions receive higher bit allocation and less aggressive compression, while non-salient regions receive lower bit allocation and more aggressive compression. This is controlled through saliency detection and region-based encoding parameters, improving visual perception quality without significantly complicating the encoding process.
Solution Approach 2:
The patent changes compression parameters dynamically based on regional saliency. Different quantization parameters, bit allocation strategies, and encoding precision are applied to salient versus non-salient regions. This parameter differentiation allows the system to optimize visual quality where it matters most while maintaining encoding efficiency.
3Manufacturing precision
If more bits are allocated to capture details in all regions, then image quality is improved, but storage and bandwidth requirements increase
Solution Approach 1:
The patent applies local quality by allocating more bits specifically to salient regions rather than uniformly across the entire image. This targeted bit allocation ensures that important visual information is preserved with high quality while non-salient regions use fewer bits, thereby improving overall image quality without proportionally increasing storage and bandwidth requirements.
Solution Approach 2:
The patent segments the image into salient and non-salient regions and applies different bit allocation strategies to each segment. Salient regions receive higher bit rates to preserve detail, while non-salient regions receive lower bit rates. This segmentation-based approach optimizes the trade-off between image quality and storage/bandwidth requirements by focusing resources where they provide the most visual value.
Data Source
AI summary
The present application relates to a multi-stage encoder/decoder system that provides image compression using hierarchical auto-regressive models and saliency-based masks. The multi-stage encoder/decoder system includes a first stage and a second stage of a trained image compression network, such that the second stage, based on the image compression performed by the first stage, identify certain redundancies that can be removed from the bit string to reduce the storage and bandwidth requirements. Additionally, by using saliency-based masks, distortions in different sections of the image can be weighted differently to further improve the image compression performance.


