Semantic Segmentation Metrics for Content-Specific Image Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing visual fidelity metrics for image compression do not effectively assess the preservation of specific types of content, such as text or geometric shapes, in digital images, as they provide only general image quality indicators.
Innovation Solution
Implementing a machine learning model to infer segmentation masks that isolate specific content types, allowing for the calculation of content-specific visual fidelity metrics to optimize image compression schemes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If general visual fidelity metrics are used to assess image compression quality, then overall image quality can be evaluated, but specific content types (text, geometric shapes) cannot be assessed separately
Solution Approach 1:
The patent applies segmentation by dividing the image into different content types using semantic segmentation models. Each content type (text, geometric shapes, natural images) is segmented separately, allowing independent quality assessment for each type rather than evaluating the entire image as a single unit.
Solution Approach 2:
The patent implements local quality assessment by calculating visual fidelity metrics separately for each segmented content type. This allows different quality standards and metrics to be applied to different regions based on their content type, with text requiring higher fidelity than natural images.
2Quantity of substance
If lossy compression techniques are applied to reduce data size, then bandwidth requirements are reduced, but image quality deteriorates
Solution Approach 1:
The patent applies different compression techniques and quality parameters to different content types within the same image. Text regions use lossless or high-fidelity lossy compression, while natural image regions use more aggressive compression, optimizing the balance between data size and quality for each region.
Solution Approach 2:
The patent dynamically adjusts compression parameters based on content type detection. Semantic segmentation identifies different content types, and compression parameters (quality factor, bitrate, technique selection) are changed accordingly to preserve important content while reducing overall data size.
3Manufacturing precision
If different compression techniques are used for different content types, then visual fidelity for specific content is improved, but system complexity increases
Solution Approach 1:
The patent uses semantic segmentation models to automatically divide images into content types, enabling differentiated compression processing. This segmentation approach manages complexity by organizing the image into distinct regions that can be processed independently with appropriate techniques.
Solution Approach 2:
The patent introduces semantic segmentation models as an intermediary layer between the original image and the compression process. This intermediary analyzes content types and guides the selection of compression techniques, simplifying the overall system architecture by centralizing the decision-making logic.
Data Source
AI summary
This disclosure provides methods, devices, and systems for image compression. The present implementations more specifically relate to systems and techniques for selecting an image compression scheme for a given type of content or application. An image encoder may encode an image based on an image compression scheme. In some aspects, the image encoder may infer first and second segmentation masks from the original image and the encoded image, respectively, based a machine learning model. The machine learning model may be trained to extract one or more types of content from input images so that the segmentation masks include only the extracted content (and exclude any other types of content) from the images. The image encoder may further calculate a visual fidelity metric for the encoded image based on the masks and selectively transmit the encoded image over a communication channel based at least in part on the visual fidelity metric.


