AI Image Encoding and Decoding With ROI Scaling for Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image encoding and decoding technologies struggle to achieve high compression rates for high-resolution/high-quality images without significantly degrading image quality.
Innovation Solution
An AI-based encoding and decoding method that utilizes a scaling neural network to identify and process object regions of interest, applying different neural network settings for varying image parts, and includes AI data for accurate upscaling during decoding to maintain image quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of substance
If conventional image encoding is used to compress high-resolution images, then compression rate is improved, but image quality is significantly degraded
Solution Approach 1:
The image is divided into multiple regions of interest (ROIs) and non-ROI areas. Different encoding strategies are applied to each region: AI-based super-resolution is applied to ROI regions to maintain high quality, while conventional compression is applied to non-ROI areas to maximize compression. This segmentation allows the system to achieve high overall compression rates while preserving critical image quality in important regions.
Solution Approach 2:
Different quality levels and encoding methods are applied to different parts of the image based on their importance. ROI regions receive high-quality AI super-resolution processing with minimal compression, while non-ROI regions undergo aggressive compression. This local quality differentiation resolves the contradiction by maintaining high quality where needed while achieving overall compression efficiency.
2Manufacturing precision
If AI super-resolution is applied to the entire image, then image quality is maintained, but compression rate decreases
Solution Approach 1:
The image is segmented into ROI and non-ROI regions. AI super-resolution is selectively applied only to ROI regions where quality is critical, while non-ROI regions use conventional compression. This selective application maintains image quality in important areas while achieving overall compression rate improvements.
Solution Approach 2:
Instead of applying AI super-resolution to the entire image (excessive action), the system applies it only to necessary ROI regions (partial action). This partial application achieves the minimum necessary quality maintenance while maximizing compression efficiency in non-critical areas.
3Manufacturing precision
If different neural network settings are used for different image regions, then image quality in ROIs is improved, but processing complexity increases
Solution Approach 1:
The image processing system segments the image into ROI and non-ROI regions and applies different neural network configurations to each. ROI regions use neural networks optimized for high-quality super-resolution, while non-ROI regions use simpler processing. This segmentation enables differentiated quality processing while managing complexity through region-based strategies.
Solution Approach 2:
The neural network settings are dynamically selected based on region importance. The system adjusts network complexity, layer depth, and processing parameters according to whether a region is marked as ROI or not. This dynamic adaptation allows high quality in ROIs while reducing processing complexity in non-ROI areas.
Data Source
AI summary
An artificial intelligence (AI) encoding apparatus includes a memory storing one or more instructions, and a processor configured to execute the one or more instructions stored in the memory to identify an object region of interest in an original image, obtain, from the original image, a plurality of original part images respectively including the object region of interest and a non-interest region, obtain a plurality of first images by performing AI scaling on the plurality of original part images through a scaling neural network (NN) that is configured to operate with NN setting information selected from among a plurality of pieces of NN setting information, at least based on whether the plurality of original part images include the object region of interest or the non-interest region, generate image data by encoding the plurality of first images, and transmit the image data, and AI data including information related to the AI scaling.


