Image Reconstruction Learning for Reliable Features Without GAN Artifacts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image reconstruction methods using generative adversarial networks (GANs) often reproduce unnecessary fine patterns and fail to improve visual quality perceived by human vision, despite improving recognition rates and reducing distortion.
Innovation Solution
An information processing system utilizing multiple machine learning models to identify and reconstruct image features, with a combined loss function that enhances reliability and recognition accuracy by adjusting parameter sets to minimize distortion and error.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a GAN is applied to extreme image compression to improve recognition rate and reduce distortion, then the recognition rate is improved and distortion is reduced, but unnecessary fine patterns are reproduced and visual quality perceived by human vision is not improved
Solution Approach 1:
The patent extracts and separates the semantic information from the fine pattern information. By using a semantic encoder that processes only semantic information and a separate fine pattern encoder that processes fine pattern information, the system can control what information is preserved during compression. The semantic encoder maintains recognition accuracy while the fine pattern encoder can be controlled to avoid reproducing unnecessary fine patterns, thus resolving the contradiction between improving recognition rate and avoiding unnecessary fine patterns.
Solution Approach 2:
The patent segments the image processing into distinct functional components: a semantic encoder for semantic information, a fine pattern encoder for fine pattern information, and a separate decoder that combines these independently processed components. This segmentation allows independent optimization of semantic fidelity (for recognition rate) and fine pattern control (to avoid unnecessary patterns), resolving the contradiction between these two requirements.
2Reliability
If additional information on the identification target is provided to the decoder to improve recognition rate, then the recognition rate is improved, but the device complexity increases
Solution Approach 1:
The semantic encoder and fine pattern encoder are designed to process different types of information (semantic and fine pattern) using the same architectural framework and processing pipeline. This universality allows the system to improve recognition rate through semantic information processing without requiring separate dedicated components, thus reducing the increase in device complexity while maintaining improved recognition rate.
Data Source
AI summary
A model learning device determines a first machine learning model so as to further increase a combined loss function obtained by combining: a first loss function indicating the level of change in the reliability of a second image feature in a feature region of a reconstructed image, from the reliability of a first image feature in a feature region of the original image; and a second loss function indicating the level of recognition error. In addition, the model learning device: determines, so as to further reduce the combined loss function, respective parameter sets for a second machine learning model used in the generation of compressed data, and a third machine learning model used in the generation of the reconstructed image from the compressed data; and determines a parameter set for the fourth machine learning model in common with that for the first machine learning model.


