Image Transformer Training for Accurate Synthetic Image Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning recognizers face challenges in achieving high recognition accuracy due to the high cost and difficulty in collecting and labeling real-domain image data, while simulation-domain data is easily generated but lacks accuracy.
Innovation Solution
A method involving training a recognizer using composite images, generating labeled real images through an image transformer, and iteratively improving the recognizer's performance using a data set of paired real and composite images, reducing the need for costly real-image capture and labeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If image data is collected from the real domain using cameras, then recognition accuracy is improved, but data collection cost and labeling cost increase significantly
Solution Approach 1:
The patent creates virtual copies of real images through composite image generation. Instead of collecting expensive real-domain images, the system generates synthetic composite images that replicate real image characteristics, thereby reducing data collection costs while maintaining recognition accuracy
Solution Approach 2:
The system transforms real images into composite images by changing domain parameters. The image transformer model converts real-domain image data into composite-domain representations, enabling the use of cheaper synthetic data while preserving the essential features needed for accurate recognition
2Quantity of substance
If composite images from simulation domain are used for training, then data generation cost is reduced, but recognition accuracy deteriorates
Solution Approach 1:
The patent introduces an image transformer model as an intermediary between real images and composite images. This intermediary transforms real image data into composite image format, allowing the system to use cheap composite images for training while maintaining the quality characteristics of real images, thus resolving the accuracy problem
Solution Approach 2:
The system performs preliminary transformation of real image characteristics into composite image format before actual recognition training. By pre-processing real image data through the image transformer, the composite images are prepared in advance with the necessary quality attributes, ensuring accurate recognition results
3Reliability
If more real images are collected and labeled to improve recognizer performance, then recognition accuracy is improved, but time and resources required increase
Solution Approach 1:
The system creates virtual copies of labeled real images through composite image generation. Instead of spending time collecting and labeling additional real images, the system generates synthetic composite images that inherit the labeled information structure, dramatically reducing data preparation time while maintaining recognizer performance
Solution Approach 2:
The image transformer enables the system to self-generate training data in the composite domain from existing real image labels. The system serves itself by automatically creating labeled composite images without manual intervention, eliminating the time-consuming manual labeling process while improving data availability
Data Source
AI summary
A machine learning method according to this embodiment trains a recognizer using a composite image, acquires a labeled real image based on an image captured by a sensor, stores, when at least one of the results of the recognition when the labeled real image is input to the recognizer match the label, results of the recognition performed by the recognizer, performs machine learning using a data set group including a plurality of the data sets, a real image and a composite image forming a pair in each of the data sets, thereby generating an image transformer, generating a labeled real image as a result of the image transformer transforming the composite image into the real image, and trains the recognizer based on the labeled real image.


