Image Transformer Training for Accurate Synthetic Image Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning recognizers face challenges in achieving high recognition accuracy due to the high cost and difficulty in collecting and labeling real-domain image data, while simulation-domain data is easily generated but lacks accuracy.

Innovation Solution

A method involving training a recognizer using composite images, generating labeled real images through an image transformer, and iteratively improving the recognizer's performance using a data set of paired real and composite images, reducing the need for costly real-image capture and labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If image data is collected from the real domain using cameras, then recognition accuracy is improved, but data collection cost and labeling cost increase significantly

Engineering Contradiction:
Improverecognition accuracyVSAvoiddata collection cost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent creates virtual copies of real images through composite image generation. Instead of collecting expensive real-domain images, the system generates synthetic composite images that replicate real image characteristics, thereby reducing data collection costs while maintaining recognition accuracy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transforms real images into composite images by changing domain parameters. The image transformer model converts real-domain image data into composite-domain representations, enabling the use of cheaper synthetic data while preserving the essential features needed for accurate recognition

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If composite images from simulation domain are used for training, then data generation cost is reduced, but recognition accuracy deteriorates

Engineering Contradiction:
Improvedata generation costVSAvoidrecognition accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent introduces an image transformer model as an intermediary between real images and composite images. This intermediary transforms real image data into composite image format, allowing the system to use cheap composite images for training while maintaining the quality characteristics of real images, thus resolving the accuracy problem

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary transformation of real image characteristics into composite image format before actual recognition training. By pre-processing real image data through the image transformer, the composite images are prepared in advance with the necessary quality attributes, ensuring accurate recognition results

Inventive Principle:
Principle #10Preliminary action

3Reliability

If more real images are collected and labeled to improve recognizer performance, then recognition accuracy is improved, but time and resources required increase

Engineering Contradiction:
Improverecognizer performanceVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system creates virtual copies of labeled real images through composite image generation. Instead of spending time collecting and labeling additional real images, the system generates synthetic composite images that inherit the labeled information structure, dramatically reducing data preparation time while maintaining recognizer performance

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The image transformer enables the system to self-generate training data in the composite domain from existing real image labels. The system serves itself by automatically creating labeled composite images without manual intervention, eliminating the time-consuming manual labeling process while improving data availability

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12573074B2Machine learning method, machine learning system, and program
Publication Date: 2026.03.10 TOYOTA JIDOSHA KK
  • US12573074B2 patent drawing
  • US12573074B2 patent drawing
  • US12573074B2 patent drawing

AI summary

A machine learning method according to this embodiment trains a recognizer using a composite image, acquires a labeled real image based on an image captured by a sensor, stores, when at least one of the results of the recognition when the labeled real image is input to the recognizer match the label, results of the recognition performed by the recognizer, performs machine learning using a data set group including a plurality of the data sets, a real image and a composite image forming a pair in each of the data sets, thereby generating an image transformer, generating a labeled real image as a result of the image transformer transforming the composite image into the real image, and trains the recognizer based on the labeled real image.