Adaptive Training Dataset Generation for Defect Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In the manufacturing field, it is challenging to define classes and perform labeling operations for training datasets used in machine learning models, especially when predicting defects in equipment, due to subjectivity and high costs associated with verifying completed training datasets.

Innovation Solution

A method and apparatus for generating an adaptive training image dataset by preparing initial and verification datasets, using a learning algorithm to generate models, and updating classes based on similarity calculations between images and their predicted classes, thereby improving prediction accuracy and reducing verification costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a training dataset is configured in advance with predefined classes and labeling operations, then the dataset can be used as input for learning algorithms, but it requires significant money and man-hours for verification and does not adapt to unpredictable defects

Engineering Contradiction:
Improveadaptability to defect typesVSAvoidverification time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-calculating similarity metrics between training images and verification images, and pre-defining class structures. This allows the verification process to quickly compare predicted classes against pre-computed similarities rather than performing complex analyses during verification, significantly reducing verification time while maintaining adaptability to various defect types

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary verification process that acts as a mediator between the training dataset creation and the actual learning algorithm execution. This verification stage uses similarity calculations as an intermediary mechanism to objectively determine class assignments, reducing the need for expensive manual verification while adapting to unpredictable defect types that may emerge during equipment operation

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple users perform labeling operations to create a training dataset, then diverse perspectives are incorporated, but subjectivity is reflected in the labeling which reduces reliability

Engineering Contradiction:
Improvediversity of labeling perspectivesVSAvoidlabeling consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system implements feedback mechanisms where the learning model's predictions are continuously compared against the verification dataset, and similarity metrics provide feedback on classification accuracy. This feedback loop allows the system to objectively adjust and refine class assignments, maintaining the diversity of perspectives from multiple labelers while eliminating subjectivity through quantitative similarity-based verification

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent transforms the subjective labeling process into an objective parameter-based system by calculating similarity parameters between images and predefined class representatives. This parameter change from subjective human judgment to objective similarity metrics allows diverse labeling perspectives to be incorporated while ensuring consistency through mathematical comparison rather than human subjectivity

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If extensive manual verification is performed on the training dataset, then labeling accuracy is improved, but the cost and time requirements increase significantly

Engineering Contradiction:
Improvelabeling accuracyVSAvoidverification time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies partial verification action by performing similarity-based verification only on critical aspects of the training dataset rather than exhaustive manual verification of all images. This selective verification approach maintains high labeling accuracy for the most important classification decisions while significantly reducing the overall time and cost burden of verification

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent substitutes the mechanical system of manual human verification with an automated computational system that calculates similarity metrics between images and class representatives. This substitution replaces time-consuming manual inspection with rapid computational comparison, maintaining high labeling accuracy through objective mathematical measures while dramatically reducing verification time and costs

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20230138430A1Method and apparatus for generating an adaptive training image dataset
Publication Date: 2023.05.04 SYSTEM ENGINEERING MEGA SOLUTION CO LTD
  • US20230138430A1 patent drawing
  • US20230138430A1 patent drawing
  • US20230138430A1 patent drawing

AI summary

A method for generating an adaptive training image dataset capable of improving prediction accuracy is provided. The method comprises preparing a first image dataset, wherein the first image dataset includes a plurality of first images and a first class corresponding to each of the plurality of first images, performing a learning algorithm on the first image dataset to generate a first learning model, preparing a second image dataset, wherein the second image dataset includes a plurality of second images and a second class corresponding to each of the plurality of second images, inputting the second image to the first learning model to obtain a prediction class corresponding to the second image, determining which class of image the second image is similar to in case of the second class corresponding to the second image being different from the prediction class, updating the second image dataset by updating a class corresponding to the second image according to the determination result, and performing a learning algorithm on the first image dataset and the updated second image dataset to generate a second learning model.