Defective Training Data Collection Using Feature-Based Stop Criteria
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing training data collection systems for machine learning in inspection devices fail to adequately evaluate the quality and amount of defective product data, leading to insufficient learning of classification models and unnecessary data collection.
Innovation Solution
A method for collecting defective product data using feature extraction and calculation of state sums and index values to quantify data quality and amount, ensuring high accuracy and optimal termination of data collection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If defective product data is collected to improve classification model accuracy, then the quality and amount of training data increase, but the collection time and man-hours increase
Solution Approach 1:
The patent implements a feedback mechanism where the classification model evaluates the quality of collected defective product data in real-time. The system calculates evaluation indices based on feature quantities extracted from the data, and uses this feedback to determine whether to continue or terminate data collection. This ensures that data collection stops at the optimal point when sufficient quality and quantity are achieved, preventing unnecessary time consumption while ensuring model accuracy.
Solution Approach 2:
The patent replaces manual evaluation of data quality with an automated computer-based evaluation system. Instead of manually assessing whether collected defective product data is sufficient, the system automatically extracts feature quantities, calculates evaluation indices, and makes termination decisions. This substitution dramatically reduces the time and labor required for data collection while maintaining scientific rigor.
2Reliability
If data collection continues to ensure sufficient quality and amount, then training data quality improves, but unnecessary collection occurs wasting resources
Solution Approach 1:
The system continuously monitors the quality and quantity of collected defective product data through automated evaluation indices. When these indices reach predetermined thresholds indicating sufficient data quality and amount, the system automatically terminates collection. This feedback loop prevents both insufficient data collection (which would harm model accuracy) and excessive collection (which would waste resources), optimizing both data quality and collection efficiency.
Solution Approach 2:
The patent changes the evaluation parameters from subjective manual assessment to objective quantitative metrics. By extracting feature quantities from the defective product data and calculating evaluation indices, the system transforms the quality assessment into measurable parameters with clear thresholds. This allows for precise determination of when data collection should terminate, improving efficiency while ensuring adequate quality.
3Measurement precision
If manual evaluation of training data quality is performed, then data quality can be assessed, but the evaluation process is subjective and inconsistent
Solution Approach 1:
The patent replaces subjective manual evaluation with an automated computer-based evaluation system that objectively measures data quality. The system extracts feature quantities from the defective product data and calculates evaluation indices using predetermined algorithms, eliminating human subjectivity and inconsistency. While this increases computational complexity, it dramatically improves measurement precision and reliability of the quality assessment.
Data Source
AI summary
The present invention relates to a collecting method for training data for collecting defective product data as training data for learning a classification model that classifies the inspected object as a normal product or an abnormal product, including: collecting many pieces of defective product data (step 1 of FIG. 3); extracting a plurality of feature quantities respectively from many pieces of the defective product data (step 2); calculating a state sum for every feature quantity of the plurality of feature quantities that have been extracted, for the many pieces of the defective product data (step 3); calculating, as an index value, a logarithmic sum of a plurality of state sums that have been calculated (step 4); and ending collecting the defective product data, in a case where the calculated index value is equal to or greater than a predetermined target value (step 5).


