Confidence-Based Data Filtering for Optical Inspection Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in efficiently collecting, storing, and utilizing low confidence data for machine learning systems in optical inspection, which is crucial for improving the accuracy of artificial intelligence systems.

Innovation Solution

A method that involves collecting measurement data, determining a confidence value associated with the data, and storing the data if the confidence value is below a threshold, with the option to use this data for training machine learning systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If all measurement data is collected and stored for machine learning training, then the quantity of training data increases, but storage resources and data processing complexity increase proportionally

Engineering Contradiction:
Improvequantity of training dataVSAvoiddata processing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent applies local quality by differentiating data based on confidence values. Instead of treating all data uniformly, the system identifies and prioritizes storage of low-confidence data (data where the inspection system is uncertain about defect presence). This selective approach focuses resources on the most valuable training data while avoiding redundant storage of high-confidence data, thus increasing effective training data quantity without proportionally increasing storage complexity

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the parameter of data selection by using confidence values as a filtering criterion. By setting a confidence threshold parameter, the system automatically identifies which data should be stored for training purposes. This parameter-based selection transforms the data collection process from storing everything to storing only data that meets specific confidence criteria, resolving the contradiction between data quantity and processing complexity

Inventive Principle:
Principle #35Parameter changes

2Loss of energy

If low confidence data is selectively stored based on confidence thresholds, then storage efficiency improves, but the complexity of determining confidence values increases

Engineering Contradiction:
Improvestorage efficiencyVSAvoidconfidence determination complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The inspection system performs self-service by automatically generating confidence values for its own measurement data. The machine learning model inherently produces confidence metrics as part of its normal operation, and the system uses these self-generated confidence values to automatically filter and prioritize storage decisions. This eliminates the need for external or separate confidence determination mechanisms, improving storage efficiency without adding significant complexity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback by using the machine learning model's output confidence values to guide data storage decisions. The confidence determination is integrated into the existing inspection workflow, where the model's confidence assessment feeds directly into the storage selection process. This feedback loop allows the system to efficiently identify valuable training data using existing model capabilities, avoiding the need for separate complex confidence determination systems

Inventive Principle:
Principle #23Feedback

3Measurement precision

If confidence threshold filtering is applied to measurement data, then the quality of training data improves, but the time required to process and evaluate confidence values increases

Engineering Contradiction:
Improvequality of training dataVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by determining confidence values at the point of data generation, immediately when the machine learning model processes the measurement data. Rather than adding a separate post-processing step to evaluate confidence, the system integrates confidence determination into the primary inspection workflow. This preliminary assessment of data quality allows for immediate filtering decisions, improving training data quality without adding significant processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent merges the confidence determination function with the existing machine learning inspection process. The same model that performs defect detection also generates confidence values, combining multiple functions into a single processing step. This merging eliminates the need for separate confidence evaluation processes, allowing quality filtering to be applied without proportionally increasing processing time

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12278940B2Automated inspection data collection for machine learning applications
Publication Date: 2025.04.15 QCIFY INC
  • US12278940B2 patent drawing
  • US12278940B2 patent drawing
  • US12278940B2 patent drawing

AI summary

A method includes collecting measurement data of a sample, determining a confidence value associated with the measurement data, determining if the confidence value is less than a confidence threshold value, and causing the measurement data to be stored in a memory device if the confidence value is less than the confidence threshold value. The method further includes utilizing the measurement data for training of a machine learning system, thereby updating the operation of the machine learning system. The measurement data is measured by an inspection device located at a first location. In one example, the measurement data is a captured image captured by an optical inspector. The confidence value associated with the captured image represents the confidence of the machine learning system regarding what is displayed in the captured image.