Neural Network Classification Model Training Using Uncertainty-Based Data Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for training deep learning models face challenges in updating models effectively when data changes, particularly in industrial applications where sensor data is used, and there is a need for efficient processing and labeling of data to enhance classification model performance.
Innovation Solution
A method involving a computer program executed by processors to determine uncertainty and similarity levels of data, selectively label data based on these levels, and update classification models using the labeled data, optimizing data selection and labeling to improve model performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all data in the dataset are labeled to improve model performance, then classification accuracy is improved, but labeling costs and time consumption increase significantly
Solution Approach 1:
The patent applies partial action by selectively labeling only a subset of data rather than all data. It uses uncertainty quantification to identify data points that最需要 labeling (those with high uncertainty), and labels only those specific portions while leaving other data unlabeled, thus achieving good classification accuracy without the time cost of labeling everything
Solution Approach 2:
The system performs self-service by automatically identifying which data points require labeling through uncertainty quantification. The model itself determines its own learning needs by calculating uncertainty levels for different data points, eliminating the need for manual assessment of which data should be labeled
2Measurement precision
If data with high uncertainty are selected for labeling to improve model accuracy, then classification performance is enhanced, but labeling costs increase if not properly optimized
Solution Approach 1:
The patent implements partial action by labeling only the specific portion of data that has high uncertainty and provides the most value for model improvement. Rather than labeling a fixed proportion or all uncertain data, it selectively labels only those data points that will most effectively reduce uncertainty and improve classification accuracy
Solution Approach 2:
The system dynamically adjusts the selection criteria for labeling based on uncertainty levels. It changes the parameter of data selection from random or uniform sampling to uncertainty-based sampling, where data points are selected for labeling based on their calculated uncertainty values, ensuring optimal use of labeling resources
3Adaptability or versatility
If a large amount of data is labeled to handle data distribution changes, then model adaptability is improved, but processing time and costs increase
Solution Approach 1:
The system enables self-service by automatically detecting when data distribution changes occur and identifying which specific data points should be relabeled. The model monitors its own uncertainty levels and automatically determines when and what to update, eliminating the need for manual intervention or comprehensive relabeling when adapting to new data patterns
Solution Approach 2:
The patent applies partial action by updating only the specific data points that show high uncertainty due to distribution changes, rather than relabeling the entire dataset. This selective updating approach maintains model adaptability while significantly reducing the processing time and resources required compared to full dataset relabeling
Data Source
AI summary
Disclosed is a non-transitory computer readable medium storing a computer program. When the computer program is executed by one or more processors of a computing device, the computer program performs the following operations for processing data, and the operations may include: determining an uncertainty level with respect to labeling criteria for each of one or more data included in a dataset; determining a similarity level for one or more data included in a data subset; and selecting at least some of data included in the dataset based on the uncertainty level and the similarity level, and additionally labeling the selected data.


