Active Data Generation Using Error and Uncertainty Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating data for machine learning predictors are time-consuming, costly, and lack sufficient coverage of the data realm, especially since they only consider difficult-to-predict data points without accounting for prediction reliability, and are limited to 0/1 fields, which do not indicate error degrees.
Innovation Solution
A method that trains a first predictor on available data, determines its prediction errors, uses a second predictor to assess anticipated errors and uncertainty, and generates new data based on a description maximizing the combination of these factors, allowing for targeted and efficient data generation with improved reliability and coverage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If automated methods select data points to be generated based only on difficulty to predict, then data generation can be automated, but prediction reliability is not accounted for and sufficient coverage of the data realm cannot be ensured
Solution Approach 1:
The patent transitions from binary error classification (0/1 fields) to continuous error standards, allowing the system to evaluate prediction reliability on a spectrum. This enables the automated selection process to consider both difficulty to predict and prediction reliability simultaneously by quantifying uncertainty continuously rather than categorically.
Solution Approach 2:
The patent introduces uncertainty quantification as an intermediary metric between the predictor and the data selection process. This intermediary layer allows the system to evaluate not just whether predictions are wrong, but how uncertain they are, enabling more nuanced automated data point selection that balances automation with reliability considerations.
2Loss of time
If data generation is minimized to reduce manual work, then expert effort is reduced, but coverage of the data realm may be insufficient
Solution Approach 1:
The system performs self-service by automatically identifying and generating data points that maximize both uncertainty reduction and data realm coverage. The automated method selects which data points to generate based on uncertainty metrics, eliminating the need for expert intervention while ensuring comprehensive coverage through systematic exploration of the data space.
Solution Approach 2:
The patent implements feedback loops where the predictor's uncertainty estimates inform the selection of new data points, which are then used to retrain and improve the predictor. This continuous feedback cycle ensures that minimal manual work is required while systematically improving both prediction reliability and data realm coverage through iterative refinement.
3Device complexity
If 0/1 fields are used to classify errors, then error classification is simple, but error degrees cannot be indicated and prediction quality estimation is limited
Solution Approach 1:
The patent changes the parameter representation from discrete binary values (0/1) to continuous uncertainty scores. This transformation maintains simplicity in the classification process while dramatically improving measurement precision, as the continuous scale allows for nuanced differentiation between various degrees of prediction error and confidence levels.
Data Source
AI summary
The invention relates to a method for generating data on the basis of data already available having individual annotated data points, the method comprising:training a first predictor on the basis of data already available;determining a prediction error of the first predictor for each data point;training a second predictor to determine an anticipated prediction error of the first predictor and an uncertainty;determining a data description which maximizes a combination of anticipated prediction error and uncertainty; andgenerating data on the basis of the previously determined data description.


