Active Data Generation Using Error and Uncertainty Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating data for machine learning predictors are time-consuming, costly, and lack sufficient coverage of the data realm, especially since they only consider difficult-to-predict data points without accounting for prediction reliability, and are limited to 0/1 fields, which do not indicate error degrees.

Innovation Solution

A method that trains a first predictor on available data, determines its prediction errors, uses a second predictor to assess anticipated errors and uncertainty, and generates new data based on a description maximizing the combination of these factors, allowing for targeted and efficient data generation with improved reliability and coverage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automated methods select data points to be generated based only on difficulty to predict, then data generation can be automated, but prediction reliability is not accounted for and sufficient coverage of the data realm cannot be ensured

Engineering Contradiction:
Improveautomation of data point selectionVSAvoidprediction reliability
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent transitions from binary error classification (0/1 fields) to continuous error standards, allowing the system to evaluate prediction reliability on a spectrum. This enables the automated selection process to consider both difficulty to predict and prediction reliability simultaneously by quantifying uncertainty continuously rather than categorically.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces uncertainty quantification as an intermediary metric between the predictor and the data selection process. This intermediary layer allows the system to evaluate not just whether predictions are wrong, but how uncertain they are, enabling more nuanced automated data point selection that balances automation with reliability considerations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If data generation is minimized to reduce manual work, then expert effort is reduced, but coverage of the data realm may be insufficient

Engineering Contradiction:
Improvemanual work of expertVSAvoidcoverage of data realm
Core Design Contradiction:
Loss of timeVSArea of stationary object

Solution Approach 1:

The system performs self-service by automatically identifying and generating data points that maximize both uncertainty reduction and data realm coverage. The automated method selects which data points to generate based on uncertainty metrics, eliminating the need for expert intervention while ensuring comprehensive coverage through systematic exploration of the data space.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements feedback loops where the predictor's uncertainty estimates inform the selection of new data points, which are then used to retrain and improve the predictor. This continuous feedback cycle ensures that minimal manual work is required while systematically improving both prediction reliability and data realm coverage through iterative refinement.

Inventive Principle:
Principle #23Feedback

3Device complexity

If 0/1 fields are used to classify errors, then error classification is simple, but error degrees cannot be indicated and prediction quality estimation is limited

Engineering Contradiction:
Improveerror classification simplicityVSAvoiderror degree indication
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent changes the parameter representation from discrete binary values (0/1) to continuous uncertainty scores. This transformation maintains simplicity in the classification process while dramatically improving measurement precision, as the continuous scale allows for nuanced differentiation between various degrees of prediction error and confidence levels.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20220237514A1Active Data Generation Taking Uncertainties Into Consideration
Publication Date: 2022.07.28 CONTINENTAL AUTOMOTIVE TECHNOLOGIES GMBH
  • US20220237514A1 patent drawing
  • US20220237514A1 patent drawing
  • US20220237514A1 patent drawing

AI summary

The invention relates to a method for generating data on the basis of data already available having individual annotated data points, the method comprising:training a first predictor on the basis of data already available;determining a prediction error of the first predictor for each data point;training a second predictor to determine an anticipated prediction error of the first predictor and an uncertainty;determining a data description which maximizes a combination of anticipated prediction error and uncertainty; andgenerating data on the basis of the previously determined data description.