Machine Learning Model Divergence Assessment via Synthetic Data Indexing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for assessing the risk of error in machine learning models do not adequately account for the choice of training data and its distribution, leading to potential divergence in outcomes when faced with new, unseen situations, as they assume the training data is representative of all possible data and perfectly random.

Innovation Solution

A computer-implemented method that trains a model on a first set of observations, generates a second set of synthetic observations based on theoretical distributions, and indexes both sets to allow querying for potential divergence by comparing the density of tagged data against theoretically possible data, providing a rough estimate of model accuracy across all possible observations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If cross-validation is used to assess model error, then the model performance can be measured on test data, but all available tagged data cannot be used for training and the assessment is limited to observations similar to the test set

Engineering Contradiction:
Improvemodel performance measurementVSAvoidtraining data usage
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent generates synthetic copies of observations by randomly sampling from the joint probability distribution of input variables, creating artificial training data that expands the available dataset without requiring additional real-world tagged data. This allows the model to be trained on a much larger effective dataset while still assessing performance on the original tagged test set.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary generation of synthetic observations and training of the model on this expanded dataset before the actual performance assessment. This preliminary action of creating synthetic training data allows the model to learn from a more comprehensive distribution, improving its generalization capability while enabling thorough assessment on the original test data.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If delta and gamma tests are used to estimate model error, then the error can be estimated based on training samples alone, but the method applies only to smooth models and requires high data density

Engineering Contradiction:
Improveerror estimationVSAvoidmodel type applicability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the fundamental parameter of data representation by generating synthetic observations that capture the joint probability distribution of input variables. This allows the error estimation method to work with any model type (not just smooth models) because the synthetic data preserves the statistical properties of the original data without requiring the model to be smooth or continuous.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

By creating synthetic copies of the training data through random sampling from the joint distribution, the patent enables error estimation that is independent of model smoothness requirements. The synthetic data captures the essential statistical structure needed for error estimation while being applicable to any model type including decision trees and ensemble methods.

Inventive Principle:
Principle #26Copying

3Productivity

If the training data is assumed to be representative of all possible data, then the model can be trained efficiently, but the assessment of potential divergence for unseen situations is inadequate

Engineering Contradiction:
Improvetraining efficiencyVSAvoidoutcome accuracy for new data
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary generation of synthetic observations that explicitly explore the space of all possible data combinations based on the joint probability distribution. This preliminary action creates a comprehensive test set that covers unseen situations, allowing the model to be assessed for potential divergence before deployment while maintaining efficient training on the original data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the data assessment process into two distinct parts: training on the original tagged data and assessment on synthetic observations. This segmentation allows the model to be trained efficiently on real data while being separately evaluated for reliability on unseen situations, addressing both productivity and reliability concerns.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11176481B2Evaluation of a training set
Publication Date: 2021.11.16 DASSAULT SYSTEMES SA
  • US11176481B2 patent drawing
  • US11176481B2 patent drawing
  • US11176481B2 patent drawing

AI summary

A computer-implemented method for assessing a potential divergence of an outcome predicted by a machine learning system including training a model on a first set of observations, each observation being associated with a target value, randomly generating a second set of observations, applying the trained model to the second set thereby obtaining a target value associated with each observation of the second set, indexing the first and second sets of observations and their associated target values into an index, receiving a first query allowing a selection of a subset of the first and second sets of observations, generating a second query, generating a third query that comprises the first query and an additional constraint, querying the index using the second and third queries, and returning a response to the second and third queries.