Machine Learning Model Divergence Assessment via Synthetic Data Indexing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for assessing the risk of error in machine learning models do not adequately account for the choice of training data and its distribution, leading to potential divergence in outcomes when faced with new, unseen situations, as they assume the training data is representative of all possible data and perfectly random.
Innovation Solution
A computer-implemented method that trains a model on a first set of observations, generates a second set of synthetic observations based on theoretical distributions, and indexes both sets to allow querying for potential divergence by comparing the density of tagged data against theoretically possible data, providing a rough estimate of model accuracy across all possible observations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If cross-validation is used to assess model error, then the model performance can be measured on test data, but all available tagged data cannot be used for training and the assessment is limited to observations similar to the test set
Solution Approach 1:
The patent generates synthetic copies of observations by randomly sampling from the joint probability distribution of input variables, creating artificial training data that expands the available dataset without requiring additional real-world tagged data. This allows the model to be trained on a much larger effective dataset while still assessing performance on the original tagged test set.
Solution Approach 2:
The patent performs preliminary generation of synthetic observations and training of the model on this expanded dataset before the actual performance assessment. This preliminary action of creating synthetic training data allows the model to learn from a more comprehensive distribution, improving its generalization capability while enabling thorough assessment on the original test data.
2Measurement precision
If delta and gamma tests are used to estimate model error, then the error can be estimated based on training samples alone, but the method applies only to smooth models and requires high data density
Solution Approach 1:
The patent changes the fundamental parameter of data representation by generating synthetic observations that capture the joint probability distribution of input variables. This allows the error estimation method to work with any model type (not just smooth models) because the synthetic data preserves the statistical properties of the original data without requiring the model to be smooth or continuous.
Solution Approach 2:
By creating synthetic copies of the training data through random sampling from the joint distribution, the patent enables error estimation that is independent of model smoothness requirements. The synthetic data captures the essential statistical structure needed for error estimation while being applicable to any model type including decision trees and ensemble methods.
3Productivity
If the training data is assumed to be representative of all possible data, then the model can be trained efficiently, but the assessment of potential divergence for unseen situations is inadequate
Solution Approach 1:
The patent performs preliminary generation of synthetic observations that explicitly explore the space of all possible data combinations based on the joint probability distribution. This preliminary action creates a comprehensive test set that covers unseen situations, allowing the model to be assessed for potential divergence before deployment while maintaining efficient training on the original data.
Solution Approach 2:
The patent segments the data assessment process into two distinct parts: training on the original tagged data and assessment on synthetic observations. This segmentation allows the model to be trained efficiently on real data while being separately evaluated for reliability on unseen situations, addressing both productivity and reliability concerns.
Data Source
AI summary
A computer-implemented method for assessing a potential divergence of an outcome predicted by a machine learning system including training a model on a first set of observations, each observation being associated with a target value, randomly generating a second set of observations, applying the trained model to the second set thereby obtaining a target value associated with each observation of the second set, indexing the first and second sets of observations and their associated target values into an index, receiving a first query allowing a selection of a subset of the first and second sets of observations, generating a second query, generating a third query that comprises the first query and an additional constraint, querying the index using the second and third queries, and returning a response to the second and third queries.


