Model Test Data Validation via Perturbation Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing approaches to generate synthetic data for machine learning models by perturbing existing training data lack the ability to determine the applicability of the perturbed data for specific use cases and domain-specific constraints, leading to erroneous evaluations of the generated models.
Innovation Solution
The proposed method generates and executes model-specific test cases that consider the context in which the models are used, validating perturbed data against metadata and quality metrics to ensure relevance and applicability, thereby reducing the amount of data needed for assessment and improving processing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If perturbed training data is used for model evaluation, then data availability is improved, but data quality and relevance deteriorate
Solution Approach 1:
The patent introduces an intermediary validation system that acts as a mediator between perturbed training data and model evaluation. This validation system checks whether perturbed data meets quality thresholds and contextual relevance criteria before being used for evaluation, thus resolving the contradiction by filtering low-quality perturbed data while maintaining data availability.
Solution Approach 2:
The patent changes the parameters of perturbed data through validation processes, adjusting quality metrics and relevance scores to determine suitability for model evaluation. By dynamically adjusting these parameters based on validation results, the system maintains data availability while ensuring adequate data quality.
2Reliability
If comprehensive test cases are generated with multiple hyperparameters, then model validation thoroughness is improved, but processing time and resources increase
Solution Approach 1:
The patent applies partial action by generating test cases selectively rather than exhaustively. It creates a sufficient subset of test cases that cover critical validation scenarios without requiring complete coverage of all possible hyperparameter combinations, thus achieving adequate model validation thoroughness while reducing processing time.
Solution Approach 2:
The patent performs preliminary validation of perturbed data before generating test cases. This preliminary action filters out inadequate data early in the process, preventing wasted processing time on low-quality data while maintaining thoroughness in validating models with high-quality test data.
3Productivity
If perturbed data is used without validation, then data processing speed is improved, but evaluation accuracy deteriorates
Solution Approach 1:
The patent performs preliminary validation of perturbed data before it is used for model evaluation. This preliminary action quickly checks data quality and relevance criteria, enabling fast processing while ensuring that only adequately validated data proceeds to evaluation, thus maintaining both processing speed and evaluation accuracy.
Solution Approach 2:
The patent implements a feedback mechanism where validation results of perturbed data influence whether the data is used for evaluation. This feedback loop ensures that data processing speed is maintained through automated validation while evaluation accuracy is preserved by rejecting inadequate data based on validation feedback.
Data Source
AI summary
Embodiments of the present invention provide computer-implemented methods, computer program products and computer systems. Embodiments of the present invention can, in response to receiving information, generate a data profile for a model that includes metadata for data requirements, model specific requirements, and data quality metrics. Embodiments of the present invention can generate one or more perturbations for training data associated with the received information and validate at least one perturbation of the one or more perturbations of training data as relevant test data based, at least in part on context associated with the model. Embodiments of the present invention can then generate one or more test scenarios based on the at least one validated perturbation and varying hyperparameters of the model and generate a test report based on an execution of at least one generated test scenario of the generated one or more test scenarios.


