Model Test Data Validation via Perturbation Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing approaches to generate synthetic data for machine learning models by perturbing existing training data lack the ability to determine the applicability of the perturbed data for specific use cases and domain-specific constraints, leading to erroneous evaluations of the generated models.

Innovation Solution

The proposed method generates and executes model-specific test cases that consider the context in which the models are used, validating perturbed data against metadata and quality metrics to ensure relevance and applicability, thereby reducing the amount of data needed for assessment and improving processing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If perturbed training data is used for model evaluation, then data availability is improved, but data quality and relevance deteriorate

Engineering Contradiction:
Improvedata availabilityVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary validation system that acts as a mediator between perturbed training data and model evaluation. This validation system checks whether perturbed data meets quality thresholds and contextual relevance criteria before being used for evaluation, thus resolving the contradiction by filtering low-quality perturbed data while maintaining data availability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameters of perturbed data through validation processes, adjusting quality metrics and relevance scores to determine suitability for model evaluation. By dynamically adjusting these parameters based on validation results, the system maintains data availability while ensuring adequate data quality.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If comprehensive test cases are generated with multiple hyperparameters, then model validation thoroughness is improved, but processing time and resources increase

Engineering Contradiction:
Improvemodel validation thoroughnessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by generating test cases selectively rather than exhaustively. It creates a sufficient subset of test cases that cover critical validation scenarios without requiring complete coverage of all possible hyperparameter combinations, thus achieving adequate model validation thoroughness while reducing processing time.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent performs preliminary validation of perturbed data before generating test cases. This preliminary action filters out inadequate data early in the process, preventing wasted processing time on low-quality data while maintaining thoroughness in validating models with high-quality test data.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If perturbed data is used without validation, then data processing speed is improved, but evaluation accuracy deteriorates

Engineering Contradiction:
Improvedata processing speedVSAvoidevaluation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary validation of perturbed data before it is used for model evaluation. This preliminary action quickly checks data quality and relevance criteria, enabling fast processing while ensuring that only adequately validated data proceeds to evaluation, thus maintaining both processing speed and evaluation accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a feedback mechanism where validation results of perturbed data influence whether the data is used for evaluation. This feedback loop ensures that data processing speed is maintained through automated validation while evaluation accuracy is preserved by rejecting inadequate data based on validation feedback.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12340285B2Testing models in data pipeline
Publication Date: 2025.06.24 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12340285B2 patent drawing
  • US12340285B2 patent drawing
  • US12340285B2 patent drawing

AI summary

Embodiments of the present invention provide computer-implemented methods, computer program products and computer systems. Embodiments of the present invention can, in response to receiving information, generate a data profile for a model that includes metadata for data requirements, model specific requirements, and data quality metrics. Embodiments of the present invention can generate one or more perturbations for training data associated with the received information and validate at least one perturbation of the one or more perturbations of training data as relevant test data based, at least in part on context associated with the model. Embodiments of the present invention can then generate one or more test scenarios based on the at least one validated perturbation and varying hyperparameters of the model and generate a test report based on an execution of at least one generated test scenario of the generated one or more test scenarios.