Machine Learning Quality Assurance via Training Dataset Distribution Patterns

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning data analysis systems are ill-suited to accurately, efficiently, and consistently perform predictive data analysis, particularly in domains with high-dimensional categorical feature spaces and high cardinality, and face challenges in creating ground truth labels due to subjectivity and resource intensity.

Innovation Solution

The proposed solution involves generating an augmented training dataset by randomly modifying a subset of binary labels based on a probability value, creating a graphical distribution pattern based on accuracy scores, and calculating a quality score by comparing this pattern to a predefined one, thereby improving the quality assurance of machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine learning systems are used for predictive data analysis in high-dimensional categorical feature spaces, then the systems can process data, but they fail to accurately, efficiently, and consistently perform predictive data analysis

Engineering Contradiction:
Improvepredictive accuracyVSAvoidcomputational efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary actions by generating multiple augmented training datasets with randomly modified binary labels before final model training. This pre-processing step creates a distribution pattern that guides subsequent training, improving both accuracy and efficiency by preventing the model from learning incorrect patterns early in the training process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes parameters by randomly modifying binary labels in augmented training datasets based on probability values. This parameter modification creates variations in training data that help the model generalize better to high-dimensional categorical feature spaces, resolving the contradiction between accuracy and efficiency

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If ground truth labels are created manually to ensure accuracy, then labeling quality improves, but the process becomes highly subjective and resource-intensive

Engineering Contradiction:
Improvelabel accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The system applies self-service by automatically generating augmented training datasets with randomly modified labels based on predefined probability distributions. This eliminates the need for manual ground truth creation while maintaining label quality, as the systematic random modification follows statistical principles rather than subjective human judgment

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system creates copies of the training dataset with systematic modifications to binary labels. These augmented copies serve as alternative training sources that reduce reliance on manual labeling while preserving the essential patterns needed for accurate predictions

Inventive Principle:
Principle #26Copying

3Productivity

If more computing resources are allocated to traditional machine learning systems, then processing capacity increases, but accuracy and consistency in high-dimensional categorical feature spaces remain insufficient

Engineering Contradiction:
Improveprocessing capacityVSAvoidprediction consistency
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary distribution pattern analysis on augmented training datasets before final model training. This preliminary action ensures that even with increased computing resources, the model trains on statistically sound data distributions, improving prediction consistency alongside processing capacity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback by comparing distribution patterns of augmented training datasets against expected patterns. This feedback mechanism ensures that increased computational resources are used effectively to maintain both processing capacity and prediction consistency through continuous validation

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250077957A1Quality assurance for machine learning using distribution patterns related to training datasets
Publication Date: 2025.03.06 OPTUM INC
  • US20250077957A1 patent drawing
  • US20250077957A1 patent drawing
  • US20250077957A1 patent drawing

AI summary

Various embodiments of the present disclosure provide quality assurance for machine learning using distribution patterns related to training datasets. In one example, an embodiment provides for generating an augmented training dataset of a plurality of augmented training datasets for a machine learning model by randomly modifying a subset of binary labels of a training dataset for the machine learning model based on a probability value, generating a graphical distribution pattern for the training dataset based on a plurality of accuracy scores of the plurality of augmented training datasets, and generating a quality score for the training dataset based on a comparison between the graphical distribution pattern and a predefined graphical distribution pattern.