AutoML Data Leakage Detection via Subprime Classifier Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated Machine Learning (AutoML) systems face challenges in detecting and addressing data leakage, where information from the training dataset inadvertently influences the test dataset, leading to unreliable model accuracy, especially in complex datasets with thousands of features, where manual detection is impractical.

Innovation Solution

A method involving one-to-many classification, where subprime classifiers are trained with each feature as the target variable, and statistical analysis is performed to identify outlier features contributing to data leakage, which are then removed from the training data to prevent leakage, using techniques like Cook distance, skewness, kurtosis, and One-class Support Vector Machine (SVM) for anomaly detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual detection methods are used to identify data leakage, then detection accuracy may be high, but the process becomes impractical and time-consuming for complex datasets with thousands of features

Engineering Contradiction:
Improvedata leakage detection accuracyVSAvoiddetection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-diagnosis by automatically training subprime classifiers for each feature and using statistical analysis to identify data leakage issues without requiring manual expert intervention. The automated pipeline detects and flags problematic features independently

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the parameter of feature evaluation by training multiple subprime classifiers with different configurations and using statistical metrics (Cook distance, skewness, kurtosis) to identify outliers, transforming the detection process from manual review to automated statistical analysis

Inventive Principle:
Principle #35Parameter changes

2Productivity

If automated classification is used to process thousands of features, then productivity increases, but the complexity of the system increases

Engineering Contradiction:
Improvefeature processing speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex task of data leakage detection into smaller sub-tasks by training individual subprime classifiers for each feature separately. This divides the monolithic detection problem into manageable feature-level analyses that can be processed in parallel

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces statistical analysis metrics (Cook distance, skewness, kurtosis) as intermediaries between the subprime classifier outputs and the final data leakage identification. These statistical measures serve as mediators that translate classifier performance into actionable leakage detection signals

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If subprime classifiers are trained for each feature to detect data leakage, then detection capability improves, but computational resources and training time increase

Engineering Contradiction:
Improvedata leakage detection reliabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system uses partial action by training subprime classifiers only for detection purposes rather than for final prediction. Each subprime classifier is trained minimally to assess feature leakage risk, not to achieve optimal predictive performance, reducing overall computational burden

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11847544B2Preventing data leakage in automated machine learning
Publication Date: 2023.12.19 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11847544B2 patent drawing
  • US11847544B2 patent drawing
  • US11847544B2 patent drawing

AI summary

A mechanism is provided in a data processing system for preventing data leakage in automated machine learning. The mechanism receives a data set comprising a label for a target variable for a classifier machine learning model and a set of features. For each given feature in the set of features, the mechanism trains a subprime classifier model using the given feature as a target variable and remaining features as independent input features, tests the subprime classifier model, and records results of the subprime classifier model. The mechanism performs statistical analysis on the recorded results to identify an outlier result corresponding to an outlier subprime classifier model. The mechanism identifies a outlier feature within the set of features corresponding to the subprime classifier model, removes the identified outlier feature from the set of features to form a modified set of features, and trains the classifier machine learning model using the label for the target variable and the modified set of features.