Adversarial Autoencoder for ML Data Drift Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning (ML) models face performance degradation when real-world input data shifts from the training data distribution, leading to inaccurate predictions and delayed retraining due to manual and subjective analysis.

Innovation Solution

The use of an Adversarial Autoencoder (AAE) to automatically identify data changes, such as skew, anomalies, and drift, between training and real-world input data, enabling continuous and prompt retraining of ML models with synthesized data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual analysis is used to determine retraining needs, then subjective judgment can be applied, but time consumption and delay increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidretraining delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automatic self-detection of data distribution changes through the AAE model, which continuously monitors input data and identifies when retraining is needed without human intervention, thereby eliminating time delays while maintaining detection accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical analysis with an automated computational system using AAE models that automatically detect data drift and trigger retraining processes, substituting human judgment with algorithmic detection to reduce time loss

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If manual analysis is used to determine retraining needs, then complex subjective evaluation can be performed, but automation level decreases

Engineering Contradiction:
Improveflexibility in judgmentVSAvoidautomated detection
Core Design Contradiction:
Adaptability or versatilityVSExtent of automation

Solution Approach 1:

The AAE model performs automatic self-detection of data distribution changes, enabling the system to autonomously determine when retraining is needed without relying on manual analysis, thus increasing automation while maintaining adaptability through continuous learning

Inventive Principle:
Principle #25Self-service

3Reliability

If retraining is delayed, then resource consumption is reduced, but model accuracy degrades

Engineering Contradiction:
Improvemodel accuracyVSAvoidretraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system implements continuous feedback monitoring through the AAE model that detects data distribution changes in real-time and automatically triggers retraining when accuracy degradation is anticipated, maintaining high model reliability while optimizing retraining timing to improve overall productivity

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12307739B2Image data synthesis using adversarial autoencoders for continual training of models
Publication Date: 2025.05.20 HEWLETT PACKARD ENTERPRISE DEV LP
  • US12307739B2 patent drawing
  • US12307739B2 patent drawing
  • US12307739B2 patent drawing

AI summary

Systems and methods are provided for retraining machine learning (ML) models. Examples may automatically identify skewed, anomalous, and/or drift occurrence data in real-world input data. By automatically identifying such data, examples can reduce subjectivity in ML model retraining as well as reduce time spent determining a need to retrain a ML model. Accordingly, a determination can be made objectively by a computing system or device according to computer-implemented instructions. Additionally, examples may automatically isolate and transfer data relevant to the retraining of a ML model to a training environment for retraining the ML model using real-world input data. Examples also synthesize large samples of data for use in retraining a ML model. The synthesized data may be generated based on the isolated and transferred data and can be used in place of actual real-world input data to reduce a corresponding delay.