Residual Data Augmentation for Neural Network Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The scarcity of labeled data in industrial domains, particularly for anomaly detection, leads to limitations in training deep neural networks, as collecting and labeling anomalous data is costly and time-consuming, and simulation data often fails to accurately represent real-world scenarios, resulting in models that may overfit or fail to capture high-order feature interactions.
Innovation Solution
A method for generating realistic data by augmenting simulation data with new samples based on residual data between real and simulated processes, using techniques such as Short-Time Fourier Transformation, Principal Component Analysis, multivariate Gaussian methods, Kernel Density Estimation, Variational Auto-Encoders, and Generative Adversarial Networks to create a distribution of discrepancies, enabling the training of more powerful data evaluation models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If simulation data is used for training anomaly detection models, then data availability is improved, but the reality gap causes poor performance on real data
Solution Approach 1:
The patent introduces residual data as an intermediary component that bridges simulation data and real data. The residual data captures the discrepancy between simulated and real processes, and when added to simulation data, creates augmented training data that maintains both abundance and realism, thereby resolving the reality gap problem
Solution Approach 2:
The patent creates a composite training data structure by combining simulation data and residual data into augmented training data. This composite approach leverages the advantages of both data sources: the quantity and controllability of simulation data, and the realism and authenticity of residual data, resulting in training data that is both abundant and representative of real-world scenarios
2Reliability
If real anomalous data is collected for training, then data quality is improved, but cost and time consumption increase significantly
Solution Approach 1:
The patent creates copies of real data characteristics through residual data extraction. Instead of collecting and labeling extensive real anomalous data, the method captures the essential discrepancy patterns from limited real data and applies them to generate synthetic residual data, thereby replicating real data quality without the associated collection and labeling costs
Solution Approach 2:
The system uses the limited real data available to automatically generate residual data that captures real-world characteristics. This self-service approach allows the method to extract and utilize real data patterns without requiring extensive manual annotation or collection efforts, reducing both time and resource investment
3Reliability
If simpler anomaly classification models are used, then overfitting is reduced, but the ability to capture high-order feature interactions is limited
Solution Approach 1:
The patent changes the parameter of data quality and diversity by introducing residual data augmentation. This transformation of the training data parameters enables complex models to be trained effectively without overfitting, as the residual data provides realistic variations and patterns that improve generalization while maintaining the model's capacity to capture high-order feature interactions
Data Source
AI summary
Various embodiments of the teachings herein include a computer implemented sample preparation method for generating a new sample of data for augmenting simulation data to generate realistic data to be applied for training of a data evaluation model. The method may include generating the new sample based on an output data set sampled from a model of an input data set based on residual data. The residual data are based on real data of a real process and simulated data of a simulated process corresponding to the real process.


