Synthetic Data Generation for Fault Classification Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in training machine learning models for fault classification in industrial automation is the scarcity of labeled training data, which is time-consuming and impractical to gather manually.
Innovation Solution
Utilizing generative artificial intelligence to synthesize labeled training data based on a small amount of initial labeled data, including lab measurements, simulation data, and domain knowledge, to train machine learning models for fault classification in industrial automation devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data labeling is used to train machine learning models for fault classification, then the model accuracy can be improved with sufficient labeled data, but the time consumption and resource burden increase significantly
Solution Approach 1:
The patent uses generative AI models to create synthetic copies of labeled fault data. The generative model learns from a small set of real labeled data and generates numerous synthetic fault samples that replicate the characteristics of real fault data, thereby obtaining large amounts of labeled training data without manual annotation.
Solution Approach 2:
The patent replaces expensive and time-consuming manual data labeling with automated synthetic data generation. The synthetic data acts as a disposable substitute that can be generated efficiently at scale, eliminating the need for costly expert annotation while providing sufficient training data for model accuracy.
2Quantity of substance
If extensive manual data labeling is performed, then sufficient training data can be gathered, but the burden on experts and data scientists increases
Solution Approach 1:
The system enables self-service data generation where the generative AI model automatically creates synthetic labeled data without requiring expert intervention. The model serves itself by learning from initial labeled data and autonomously generating additional training data, eliminating the need for experts to manually annotate each sample.
Solution Approach 2:
The generative model creates synthetic copies of fault data that replicate the essential characteristics of real fault data. These synthetic copies provide the necessary quantity of labeled data while requiring no expert review or annotation, as the generation process is fully automated.
3Measurement precision
If real fault data is collected from industrial automation devices, then accurate training data is obtained, but the devices experience downtime and operational disruption
Solution Approach 1:
The patent generates synthetic fault data that replicates the characteristics of real fault data without requiring actual fault occurrence. The generative model creates virtual fault samples that maintain statistical properties and diagnostic value while avoiding the need to disrupt industrial automation device operation.
Solution Approach 2:
The synthetic data serves as a disposable alternative to real fault data collection. Instead of disrupting devices to capture fault conditions, the system uses computationally generated data that provides equivalent training value without the operational costs and downtime associated with real fault monitoring.
Data Source
AI summary
A method including generating synthetic labelled training data based on a set of labelled sensor data associated with one or more industrial automation devices; and training, based at least on the synthetic labelled training data, one or more machine learning models for fault classification associated with the one or more industrial automation devices.


