Training Data Filtering for Accurate Anomaly Detection Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine training models struggle to accurately detect anomaly values from operation data due to the inclusion of anomaly data during training, necessitating the generation of training data that includes only normal data.
Innovation Solution
A training data generation apparatus that acquires operation data, divides it into temporary training and test data, detects anomaly values, and generates training data by excluding anomaly periods, ensuring the final data set consists only of normal data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If operation data including anomaly data is used as training data, then the machine training model can be trained with sufficient data, but the model fails to accurately detect anomaly values
Solution Approach 1:
The patent extracts and removes anomaly data from the operation data before generating training data. The anomaly detection unit identifies anomaly periods in the operation data, and the training data generation unit excludes these anomaly periods to create clean training data. This extraction of harmful anomaly data resolves the contradiction by maintaining sufficient training data quantity while eliminating the negative impact of anomaly data on detection accuracy.
Solution Approach 2:
The patent introduces an intermediary process (anomaly detection and filtering mechanism) between the raw operation data and the training data. The temporary training data and temporary test data are generated as intermediaries to detect anomaly periods, which then inform the creation of final clean training data. This intermediary process ensures that only normal operation data is included in the training set, resolving the contradiction between data quantity and detection accuracy.
2Measurement precision
If anomaly data is excluded from training data, then anomaly detection accuracy is improved, but the amount of available training data decreases
Solution Approach 1:
The patent segments the operation data into multiple temporary training data sets and temporary test data sets. By dividing the data and performing anomaly detection on each segment, the system can identify and exclude only the anomaly periods while retaining the majority of normal operation data. This segmentation approach maintains sufficient training data quantity while ensuring high detection accuracy by removing only the harmful anomaly portions.
Solution Approach 2:
The patent performs preliminary anomaly detection on temporary training data and temporary test data before generating the final training data. This preliminary action identifies anomaly periods in advance, allowing the training data generation unit to exclude only those specific anomaly periods. This approach ensures that the final training data contains only normal operation data while maintaining sufficient quantity for effective model training.
Data Source
AI summary
According to one embodiment, a training data generation apparatus includes a processor. The processor acquires operation data related to an operation state of a device in a predetermined period. The processor divides the operation data into at least temporary training data and temporary test data. The processor detects an anomaly value from the temporary test data based on the temporary training data. The processor generates training data from the operation data by excluding an anomaly period in which the anomaly value is detected from the predetermined period.


