Industrial internet time series data anomaly detection method based on WEAGNN

By combining TMSMOTE data enhancement and wavelet decomposition with graph neural networks, the dataset imbalance and feature capture problems of traditional models in time series data anomaly detection are solved, achieving higher detection accuracy and generalization capability, which is suitable for time series data anomaly detection in the Industrial Internet.

CN120670832APending Publication Date: 2025-09-19GUILIN UNIV OF ELECTRONIC TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510765254.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional deep learning models find it difficult to simultaneously capture the features of different time scales and the complex relationships between different indicators in time series data anomaly detection, resulting in problems of missed reports and false positives in anomaly detection results, especially in the industrial Internet where data sets are severely unbalanced.

Method used

TMSMOTE is used for data enhancement, combined with wavelet decomposition and graph neural network (WEAGNN) model, and spatial dependencies are modeled through graph convolution technology and graph attention mechanism. At the same time, a contrastive learning mechanism is introduced to model positive and negative sample patterns to capture local mutations and global trend characteristics of time series data.

Benefits of technology

It improves the accuracy and generalization ability of anomaly detection, effectively copes with complex time series data anomaly detection tasks, and reduces missed reports and false positives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670832A_ABST
    Figure CN120670832A_ABST
Patent Text Reader

Abstract

The invention relates to the field of time series data anomaly detection, deep learning, graph neural network, wavelet decomposition and industrial internet protection, in particular to an industrial internet time series data anomaly detection method based on WEAGNN. According to the method, wavelet decomposition and data enhancement are combined to solve the problem of imbalance of a time series data set, data time dimension features are captured at the same time, and then local mutation features and global trend features of time series data are captured from a spatial dimension by applying technologies such as a graph neural network and comparative learning. Therefore, higher anomaly detection accuracy and stronger generalization ability are provided, and complex time series data anomaly detection tasks are effectively coped with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of time series data anomaly detection, deep learning, graph neural networks, wavelet decomposition and industrial Internet protection, and specifically to a method for detecting anomaly in industrial Internet time series data based on WEAGNN. Background Art

[0002] Time series data anomaly detection is a process of analyzing chronologically arranged data sequences to identify abnormal points or patterns that do not conform to normal patterns or rules. It is widely used in industrial equipment fault diagnosis, financial risk warning, network security monitoring and other fields to timely discover potential problems and take corresponding measures.

[0003] In the era of the Industrial Internet, the massive amounts of data generated by systems contain critical operational information, such as equipment failures and anomalies. With the advancement of intelligent technologies and the increasing complexity of devices, the security challenges facing the Industrial Internet are becoming increasingly severe. Against this backdrop of surging data volumes, time series data anomaly detection technology is particularly important in the Industrial Internet. Time series data encompasses a wide range of aspects of industrial processes, including temperature, pressure, network traffic, and medical monitoring data. It contains critical information about system status, behavioral patterns, and potential risks. Time series data anomaly detection technology can promptly detect abnormal patterns, provide early warnings of potential failures, and prevent losses, making it crucial for ensuring stable system operation. Despite this, this technology still faces numerous challenges in practical industrial applications.

[0004] In the current field of time series data anomaly detection, a particularly prominent issue is dataset imbalance. This imbalance means that the number of normal data samples far outnumbers abnormal samples, resulting in poor performance of detection models when faced with abnormal samples. This phenomenon is particularly prominent in the complex industrial internet, where the diversity of devices and protocols leads to large data discrepancies and a scarcity of abnormal samples.

[0005] In recent years, deep learning technology has become a powerful tool for automatically detecting anomalous or malicious activity, leveraging its ability to exploit inherent patterns and relationships in data to distinguish normal from abnormal behavior. However, traditional deep learning techniques also have limitations. In anomaly detection in time series data, two types of features are typically present: short-term data fluctuations and anomalies in long-term trends. Capturing both of these features simultaneously is crucial in the field of anomaly detection.

[0006] Furthermore, time series data contains complex interdependencies between different indicators. Convolutional neural networks are widely used to capture local spatial relationships in time series data, but they struggle to model the complex interrelationships between multiple variables. On the other hand, while recurrent neural networks can capture historical information about time series, they perform poorly when processing long time series. While the Transformer architecture, with its self-attention mechanism, is better able to handle long-range dependencies, this can also lead to the loss of information about short-term data mutations.

[0007] In summary, traditional deep learning models often struggle to simultaneously account for features at different time scales, and fail to effectively capture the complex interrelationships between different metrics. This ultimately leads to serious omissions and false positives in anomaly detection results. Summary of the Invention

[0008] The purpose of this paper is to provide a WEAGNN-based anomaly detection method for industrial Internet time series data. This method aims to address the imbalanced dataset problem in the field of time series data anomaly detection and the serious omissions and false positives in traditional anomaly detection results.

[0009] To achieve the above objectives, the present invention provides an industrial Internet time series data anomaly detection method based on WEAGNN, comprising the following steps:

[0010] TMSMOTE is used to perform data augmentation on imbalanced datasets.

[0011] The enhanced data is divided into training set and test set and wavelet decomposition is performed on each set to extract the temporal characteristics of the data.

[0012] The training dataset is input into the EAGNN architecture to apply graph convolution technology and graph attention mechanism to model spatial dependencies, while the contrastive learning mechanism is introduced to model positive and negative sample patterns, and finally a trained WEAGNN model is obtained.

[0013] The above test data set is input into the trained WEAGNN model to obtain the anomaly detection result R.

[0014] The present invention provides an industrial Internet time series data anomaly detection method based on WEAGNN. By combining wavelet decomposition, data enhancement, graph neural network and contrastive learning technologies, it solves the imbalance problem of time series data sets, and simultaneously captures the local mutation characteristics and global trend characteristics of time series data from the two dimensions of time and space, thereby providing higher anomaly detection accuracy and stronger generalization ability, and effectively coping with complex time series data anomaly detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 A flowchart of a method for detecting anomalies in industrial Internet time series data based on WEAGNN is provided in an embodiment of the present invention.

[0017] Figure 2 It is a flow chart of the solution of the present invention.

[0018] Figure 3 This is a comparison chart of the Accuracy, Precision, Recall and F1 score experimental results of each method in the specific embodiment of the present invention. DETAILED DESCRIPTION

[0019] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0020] See also Figure 1 , the present invention proposes an industrial Internet time series data anomaly detection method based on WEAGNN, which includes the following steps:

[0021] S1: Use TMSMOTE to perform data augmentation on the imbalanced dataset;

[0022] S2: Divide the enhanced data into training and test sets and perform wavelet decomposition on each set to extract the temporal characteristics of the data;

[0023] S3: The training dataset is input into the EAGNN architecture to apply graph convolution technology and graph attention mechanism to model spatial dependencies. At the same time, a contrastive learning mechanism is introduced to model positive and negative sample patterns, and finally a trained WEAGNN model is obtained.

[0024] S4: Input the above test data set into the trained WEAGNN model to obtain the anomaly detection result R.

[0025] Specifically, in step S1, the process of using TMSMOTE to perform data enhancement on an imbalanced dataset includes the following steps:

[0026] First, TMSMOTE divides time series data into fixed-length segments through a sliding window to capture local temporal features.

[0027] Secondly, Dynamic Time Warping (DTW) is used to calculate the minority class window similarity to capture the temporal dynamic patterns.

[0028] Then, principal component analysis (PCA) was performed to model multivariate dependencies, reduce dimensionality and retain the main features.

[0029] Then, synthetic minority samples are generated by interpolation in the PCA space to enhance the diversity of the dataset.

[0030] Then, a moving average filter is used to smooth the synthesized samples to ensure the continuity of the time dimension.

[0031] Finally, the generated window and the original data are integrated to form a balanced time series dataset.

[0032] In step S2, the enhanced data is divided into a training set and a test set and wavelet decomposition is performed on each set to extract the temporal characteristics of the data. The process includes the following steps:

[0033] First, for the input time series data set X={X1, X2, ..., X n}, for each time series data X i Perform wavelet decomposition to generate low-frequency components L i and high frequency components H i .

[0034] Among them, the low-frequency component L i Reflects long-term trend characteristics and captures normal operating status; high-frequency component H i Reflect short-term fluctuation characteristics and detect abnormal mutations and noise.

[0035] Secondly, extract low-frequency and high-frequency features respectively and Merge into time features

[0036] Finally, a temporal feature set T is formed to distinguish normal and abnormal signals and enhance the understanding of the intrinsic structure of time series data.

[0037] In step S3, the training dataset is input into the EAGNN architecture to apply graph convolution technology and graph attention mechanism to model spatial dependencies, and at the same time introduce contrastive learning mechanism to model positive and negative sample patterns. The process of finally obtaining the trained WEAGNN model mainly includes two operations: one is to apply graph convolution technology and graph attention mechanism to model spatial dependencies, and the other is to introduce contrastive learning mechanism to model positive and negative sample patterns.

[0038] Modeling spatial dependencies involves the following steps:

[0039] First, for the time feature set T = {T1, T2, ..., T m} and the graph structure G built based on data similarity, initialize GCN and GAT parameters θ GCN and θ GAT .

[0040] Secondly, T is processed by GCN i , extract local dependency features Capture local spatial relationships.

[0041] Then, adaptive learning rate is used to optimize GAT parameters and GAT is used to extract global abnormal features.

[0042] Then, the attention mechanism is used to enhance the focus on important features. and Merge into multi-level feature representation

[0043] Finally, the feature set H is generated to realize spatial feature extraction from local to global.

[0044] The introduction of contrastive learning mechanism to model positive and negative sample patterns includes the following steps:

[0045] First, for the multi-level feature set H = {H1, H2, ..., H n}、Positive sample pair set P and negative sample pair set N, initialize the contrast loss function L contrastive .

[0046] Secondly, for the positive sample (H i , H j ) to calculate the feature similarity sim(H i , H j ), update L contrastive To maximize the similarity of positive samples.

[0047] Then, for the negative samples (H i , H k ) to calculate the feature similarity sim(H i , H k ), update L contrastive To minimize the negative sample similarity.

[0048] Finally, by optimizing L contrastive , accurately model the normal and abnormal patterns of samples and obtain the trained model WEAGNN.

[0049] In step S4: the test data set divided in step S2 is input into the trained WEAGNN model to obtain the anomaly detection result R.

[0050] The following is further explained in conjunction with specific embodiments and execution processes:

[0051] 1. Experimental Data

[0052] In this specific embodiment, the dataset used is the public anomaly detection dataset WADI. The WADI dataset contains 172,801 data points, 9,977 of which are anomalies. The dataset also includes 127 features. To facilitate simulation experiments, 138,241 data points from the WADI dataset were selected as the training set, and 34,560 data points were selected as the test set.

[0053] 2. Implementation process

[0054] The implementation process of the anomaly detection method based on WEAGNN (pseudo-code algorithm) is shown in the following table:

[0055]

[0056] In this pseudocode algorithm, the symbol X train represents the training data set, Y test represents the test data set, R represents Test results, n represents the training epoch. The implementation process is as follows:

[0057] Algorithm pseudocode line 1: Initialize all parameters and datasets.

[0058] Algorithm pseudocode line 3: X train Input into the WEAGNN time series data anomaly detection model for modeling.

[0059] Algorithm pseudocode line 5: Y test Input into the WEAGNN model and get the result R;

[0060] Furthermore, the present invention compares the experimental results with some existing time series data anomaly detection methods. Specifically, the experimental results of the present invention's WEAGNN-based anomaly detection method are compared with those of anomaly detection methods based on IMDiffusion, gDN, CAT, NPSR, CST-GL, and FourierGNN.

[0061] The specific evaluation indicators are as follows:

[0062] Accuracy, Precision, Recall, and F1 score are used as experimental evaluation criteria to verify the effectiveness of our proposed WEAGNN-based anomaly detection method. The evaluation criteria are defined as follows.

[0063] True positive (TP): The actual case is positive and the test result is also positive.

[0064] True negative (TN): The actual condition is negative and the test result is also negative.

[0065] False positive (FP): The test result is positive when the actual result is negative.

[0066] False negative (FN): The test result is negative even though the test result is actually positive.

[0067] Accuracy: The percentage of correctly detected samples in the total samples. The formula is as follows:

[0068]

[0069] Precision: The percentage of samples that are positive and the true value is also positive. The formula is as follows:

[0070]

[0071] Recall: The proportion of positive examples in a sample that are correctly detected. Its formula can be described as follows:

[0072]

[0073] F1_score: Combines the results of Precision and Recall to comprehensively evaluate the results of anomaly detection The price is as follows:

[0074]

[0075] For simulation results, see Figure 2 ,This specific embodiment tests the indicators such as Accuracy, Precision, Recall and F1_score through simulation experiments. It is not difficult to see that the present invention is superior to the other six existing methods in terms of Accuracy, Precision, Recall and F1_score, and can therefore provide more effective time series data anomaly detection.

[0076] In summary, the present invention has obvious advantages over the anomaly detection method based on IMDiffusion, the anomaly detection method based on GDN, the anomaly detection method based on CAT, the anomaly detection method based on NPSR, the anomaly detection method based on CST-GL, and the anomaly detection method based on FourierGNN. This is because the anomaly detection method based on WEAGNN can simultaneously consider the local and long-term characteristics of the time and spatial dimensions of time series data compared to the other six methods.

[0077] The above disclosure is only a preferred embodiment of the present invention, and certainly cannot be used to limit the scope of the rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.

Claims

1. A method for detecting anomaly in industrial Internet time series data based on WEAGNN, characterized in that: The following steps are involved: TMSMOTE is used to enhance the data of the imbalanced dataset. The enhanced data is divided into training set and test set and wavelet decomposition is performed on each set to extract the temporal characteristics of the data. The training dataset is input into the EAGNN architecture to apply graph convolution technology and graph attention mechanism to model spatial dependencies, while introducing contrastive learning mechanism to model positive and negative sample patterns, and finally obtain the trained WEAGNN model. The above test data set is input into the trained WEAGNN model to obtain the anomaly detection result R.

2. The industrial Internet time series data anomaly detection method based on WEAGNN according to claim 1 is characterized in that: The implementation process of the WEAGNN-based anomaly detection model includes the following steps: TMSMOTE uses a sliding window to split time series data into fixed-length segments to capture local temporal features; Dynamic Time Warping (DTW) is used to calculate the similarity of minority class windows to capture temporal dynamic patterns; Modeling multivariate dependencies through principal component analysis (PCA), reducing dimensionality and retaining key features; Generate synthetic minority samples by interpolation in PCA space to enhance the diversity of the dataset; Moving average filtering is used to smooth the synthetic samples to ensure continuity in the time dimension; Integrate the generated window and the original data to form a balanced time series dataset.

3. The industrial Internet time series data anomaly detection method based on WEAGNN according to claim 2 is characterized in that: For the input time series data set X={X1, X2, ..., X n }, for each time series data X i Perform wavelet decomposition to generate low-frequency components L i and high frequency components H i ; Low frequency component L i Reflects long-term trend characteristics and captures normal operating status; high-frequency component H i Reflect short-term fluctuation characteristics and detect abnormal mutations and noise; Extract low-frequency and high-frequency features separately and Merge into time features Form a time feature set T to distinguish normal and abnormal signals and enhance the understanding of the intrinsic structure of time series data.

4. The industrial Internet time series data anomaly detection method based on WEAGNN according to claim 2 is characterized in that: For the time feature set T = {T1, T2, ..., T m } and the graph structure G built based on data similarity, initialize GCN and GAT parameters θ GCN and θ GAT ; Processing T through GCN i , extract local dependency features Capturing local spatial relationships; Adaptive learning rate is used to optimize GAT parameters and GAT is used to extract global abnormal features By using the attention mechanism to enhance the focus on important features, and Merge into multi-level feature representation Generate feature set H to realize spatial feature extraction from local to global, For the multi-level feature set H = {H1, H2, ..., H n }、Positive sample pair set P and negative sample pair set N, initialize the contrast loss function L contrastive ; For positive samples (H i , H j ) to calculate the feature similarity sim(H i , H j ), update L contrastive To maximize the similarity of positive samples; For negative samples (H i , H k ) to calculate the feature similarity sim(h i ,H k ), update L contrastive To minimize the similarity of negative samples; By optimizing L contrastive , accurately model the normal and abnormal patterns of samples and obtain the trained model WEAGNN.