Industrial internet of things anomaly detection method based on time-frequency domain coordination under full-mlp lightweight architecture
By employing a time-frequency domain collaborative method under a lightweight MLP architecture and a dual-branch reconstruction learning process, combined with the time-frequency domain collaborative method, reconstruction learning is carried out in both the time and frequency domains through the dual-branch reconstruction learning process. This solves the problem of achieving high timeliness and low resource consumption on resource-limited edge devices in existing technologies, thereby improving the accuracy of anomaly detection.
Patent Information
- Application Number
- CN202411506375.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-10-28
AI Technical Summary
Existing industrial IoT anomaly detection models struggle to achieve high timeliness, low resource consumption, and high accuracy on resource-constrained edge devices, and traditional methods are unable to effectively capture the complex characteristics of sensor signals.
We adopt a time-frequency domain collaborative approach under a lightweight MLP architecture. Through a dual-branch reconstruction learning process, we perform reconstruction learning in the time and frequency domains. We use the time-frequency domain reconstruction error to determine anomalies and use four parallel two-layer MLP networks to build a lightweight model.
It achieves high-precision and fast timestamp-level anomaly detection on resource-constrained edge devices, improving the accuracy of anomaly detection while achieving high timeliness and low resource consumption on resource-constrained edge devices.
Smart Images

Figure CN119513551B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of abnormal signal detection of sensors, and particularly relates to an industrial Internet of Things abnormality detection method based on time-frequency domain cooperation of a full MLP lightweight architecture. BACKGROUND
[0002] Industrial Internet of Things plays a vital role in promoting industrial intelligence and intelligent manufacturing. Specifically, it integrates a variety of heterogeneous sensors, such as pressure, vibration and angle sensors, to monitor and control the running state of industrial equipment or industrial production processes. Influenced by the industrial Internet of Things, traditional independent and relatively secure industrial systems or industrial equipment begin to interconnect with each other and interface with external networks, which makes these systems or equipment more vulnerable to external attacks and thus abnormal operation. In addition, the rapid expansion of the industrial Internet of Things also forces more sensors to be deployed to perform more fine-grained management and control of industrial systems and equipment. However, these sensors are often exposed to harsh production environments, such as humidity and darkness, making them vulnerable to external interference and mechanical failure, leading to abnormal events. In such cases, abnormal events often occur in sensors in the industrial Internet of Things, which may hinder intelligent management and maintenance of industrial equipment or production, or even cause equipment damage, property loss or personnel injury. Therefore, anomaly detection in the industrial Internet of Things is crucial and has always been a challenging research hotspot.
[0003] In order to discover abnormal events in the industrial Internet of Things, early methods rely on statistical features of sensor signals (such as distance, variance and density) to build detection models. However, these models are difficult to capture the complex features of sensor signals, resulting in insufficient accuracy. To solve this problem, people began to explore machine learning and deep learning-based models. They use a large amount of historical data sets and expensive manual labels and through supervised or semi-supervised learning to accurately capture the complex features of sensor signals, thereby improving model accuracy. However, abnormal labels in the industrial Internet of Things system are very scarce, and the cost of generating labels is very high, which makes supervised or semi-supervised deep anomaly detection models gradually lose their competitiveness. Therefore, unsupervised deep learning-based anomaly detection models [8] have become the mainstream of research in recent years.
[0004] The sensor signals of the industrial Internet of Things have some new characteristics, which also bring some new challenges to anomaly detection.
[0005] Challenge 1: Difficulty in accurately extracting complex features in sensor signals. With the increasing demand for high-quality production, traditional production processes have become increasingly complex, exhibiting various operating modes such as seasonality, periodicity, and trends. These complex production processes not only result in increasingly complex sensor signals but also require more granular intelligent management and monitoring. To this end, the Industrial Internet of Things requires the deployment of more sensors to monitor production processes more deeply and comprehensively. Therefore, the Industrial Internet of Things sensor signals collected by multiple sensors exhibit some new characteristics: more complex features, more variables, stronger variable coupling, larger data volume, and more scarce labels. These new characteristics make it more difficult to identify abnormal signals from normal sensor signals.
[0006] Challenge 2: Ignoring the high timeliness and low consumption required by models in resource-limited scenarios. Generally, the resources of control and perception devices in industrial scenarios are limited, such as weak computing power and limited storage capacity. However, fine-grained industrial intelligent control requires real-time and accurate monitoring and management of these resource-limited devices. In this context, an efficient anomaly detection model in the Industrial Internet of Things should have high timeliness, low resource consumption, and high accuracy. Only by possessing these three advantages can early and accurate detection of anomalies in resource-efficient industrial scenarios be achieved. Moreover, compared to high precision, high timeliness and low resource consumption are the abilities that an anomaly detection model or method should prioritize or possess. Therefore, an ideal Industrial Internet of Things anomaly detection model should detect sensor signal anomalies as quickly as possible while minimizing the use of computing and storage resources. However, existing models often only focus on pursuing high precision by building deep and large neural networks, ignoring the problem of resource limitation, resulting in poor timeliness and high resource consumption of existing models, and most of them can only be deployed in the cloud and cannot be deployed on resource-limited edge devices.
[0007] To address the first challenge, existing solutions focus on building bloated neural networks with deep structures and huge parameters to enhance the feature capturing ability of the model, thereby improving the model accuracy. For example, the AnomalyTransformer model and the DCdetector model use the Transformer network and the attention network as the feature extractor to extract complex features in the sensor signals. C. Tang et al. employ the gated recurrent unit network (GRU) to accurately capture short-term and long-term temporal dependencies. The DTAAD model integrates the Transformer network and the bidirectional temporal convolution network (TCN) to improve the feature extraction ability. Similarly, the MSt-Gat model and the 2P-DAs model use the graph neural network (GNN) and the generative adversarial network (GAN) to more accurately capture the features of the sensor signals. However, the huge parameters in these models often result in high CPU / GPU resource consumption, and their deep structures significantly increase the training and detection time. This results in the existing models being able to be deployed only on resourceful clouds, but not directly on resource-limited edge devices.
[0008] To address the second challenge, existing solutions are divided into two categories. The first category of solutions uses techniques such as parallel computing and federated learning to accelerate the detection model, thereby improving the timeliness of the model. However, this solution, while achieving high timeliness, requires more computing and storage resources. Therefore, this solution is difficult to deploy on resource-limited industrial Internet of Things edge devices. The second category of solutions uses lightweight components to replace the computationally intensive components in bloated deep models, thereby greatly reducing the model parameters, improving the model timeliness, and reducing the resource consumption of the model. However, these models still have deep structures, and their improvements in timeliness and resource consumption are still unsatisfactory. This results in them also being difficult to deploy and run smoothly on resource-limited edge devices.
[0009] In summary, most existing models are keen on obtaining high accuracy by building bloated neural networks with deep structures and huge parameters. As a result, these models often exhibit poor timeliness and high resource consumption, making them unsuitable for resource-limited edge industrial scenarios. SUMMARY
[0010] To address the above technical problems, the present application provides an industrial Internet of Things anomaly detection method in a full MLP lightweight architecture with time-frequency domain cooperation, which has the advantages of high accuracy, high timeliness, and low resource consumption.
[0011] To solve the above technical problems, the technical solution provided by the present application is:
[0012] An industrial internet of things anomaly detection method in time-frequency domain coordination under a full MLP lightweight architecture, the anomaly detection method constructs a double-branch reconstruction learning process from the perspectives of "local to global" and "global to local"; the reconstruction learning process of each branch utilizes two two-layer MLP networks to perform reconstruction learning in the time domain and the frequency domain respectively, and aligns the time domain and the frequency domain by using reconstruction consistency; the time-frequency domain reconstruction errors obtained in the reconstruction learning of the two branches are fused by using different weights to determine whether each timestamp is abnormal; specifically comprising the following steps:
[0013] Step S1, variable independent sampling, generating local neighbors and global neighbors for each timestamp;
[0014] Step S2, local to global time-frequency domain collaborative reconstruction learning, as the first reconstruction learning branch of the full MLP framework; employing a two-layer MLP network based on real numbers to perform "local to global" reconstruction learning in the time domain; at the same time, employing another two-layer MLP network based on complex numbers to perform "local to global" reconstruction learning in the frequency domain; aligning the time domain and the frequency domain by using the consistency of the global reconstruction values in the time domain and the frequency domain;
[0015] Step S3, global to local time-frequency domain collaborative reconstruction learning, as the second reconstruction learning branch of the full MLP framework, applying a two-layer MLP network based on real numbers to perform "global to local" reconstruction learning in the time domain; at the same time, applying another two-layer MLP network based on complex numbers to perform "global to local" reconstruction learning in the frequency domain; aligning the time domain and the frequency domain by using the consistency of the local reconstruction values in the time domain and the frequency domain;
[0016] Step S4, anomaly scoring based on time-frequency reconstruction error, taking the time-frequency domain reconstruction errors obtained in the double-branch reconstruction learning process as the basis, fusing the two errors by using different weights to generate an anomaly score for each timestamp to determine whether it is abnormal.
[0017] As a further improvement of the above technical solution:
[0018] Preferably, in step S1, the variable independent sampling process of the mth variable at the tth time is:
[0019] S1-1, local neighbor generation, taking a neighbor sequence with a length of L centered on the timestamp t as the local neighbor of the timestamp t;
[0020] S1-2, candidate global neighbor generation, taking a neighbor sequence with a length of G centered on the timestamp t as a candidate global neighbor, wherein L
[0021] S1-3, global neighbor generation, identifying the rest of the candidate neighbor sequence except the local neighbors as global neighbors of timestamp t with length (G-L);
[0022] S1-4, multivariate parallel sampling, the sampling process of multiple variables in the same industrial IoT sensor data set can be executed in parallel.
[0023] Preferably, in step S2, the specific learning process is:
[0024] S2-1, time domain reconstruction learning, using a two-layer MLP network based on real numbers to perform local-to-global reconstruction learning in the time domain; the MLP network includes an input layer with L neurons for receiving the local neighbors of each timestamp, a hidden layer with d neurons and containing ReLU activation function for capturing associated features, and an output layer with (G-L) neurons for reconstructing global neighbors; the input layer and the hidden layer form the first fully connected layer with Lxd trainable parameters; the hidden layer and the output layer form the second fully connected layer with dx(G-L) trainable parameters;
[0025] S2-2, frequency domain reconstruction learning, using another two-layer MLP network based on complex numbers to perform reconstruction learning, which consists of an input layer with L / 2 neurons, a hidden layer with d neurons and containing zReLU activation function, and an output layer with (G-L) / 2 neurons; the input layer and the hidden layer form the first fully connected layer containing (L / 2)xd parameters, and the hidden layer and the output layer form the second fully connected layer containing dx[(G-L) / 2] parameters;
[0026] S2-3, loss function, constructing a loss function consisting of three parts, including time domain reconstruction loss, frequency domain reconstruction loss and time-frequency domain consistency loss;
[0027] The first part is the reconstruction loss of global neighbors in the time domain, which is:
[0028]
[0029] where T is the number of timestamps, M is the number of variables, (G-L) is the length of global neighbors, and are the time domain reconstruction value and the actual value of the ith global neighbor of timestamp t on the mth variable, respectively;
[0030] The second part is the reconstruction loss of global neighbors in the frequency domain, which is:
[0031]
[0032] wherein, is the reconstructed value of the global neighbor in the frequency domain;
[0033] The third part is the time-frequency domain consistency loss, which is used to align the time domain and the frequency domain. The time-frequency domain consistency loss aims to minimize the difference between the reconstructed values in the time domain and the frequency domain, which is:
[0034]
[0035] 1. Preferably, in the time domain reconstruction learning, the local-to-global reconstruction learning process of the mth variable at the tth time stamp is:
[0036] S2-1-1, the local neighbor of length L is input into the first fully connected layer of the first real MLP in the first reconstruction branch to extract the time domain feature, and the extraction formula is:
[0037]
[0038] wherein, is the output of the first fully connected layer, is the training parameter of the first fully connected layer, and ReLU is the activation function;
[0039] S2-1-2, the extracted time domain feature LF t Tim is transmitted to the second fully connected layer of the first real MLP in the first reconstruction branch to reconstruct the global neighbor of the tth time stamp, and the reconstruction formula is:
[0040]
[0041] wherein, represents the global reconstructed value, represents the training parameter of the second fully connected layer.
[0042] 10. Preferably, in the frequency domain reconstruction learning, the local-to-global reconstruction process of the mth variable at the tth time stamp is as follows:
[0043] S2-2-1, the local neighbor of length L is converted into the frequency component of length L / 2 The formula is:
[0044]
[0045] wherein, each frequency component is a complex number,
[0046] S2-2-2, the frequency component LF is the frequency domain feature obtained by inputting into the first fully connected layer of the second complex MLP in the first reconstruction branch t Fre :
[0047]
[0048] wherein, is the output of the first fully connected layer, is the training parameter of the first fully connected layer; zReLU() is a complex activation function, θ z denotes the phase of the complex z;
[0049] S2-2-3, the obtained frequency domain feature LF t Fre Input the second fully connected layer of the second complex MLP in the first reconstruction branch to reconstruct the frequency component of the global neighbor:
[0050]
[0051] wherein, is the output of the second fully connected layer, is the training parameter of the second fully connected layer;
[0052] S2-2-4, apply inverse fast Fourier transform to transform the reconstructed frequency component GF t m back to the global neighbor of the time stamp t in the time domain:
[0053]
[0054] wherein, is the reconstructed value of the global neighbor generated from the frequency domain.
[0055] 11. Preferably, in the step S3, the specific learning process is:
[0056] S3-1, time domain reconstruction learning, using a two-layer MLP network based on real numbers to perform global-to-local reconstruction learning in the time domain, the MLP network is composed of an input layer with (G-L) neurons, a hidden layer with d neurons and an output layer with L neurons; the input layer and the hidden layer form a first fully connected layer containing (G-L) x d training parameters, and the hidden layer and the output layer form a second fully connected layer containing d x L training parameters;
[0057] S3-2, frequency domain reconstruction learning, using another complex-based two-layer MLP network to perform global-to-local reconstruction learning in the frequency domain, the MLP consists of an input layer with (G-L) / 2 neurons, a hidden layer with d neurons, and an output layer with L / 2 neurons; the input layer and the hidden layer form a first fully connected layer containing
[0058] [(G-L) / 2]×d training parameters, and the hidden layer and the output layer form a second fully connected layer containing d×(L / 2) training parameters;
[0059] S3-3, loss function, the loss function consists of three parts, including time domain reconstruction loss, frequency domain reconstruction loss and time-frequency domain consistency loss;
[0060] The time domain reconstruction loss is:
[0061]
[0062] where L is the length of the local neighbor, and is the reconstructed value and the true value of the local neighbor in the time domain; the frequency domain reconstruction loss is:
[0063]
[0064] where, is the reconstructed value of the local neighbor in the frequency domain;
[0065] The time-frequency domain consistency loss is:
[0066]
[0067] 12. Preferably, in the time domain reconstruction learning, the global-to-local reconstruction process of the mth variable at the tth time stamp is:
[0068] S3-1-1, input the global neighbor with length (G-L) to the first fully connected layer of the first real MLP in the second reconstruction branch to extract the time domain features:
[0069]
[0070] where, is the output of the first fully connected layer, is the training parameter of the first fully connected layer, and the ReLU activation function;
[0071] S3-1-2, input GF t Timinput into the first fully connected layer of the first complex MLP in the second reconstruction branch to generate the frequency domain feature GF
[0072]
[0073] wherein, denotes the reconstructed value of the time domain local neighbor, is the training parameter of the first fully connected layer.
[0074] 13. Preferably, in the frequency domain reconstruction learning, the global-to-local reconstruction process of the mth variable at the tth time stamp is:
[0075] S3-2-1, using the fast Fourier transform function to convert the global neighbor of length (G-L) into frequency components of length (G-L) / 2 The formula is:
[0076]
[0077] wherein each frequency component is a complex number,
[0078] S3-2-2, input the obtained frequency components into the first fully connected layer of the second complex MLP in the second reconstruction branch to generate the frequency domain feature GF t Fre The formula is:
[0079]
[0080] wherein, is the output of the first fully connected layer, is the training parameter of the first fully connected layer;
[0081] S3-2-3, input the frequency domain feature GF t Fre into the second fully connected layer of the second complex MLP in the second reconstruction branch to reconstruct the frequency components of the local neighbor The formula is:
[0082]
[0083] wherein, is the output of the second fully connected layer, is the training parameter of the second fully connected layer;
[0084] S3-2-4, using the inverse fast Fourier transform rFFT to generate the reconstructed local neighbor of the time stamp t:
[0085]
[0086] wherein, denotes the reconstructed value of the global neighbor generated from the frequency domain.
[0087] 14. Preferably, in the step S4, let the timestamp to be tested be t, and the anomaly score process be:
[0088] S4-1, the local neighbor and the global neighbor of the timestamp t are simultaneously input into two reconstruction learning branches and the reconstruction learning of M variables is performed in parallel;
[0089] S4-2, each branch will generate three loss values for each timestamp on each variable, and the sum of the three loss values is regarded as the anomaly score of the timestamp t on the variable, ScoreL(M,t) is the score of the timestamp t on the variable M in the first branch, ScoreG(M,t) is the score of the timestamp t on the variable M in the second branch;
[0090] S4-3, the average anomaly scores of the two branches on the M variables are calculated and
[0091] S4-4, the two average anomaly scores are combined with different weights to generate the final anomaly score of the timestamp t, which is:
[0092] Score t = β × ScoreL t + (1-β) × ScoreG t (21)
[0093] wherein, β ∈ [0, 1] is a preset parameter, when β = 0, only the anomaly score of the second reconstruction branch is considered, and when β = 1, only the anomaly score of the first reconstruction branch is considered;
[0094] S4-5, after obtaining the anomaly score of the timestamp t, it is determined whether the tth timestamp is abnormal:
[0095]
[0096] wherein, λ represents a preset threshold value, ranging from 0 to 1; if the tth timestamp is an abnormal timestamp; otherwise, the tth timestamp is a normal timestamp.
[0097] The industrial Internet of Things anomaly detection method under the time-frequency domain coordination of the full MLP lightweight architecture provided by the application has the following advantages compared with the prior art:
[0098] (1) The time-frequency domain cooperative industrial Internet of Things anomaly detection method under the full MLP lightweight architecture of the present application discards the traditional scheme based on deep and large neural networks, and only uses four parallel two-layer MLPs to construct a lightweight full MLP architecture, thereby realizing high-precision, high-timeliness and low-resource consumption timestamp-level anomaly detection. This is the first work of timestamp-level anomaly detection using only a full MLP lightweight architecture design. Experiments on Raspberry Pi 4b and NVIDIA Jetson Xavier NX two edge terminals prove that the LTFAD method can be easily deployed and run in resource-limited industrial scenarios.
[0099] (2) The time-frequency domain cooperative industrial Internet of Things anomaly detection method under the full MLP lightweight architecture of the present application designs a novel dual-branch reconstruction learning framework, uses a simple and shallow MLP network as the backbone, regards "global to local reconstruction learning" and "local to global reconstruction learning" as two independent branches, and simultaneously performs time domain and frequency domain cooperative learning in each branch to improve reconstruction accuracy and enhance anomaly detection accuracy. This learning framework can provide method reference for other fields such as sensor signal analysis or time series prediction. BRIEF DESCRIPTION OF DRAWINGS
[0100] Figure 1 is a general diagram of the detection method of the present application.
[0101] Figure 2 is a variable independent sampling schematic diagram of the present application.
[0102] Figure 3 is an anomaly scoring process diagram based on time-frequency reconstruction error of the present application.
[0103] Figure 4 is a running process screenshot of the detection method of the present application.
[0104] Figure 5 is a detection result diagram of the first variable from 4800 to 5000 timestamps on the MSL data set in the experimental verification of the present application.
[0105] Figure 6 is a deployment result on two industrial Internet of Things edge devices in the experimental verification of the present application.
[0106] Figure 7 is a log file screenshot of the experimental verification of the present application. DETAILED DESCRIPTION
[0107] The specific embodiments of the present application are described in detail below. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.
[0108] The industrial IoT sensor signal is a time series, usually T observations of M heterogeneous sensors on an industrial dynamic system, denoted as The time series contains two dimensions: the time dimension and the variable (spatial) dimension. In the time dimension: the series is denoted as where x t represents the observation from the mth sensor at the tth timestamp. In the variable (spatial) dimension, the series is denoted as where x m represents the observation from the mth sensor at the tth timestamp.
[0109] As shown in Figures 1 to 3 , the industrial IoT anomaly detection method of time-frequency domain coordination under the full MLP lightweight architecture of the present application, first, unlike deep and bloated traditional solutions, the present application designs a shallow lightweight full MLP architecture to realize high timeliness and low resource consumption of anomaly detection. Secondly, based on the lightweight full MLP architecture, a double-branch reconstruction network is constructed to simultaneously perform "global to local" and "local to global" reconstruction learning to improve model accuracy. Finally, time-frequency domain coordination learning is adopted in each reconstruction branch to further improve the accuracy of the model.
[0110] The industrial IoT anomaly detection method of time-frequency domain coordination under the full MLP lightweight architecture of the present application, specifically includes the following steps:
[0111] Step S1, variable independent sampling.
[0112] Fast generation of local neighbors and global neighbors for each timestamp to prepare for subsequent two-branch reconstruction learning. In order to improve speed, variable-independent sampling is performed, that is, the sampling process of each variable is independent and parallel.
[0113] As shown in Figure 2 , the present embodiment designs a simple and fast variable-independent sampling strategy for generating a local neighbor and a global neighbor for each timestamp. Variable-independent sampling means that the sampling process of each variable is independent and parallel.
[0114] Taking the mth variable at the tth timestamp as an example, the sampling process is as shown in Figure 3 . First, a neighbor sequence of length L is taken as its local neighbor centered on timestamp t. Secondly, another neighbor sequence of length G is extracted, where L < G. Finally, the remaining timestamps in the second neighbor sequence except for the local neighbor are identified as the global neighbor of timestamp t , and the length is (G-L).
[0115] Step S2, local-to-global time-frequency domain collaborative reconstruction learning.
[0116] An MLP network based on real numbers is employed to perform "local-to-global" reconstruction learning in the time domain, while another MLP network based on complex numbers is employed to perform "local-to-global" reconstruction learning in the frequency domain, and the consistency of the global reconstruction values in the time domain and the frequency domain is used to align the time domain and the frequency domain, thereby improving the detection accuracy.
[0117] The embodiment designs a timestamp-level anomaly detection model of a shallow and lightweight full-MLP architecture. However, compared with a traditional deep and large neural network, the lightweight full-MLP architecture inevitably has weaker representation ability, and thus the anomaly detection accuracy is reduced. In order to solve this problem, the embodiment improves the accuracy of anomaly detection from two aspects.
[0118] The first aspect is to optimize the design of the full-MLP architecture.
[0119] Generally, normal timestamps in a sensor signal present periodicity, seasonality and other cyclic patterns, while an anomaly is sudden and lasts for multiple timestamps, and the value difference between the anomaly and the normal timestamps is obvious. In this case, the similarity between the local neighbors (close timestamps) and the global neighbors (remote timestamps) of the anomaly timestamps is obviously different from the similarity between the local neighbors and the global neighbors of the normal timestamps. This means that the similarity between the local neighbors and the global neighbors can be used as an evaluation index for discovering anomaly timestamps from normal timestamps, so as to optimize the design of the full-MLP lightweight architecture. Restricted by the shallow full-MLP architecture, the present application selects reconstruction learning as the method for similarity calculation. Based on this, the present application designs a double-branch reconstruction network to learn the similarity between the local and global neighbors from two angles: "local-to-global reconstruction" and "global-to-local reconstruction", and in the learning of each angle, the present application uses one and only one two-layer MLP network to perform reconstruction learning.
[0120] The second aspect is to collaboratively learn the time domain and the frequency domain to enhance the representation ability of the full-MLP architecture.
[0121] Generally, the time domain and the frequency domain are two complementary perspectives for analyzing a time signal. A sensor signal with complex features is difficult to extract accurate features in the time domain, but its features are simple and easy to extract in the frequency domain. Therefore, the present application simultaneously performs the reconstruction learning process of each branch in the time domain and the frequency domain, and uses consistency loss to align the time domain and the frequency domain. Through the mode of collaborative learning of the time-frequency domain, the representation ability of the double-branch reconstruction network based on the full-MLP architecture is enhanced.
[0122] Based on the above idea, the application designs a local-to-global time-frequency domain collaborative reconstruction learning component as the first branch of the double-branch reconstruction network. In this branch, the local neighbors of each timestamp will respectively reconstruct their global neighbors in the time and frequency domains through a two-layer MLP network. The time-frequency domain reconstruction error generated thereby will be used to represent the similarity between the two kinds of neighbors.
[0123] Specifically,
[0124] S2-1, time domain reconstruction learning
[0125] This embodiment only uses a two-layer MLP network to perform local-to-global reconstruction learning in the time domain. The MLP network includes an input layer with L neurons for receiving the local neighbors of each timestamp, a hidden layer with d neurons and a ReLU activation function for capturing associated features, and an output layer with (G-L) neurons for reconstructing the global neighbors. In this MLP network, the input layer and the hidden layer form a first fully connected layer with Lxd trainable parameters. The hidden layer and the output layer form a second fully connected layer with d(G-L) trainable parameters.
[0126] Taking the t-th timestamp of the m-th variable as an example, the local-to-global reconstruction learning process is described as follows:
[0127] S2-1-1, the local neighbor of length L is input into the first fully connected layer of the first real MLP in the first reconstruction branch to extract time domain features, and the extraction formula is: The output of the first fully connected layer is:
[0128]
[0129] wherein, is the output of the first fully connected layer, is the training parameter of the first fully connected layer, and ReLU is the activation function;
[0130] S2-1-2, the extracted time domain feature LF t Tim is input into the second fully connected layer of the first real MLP in the first reconstruction branch to reconstruct the global neighbor of timestamp t, and the reconstruction formula is as follows:
[0131]
[0132] wherein, represents the global reconstruction value, represents the training parameter of the second fully connected layer.
[0133] S2-2, frequency domain reconstruction learning
[0134] Similar to the time-domain reconstruction, the frequency-domain reconstruction learns another complex-based two-layer MLP network to perform the reconstruction learning. Specifically, the MLP network consists of an input layer with L / 2 neurons, a hidden layer with d neurons and zReLU activation function, and an output layer with (G-L) / 2 neurons. In this network, the first fully connected layer contains (L / 2) x d parameters, and the second fully connected layer contains d x [(G-L) / 2] parameters. Different from the time-domain MLP, each training parameter of the frequency-domain MLP is a complex number instead of a real number.
[0135] Taking the t-th timestamp of the m-th variable as an example, the local-to-global reconstruction process is as follows:
[0136] S2-2-1, using the fast Fourier transform to convert the local neighbor of length L into frequency components of length L / 2 The formula is as follows:
[0137]
[0138] where each frequency component is a complex number,
[0139] S2-2-2, inputting the frequency components into the first fully connected layer of the second complex MLP in the first reconstruction branch to obtain the frequency domain feature LF t Fre :
[0140]
[0141] where, is the output of the first fully connected layer, is the training parameter of the first fully connected layer; zReLU() is a complex activation function, θ z represents the phase of the complex number z;
[0142] S2-2-3, inputting the obtained frequency domain feature LF t Fre into the second fully connected layer of the second complex MLP in the first reconstruction branch to reconstruct the frequency components of the global neighbor:
[0143]
[0144] where, is the output of the second fully connected layer, is the training parameter of the second fully connected layer;
[0145] S2-2-4, applying the inverse fast Fourier transform to the reconstructed frequency components GFt m Global neighbors in time domain transformed back to time stamp t:
[0146]
[0147] where, is the reconstructed value of global neighbors generated from frequency domain.
[0148] Further, the reconstruction process of each time stamp within a variable is independent of each other. Similarly, the reconstruction process of M variables is also independent, which can be calculated in parallel.
[0149] S2-3, loss function
[0150] In order to train the two MLP networks in time and frequency domains, the present application constructs a loss function consisting of three parts.
[0151] The first part is the reconstruction loss of global neighbors in time domain, defined as:
[0152]
[0153] where, TL represents the reconstruction loss in time domain, T is the number of time stamps, M is the number of variables, (G-L) is the length of global neighbors, and are the time domain reconstructed value and actual value of the ith global neighbor of time stamp t on the mth variable, respectively.
[0154] The second part is the reconstruction loss of global neighbors in frequency domain, represented as follows:
[0155]
[0156] where, FL represents the reconstruction loss in frequency domain, is the reconstructed value of global neighbors in frequency domain.
[0157] The third part is the time-frequency domain consistency loss, which is used to align the time and frequency domains. This loss aims to minimize the difference between the reconstructed values in time and frequency domains, defined as follows:
[0158]
[0159] where, TFL represents the time-frequency domain consistency loss.
[0160] Step S3, global to local time-frequency domain collaborative reconstruction learning.
[0161] A real-valued two-layer MLP network is employed to perform the "global-to-local" reconstruction learning in the time domain, while another complex-valued two-layer MLP network is employed to perform the "global-to-local" reconstruction learning in the frequency domain, and the consistency of the local reconstruction values in the time domain and the frequency domain is utilized to align the time domain and the frequency domain.
[0162] In this embodiment, a global-to-local time-frequency domain collaborative reconstruction learning component is designed as the second branch of the dual-branch reconstruction network, which utilizes the global neighbors of each timestamp to reconstruct its local neighbors, measures the similarity between the two kinds of neighbors from another perspective, and supplements the first branch.
[0163] Specifically:
[0164] S3-1, time domain reconstruction learning
[0165] A two-layer real-valued MLP network is employed to perform the global-to-local reconstruction learning in the time domain. The MLP network consists of an input layer with (G-L) neurons, a hidden layer with d neurons, and an output layer with L neurons. In this MLP, the first fully connected layer contains (G-L) x d training parameters, and the second fully connected layer contains d x L training parameters.
[0166] Taking the t-th timestamp on the m-th variable as an example, the global-to-local reconstruction process is as follows:
[0167] S3-1-1, the global neighbors of length (G-L) are input into the first fully connected layer of the first real-valued MLP in the second reconstruction branch to extract time domain features:
[0168]
[0169] wherein, is the output of the first fully connected layer, is the training parameter of the first fully connected layer, and the ReLU activation function;
[0170] S3-1-2, the GF t Tim is input into the second fully connected layer of the first real-valued MLP in the second reconstruction branch to reconstruct the local neighbors of the timestamp t:
[0171]
[0172] wherein, represents the reconstruction value of the time domain local neighbor, is the training parameter of the second fully connected layer.
[0173] S3-2, frequency domain reconstruction learning
[0174] Another two-layer complex MLP network is employed to perform the global-to-local reconstruction learning in the frequency domain. The MLP consists of an input layer with (G-L) / 2 neurons, a hidden layer with d neurons, and an output layer with L / 2 neurons. Among them, the first fully connected layer contains [(G-L) / 2]xd training parameters, and the second fully connected layer contains d x (L / 2) training parameters.
[0175] Take the mth variable at the tth timestamp as an example, the global-to-local reconstruction process is as follows:
[0176] S3-2-1, the length of the global neighbor (G-L) is converted to the frequency component of length (G-L) / 2 by using the fast Fourier transform function The formula is:
[0177]
[0178] Each frequency component is a complex number,
[0179] S3-2-2, the obtained frequency component is input into the first fully connected layer of the second complex MLP in the second reconstruction branch to generate the frequency domain feature GF t Fre The formula is:
[0180]
[0181] Among them, is the output of the first fully connected layer, is the training parameter of the first fully connected layer;
[0182] S3-2-3, the frequency domain feature GF t Fre is input into the second fully connected layer of the second complex MLP in the second reconstruction branch to reconstruct the frequency component LF of the local neighbor t m The formula is:
[0183]
[0184] Among them, is the output of the second fully connected layer, is the training parameter of the second fully connected layer;
[0185] S3-2-4, the inverse fast Fourier transform rFFT is used to generate the reconstructed local neighbor of the timestamp t:
[0186]
[0187] wherein, denotes the reconstructed value of global neighbors generated from the frequency domain.
[0188] S3-3, loss function
[0189] The loss function is composed of three parts: time domain reconstruction loss, frequency domain reconstruction loss and time-frequency domain consistency loss.
[0190] The time domain reconstruction loss is:
[0191]
[0192] wherein, TG denotes the time domain reconstruction loss, L is the length of the local neighbors, and are the reconstructed value and the true value of the local neighbors in the time domain.
[0193] The frequency domain reconstruction loss is:
[0194]
[0195] wherein, FG denotes the frequency domain reconstruction loss, is the reconstructed value of the local neighbors in the frequency domain.
[0196] The time-frequency domain consistency loss is defined as follows:
[0197]
[0198] wherein, TFG denotes the time-frequency domain consistency loss.
[0199] Step S4, abnormal score based on time-frequency reconstruction error.
[0200] The present application jointly uses the time-frequency reconstruction error to generate an abnormal score for each timestamp, so as to determine whether it is abnormal.
[0201] The present application designs an individualized abnormal score strategy jointly using the time-frequency reconstruction error to generate an abnormal score for each test timestamp. As shown in Figure 4 , taking the test timestamp t as an example, the detailed abnormal score process is as follows:
[0202] S4-1, the local neighbors and the global neighbors of the timestamp t are simultaneously input into two reconstruction branches to perform reconstruction learning in parallel. The reconstruction learning process of the M variables in the two branches is performed in parallel.
[0203] S4-2, each branch will generate three loss values for each timestamp on each variable, and the three loss values are regarded as the abnormal score of the timestamp t on the variable.
[0204] For example, in the first reconstruction branch, the timestamp t at the Mth variable generates an anomaly score as follows where and denote the local time-domain reconstruction, the frequency-domain local reconstruction, and the time-frequency domain local consistency error, respectively. Similarly, the second reconstruction branch generates an anomaly score for the timestamp t at the Mth variable where and denote the global time-domain reconstruction, the global frequency-domain reconstruction, and the time-frequency domain global consistency error, respectively.
[0205] S4-3, calculate the average anomaly scores of the two branches on the M variables and
[0206] S4-4, merge the two average anomaly scores with different weights to generate the final anomaly score of the timestamp t, defined as follows:
[0207] Score t = β × ScoreL t + (1-β) × ScoreG t (21)
[0208] where β ∈ [0, 1] is a preset parameter. When β = 0, only the anomaly score of the second reconstruction branch is considered; when β = 1, only the anomaly score of the first reconstruction branch is considered.
[0209] S4-5, after obtaining the anomaly score of the timestamp t, determine whether the tth timestamp is abnormal or not using the following formula:
[0210]
[0211] where λ represents a preset threshold value ranging from 0 to 1. If the tth timestamp is an abnormal timestamp; otherwise, the tth timestamp is a normal timestamp.
[0212] Experimental verification
[0213] (1) Experimental data set
[0214] Twelve publicly available sensor signal data sets from industrial IoT environments were selected as experimental data sets, including five single-variable data sets and seven multi-variable data sets, as shown in Table 1. The selected data sets cover multiple industrial fields such as aviation, manufacturing, water treatment, and telemetry, with the goal of ensuring fairness and impartiality in experimental evaluation.
[0215] Table 1 Twelve industrial IoT sensor signal data sets
[0216]
[0217] 1) Multivariate datasets: SKAB is a multi-sensor time series from an industrial testbed; Genesis is a five continuous and thirteen discrete signals from an industrial IoT; MSL is the operational status of multiple sensors or controllers on a Mars rover; PSM is 25-dimensional monitoring information from eBay server machines; SWaT is 51-dimensional data collected by multiple sensors in a public water treatment infrastructure; WADI is 127-dimensional monitoring data from an industrial control system; PUMP is inspection data from multiple sensors on a pump.
[0218] 2) Univariate datasets: Dodgers is traffic flow collected by inductive loop sensors near the Los Angeles Dodgers stadium Data; ECG is electrocardiogram data of an abnormal ventricular premature beat; NAB is a record of the running state of a server cluster; SensorScop e is weather monitoring data from a wireless sensor network; SVDB is 7 semi-hour electrocardiogram recordings with supraventricular arrhythmia anomalies.
[0219] (2) Comparative models and evaluation metrics
[0220] 1) Comparative models
[0221] Nine representative state-of-the-art (SOTA) methods were selected as competitors of the proposed MLP architecture lightweight time-frequency collaborative anomaly detection (LTFAD) method, as shown in Table 2. The above-mentioned comparative methods were selected from the following aspects:
[0222] In terms of model structure, DCdetector, ATF-UAD, TFMAE, PatchAD, DTAAD, and SimAD are deep anomaly detection methods with multiple layers and a large number of parameters. In contrast, DIFFI, COUTA, LODA, and LTFAD are shallow methods.
[0223] In terms of time and frequency domains, DCdetector, DIFFI, COUTA, LODA, PatchAD, DTAAD, and SimAD are anomaly detection methods in the time domain only. In contrast, ATF-UAD, TFMAE, and LTFAD are anomaly detection methods in both time and frequency domains.
[0224] In terms of network backbone, DIFFI, COUTA, and LODA are machine learning-based methods, while the remaining models are neural network-based methods.
[0225] 2) Evaluation metrics
[0226] To comprehensively evaluate the performance of the methods, four widely used evaluation metrics are selected: accuracy (ACC), precision, recall, and F-Score. Among them, accuracy (ACC) is used to measure the proportion of correctly predicted samples to the total number of samples. Precision represents how many of the timestamps predicted as normal are actually normal. Recall describes how many of the actual normal timestamps are correctly predicted. F-Score combines recall and precision to provide a more balanced performance evaluation.
[0227] In general, the range of these metrics is from 0 to 1, and the higher the value represents better performance of anomaly detection.
[0228] Table 2 Comparison of methods
[0229]
[0230]
[0231] (3) Accuracy experiment
[0232] On twelve public industrial IoT datasets, the LTFAD method is compared with nine comparison methods.
[0233] 1) Performance comparison on univariate datasets: Table 3 lists the detection results of ten methods on five univariate datasets. In the table, bold text indicates the best performance, and underlined text indicates the second best performance. From the table, the following findings can be seen:
[0234] ① Overall performance: Compared with other methods, LT-FAD, TFMAE, COUTA, DCdetector, and ATF-UAD show better overall performance. For example, on the SensorScope dataset, LTFAD achieves an ACC of 0.9921 and an F-Score of 0.9815, followed by DCdetector (0.9920 and 0.9813), while SimAD performs the worst (0.3335 and 0.3347).
[0235] ② Stability: The LTFAD method shows the most stable performance among the four indicators, followed by TFMAE, DCdetector, and ATF-UAD. For example, on the ECG, SensorScope, and NAB datasets, LTFAD achieves three first places and one second place.
[0236] ③ Time-frequency domain collaborative learning: The models based on time-frequency domain collaborative learning (LTFAD, TFMAE, and ATF-UAD) have higher detection accuracy than other models based only on time domain. Among the three time-frequency domain collaborative learning models, the LTFAD method of the present invention ranks the highest.
[0237] ④ Lightweight structure: Deep methods (PatchAD, TFMAE, DCdetector, DTAAD, ATF-UAD, SimAD) are generally superior to shallow lightweight methods (LODA, DIFFI, COUTA). However, the lightweight LTFAD method of the present invention outperforms multiple deep methods. For example, on the NAB dataset, the Precision score of LTFAD is 0.9700, the Precision score of DCdetector is 0.9659, the Precision score of TFMAE is 0.9675, the Precision score of COUTA is 0.6254, and the Precision score of DIFFI is only 0.3943.
[0238] 2) Performance comparison on multivariate datasets: Table 4 lists the detection accuracy of ten methods on seven multivariate sensor signals. From the table, similar observations can be made.
[0239] ① Overall performance: The LTFAD method exhibits the best performance, followed by TFMAE, DCdetector, PatchAD, and ATF-UAD. For example, on the SKAB dataset, LTFAD achieves an ACC of 0.9974 and a Recall of 1.0, while COUTA only obtains an ACC of 0.9121 and a Recall of 0.7775.
[0240] ② Stability: LTFAD, TFMAE, DCdetector, and PatchAD exhibit relatively stable accuracy on the four indicators. For example, on the PSM dataset, LTFAD ranks first in three indicators and second in one indicator, while DIFFI and SimAD exhibit high and low instability.
[0241] ③ Time-frequency domain collaborative learning: TFMAE, LTFAD, and ATF-UAD, which employ time-frequency collaborative learning, are superior to DIFFI, PatchAD, and DTAAD, which are pure time domain learning methods. For example, on the WADI dataset, the ACC of LTFAD and TFMAE is 0.9934, the ACC of DTAAD is 0.9533, and the ACC of DIFFI is only 0.6543.
[0242] (4) Lightweight structure: The lightweight LTFAD method outperforms multiple deep and shallow methods. For example, on the PUMP dataset, LTFAD achieves a Precision of 0.9407 and an F-Score of 0.9648, outperforming deep methods such as PatchAD (Precision 0.9311, F-Score 0.9579) and DCdetector (Precision 0.9407, F-Score 0.9606).
[0243] Table 3 Performance analysis on univariate datasets
[0244]
[0245] 3) Visual comparison: Figure 4 The detection results of all methods on the first variable of the MSL dataset from timestamp 4800 to 5000 are presented. In the figure, the black line represents the true data curve, and the red line represents the corresponding label (peak value represents anomaly, valley value represents normal state). From the figure, two findings can be obtained.
[0246] (1) Accuracy: Compared with other methods, COUTA, DCdetector, LTFAD, TFMAE, and LODA have a better match with the true label. For example, these methods can accurately identify anomalies in the range of 48700 to 48950, while other methods produce a large number of errors.
[0247] (2) False positive rate: LTFAD, DCdetector, and LTFAD perform particularly well, with significantly fewer false positives than other methods.
[0248] 4) Conclusion: The above experimental results show that the LTFAD method can accurately identify anomalies in sensor signals in an industrial Internet of Things environment. Specifically, the lightweight LTFAD method of the present application consistently ranks first or second in four indicators on multiple datasets, outperforming multiple deep methods. These findings further demonstrate the effectiveness of time-frequency domain collaborative learning in anomaly detection.
[0249] Table 4 Performance analysis on multivariate datasets
[0250]
[0251] 4) Deployability experiment
[0252] Deployment experiments were conducted on ten methods on two industrial IoT edge devices: Raspberry Pi 4b (less resource) and Jetson Xavier NX (ordinary resource). Specifically, Raspberry Pi 4b is equipped with a 1.5GHz ARM Cortex-A72 processor and 2GB RAM, while Jetson Xavier NX is equipped with a 6-core Carmel ARMv8.2 processor and 8GB RAM. In the experiment, all methods were initially trained in the cloud and then deployed on edge devices for testing. The deployment results are shown in Figure 6 .
[0253] From the figure, two observations can be made:
[0254] ① Deployability: 9 methods were successfully deployed, while PatchAD failed due to storage requirements exceeding the capabilities of both edge devices. In contrast, lightweight methods are easier to deploy and maintain.
[0255] ② Convenience: In the industrial IoT environment, frequent changes in production processes can lead to dynamic fluctuations in time signals. Therefore, methods need to be updated regularly to maintain their optimal performance. However, the update process involves fine-tuning methods using the latest data samples in the cloud and redeploying the fine-tuned methods to edge devices. In this process, lightweight methods consume fewer data samples, have lower communication overhead and labor costs compared to deep methods.
[0256] Summary: Compared with deep methods, shallow lightweight methods are easier to deploy and maintain on resource-limited industrial IoT edge devices.
[0257] 5) Timeliness experiment
[0258] The timeliness and resource consumption of the ten methods were evaluated under three deployment environments, as shown in Figure 5 and Figure 7The environments include a resourceful PC, a less resourceful Raspberry Pi 4b and a generally resourceful Jetson Xavier NX. In this experiment, the PC environment selects model parameters (Parameters), processing time of one epoch with the same batch_size (One-epoch-time), CPU and GPU usage (CPU-Usage and GPU-Usage), RAM usage (RAM-Usage) and VGA RAM usage (VGA-RAM-Usage) as evaluation indexes; the Raspberry Pi 4b environment selects detection time of 100 timestamps (100-timestamps-time), CPU-Usage and RAM-Usage as evaluation indexes; the Jetson Xavier NX environment selects processing time of 10000 timestamps (10000-timestamps-time), CPU-Usage and RAM-Usage as evaluation indexes.
[0259] Table 5 lists the comparison results of the ten methods, where "-" means no available data and "Ot" means out of memory. From the table, several important findings can be obtained:
[0260] ① On training parameters: compared with PatchAD (1230.7K) and DTAAD (123.4K), the LTFAD method has the least training parameters (75K).
[0261] ② On timeliness: in the three deployment environments, the processing time of the shallow lightweight methods (LTFAD, LODA, DIFFI and COUTA) is significantly faster than that of the deep methods (PatchAD, SimAD, TFMAE, DCdetector, ATF-UAD and DTAAD). For example, in the PC environment, the processing time of LTFAD is 37.4 seconds and that of DIFFI is only 8.8 seconds, but that of PatchAD is 461.2 seconds and that of DCdetector is 300.1 seconds.
[0262] ③ On computing consumption: compared with deep methods, the CPU and GPU usage of shallow methods is significantly lower. For example, the CPU usage of shallow DIFFI on Raspberry Pi 4b is 27%, while the CPU usage of deep DCdetector is 72.2% and that of deep ATF-UAD is 68.6%. Comparing these four shallow methods, their CPU usage is similar.
[0263] ④In terms of storage consumption: the storage resources occupied by the deep method are obviously more than those of the shallow method. For example, in the Jetson Xavier NX environment with only 8GRAM, the RAM usage of deep ATF-UAD is 7.5GB, the RAM usage of deep DCdetector is 7.1GB, the RAM usage of the LTFAD of the application is 4.5GB, and the RAM usage of the shallow DIFFI is 3.4GB.
[0264] Summary: The LTFAD method shows good timeliness and low resource consumption, ensuring its applicability to resource-limited industrial Internet of Things edge devices.
[0265] Table 6 Timeliness and resource consumption
[0266]
[0267] The experimental results on twelve public datasets show that the LTFAD method of the application is superior to multiple deep SOTA methods. The experiments on two edge terminals, Raspberry Pi 4b and NVIDIA Jetson Xavier NX, further prove that the LTFAD method can be easily deployed and run in resource-limited industrial scenarios.
[0268] The above implementation cases are only preferred embodiments of the application and do not limit the application in any form. Although the application has been disclosed as above with preferred embodiments, it is not intended to limit the application. Therefore, any simple modification, equivalent change and modification made to the above embodiments without departing from the technical solution of the application, according to the technical essence of the application, shall fall within the protection scope of the technical solution of the application.
Claims
1. An industrial Internet of Things anomaly detection method in a time-frequency domain under a full MLP lightweight architecture, characterized by, The anomaly detection method constructs a double-branch reconstruction learning process from "local to global" and "global to local" two angles; the reconstruction learning process of each branch uses two two-layer MLP networks to perform reconstruction learning in the time domain and the frequency domain respectively, and aligns the time domain and the frequency domain by using reconstruction consistency; The time-frequency domain reconstruction errors obtained in the two-branch reconstruction learning are fused by using different weights to determine whether each timestamp is abnormal; Specifically includes the following steps: Step S1, variable independent sampling, generating local neighbors and global neighbors for each timestamp; Step S2, local-to-global time-frequency domain collaborative reconstruction learning, as the first reconstruction learning branch of the full MLP framework; employ a two-layer MLP network based on real numbers to perform "local-to-global" reconstruction learning in the time domain; at the same time, employ another two-layer MLP network based on complex numbers to perform "local-to-global" reconstruction learning in the frequency domain; align the time domain and the frequency domain by using the consistency of the global reconstruction values in the time domain and the frequency domain; Step S3, global-to-local time-frequency domain collaborative reconstruction learning, as the second reconstruction learning branch of the full MLP framework, apply a two-layer MLP network based on real numbers to perform "global-to-local" reconstruction learning in the time domain; at the same time, apply another two-layer MLP network based on complex numbers to perform "global-to-local" reconstruction learning in the frequency domain; align the time domain and the frequency domain by using the consistency of the local reconstruction values in the time domain and the frequency domain; Step S4, anomaly scoring based on time-frequency reconstruction error, based on the time-frequency domain reconstruction errors obtained in the double-branch reconstruction learning, fuse the two errors by using different weights to generate an anomaly score for each timestamp to determine whether it is abnormal.
2. The method according to claim 1, wherein, In step S1, the variable independent sampling process of the mth variable at the tth time is: S1-1, local neighbor generation, take a length L neighbor sequence centered at timestamp t local neighbor of timestamp t; S1-2, candidate global neighbors generation, extract a neighbor sequence of length G centered at timestamp t as candidate global neighbors, where L < G; S1-3, global neighbor generation, identifying the remaining time stamps in the candidate neighbor sequence, except for the local neighbor, as global neighbors of time stamp t of length (G - L); S1-4, multi-variable parallel sampling, the sampling processes of multiple variables in the same industrial Internet of Things sensor data set are the same and can be executed in parallel.
3. The method according to claim 1, wherein, In step S2, the specific learning process is: S2-1, time domain reconstruction learning, using a two-layer MLP network based on real numbers to perform local-to-global reconstruction learning in the time domain; the MLP network includes an input layer with L neurons for receiving the local neighbors of each timestamp, a hidden layer with d neurons and containing a ReLU activation function for capturing associated features, and an output layer with (G-L) neurons for reconstructing the global neighbors; the input layer and the hidden layer form a first full connection layer with Lxd trainable parameters; the hidden layer and the output layer form a second full connection layer with dx(G-L) trainable parameters; S2-2, frequency domain reconstruction learning, reconstruction learning is performed using another complex-based two-layer MLP network, which consists of an input layer with L / 2 neurons, a hidden layer with d neurons and containing zReLU activation function, and an output layer with (G-L) / 2 neurons; the input layer and the hidden layer form a first fully connected layer containing (L / 2)×d parameters, and the hidden layer and the output layer form a second fully connected layer containing d×[(G-L) / 2] parameters; S2-3, loss function, a loss function consisting of three parts is constructed, including time domain reconstruction loss, frequency domain reconstruction loss and time-frequency domain consistency loss; The first part is the reconstruction loss of the global neighbor in the time domain, which is: where T is the number of timestamps, M is the number of variables, (G-L) is the length of global neighbors, and are the time-domain reconstructed value and the actual value of the i-th global neighbor of the m-th variable at timestamp t, respectively. The second part is the reconstruction loss of the global neighbor in the frequency domain, which is: wherein is the reconstructed value of the global neighbor in the frequency domain; The third part is the time-frequency domain consistency loss, which is used to align the time domain and the frequency domain, and the time-frequency domain consistency loss aims to minimize the difference between the reconstructed values in the time domain and the frequency domain, which is:
4. The method according to claim 3, wherein, In the time domain reconstruction learning, the local to global reconstruction learning process of the mth variable at the tth time stamp is: S2-1-1, a local neighbor of length L The time domain features are extracted into the first fully connected layer of the first real MLP in the first reconstruction branch, and the extraction formula is: wherein, is the output of the first fully connected layer, is the training parameter of the first fully connected layer, and ReLU is the activation function. S2-1-2, reconstructing the extracted time-domain features LF t Tim Reconstruct the global neighbors of the time stamp t in the first real MLP of the first reconstruction branch, and the reconstruction formula is: wherein, represents the global reconstruction value, represents the training parameters of the second fully connected layer.
5. The method according to claim 4, wherein, In the frequency domain reconstruction learning, the local to global reconstruction process of the mth variable at the tth time stamp is as follows: S2-2-1, using a fast Fourier transform to convert the local neighbors of length L into frequency components of length L / 2 S2-2-2, using a fast Fourier transform to convert the local neighbors of length L / 2 into frequency components of length L The formula is: where each frequency component is a complex number, S2-2-2, the frequency component input into a first fully connected layer of a second complex MLP in the first reconstruction branch to obtain a frequency domain feature LF t Fre : wherein, is the output of the first fully connected layer, is the training parameter of the first fully connected layer; zReLU() is a complex activation function, θ z denotes the phase of the complex number z. S2-2-3, the obtained frequency domain features LF t Fre input the second fully connected layer of the second complex MLP in the first reconstruction branch to reconstruct the frequency components of the global neighbors: wherein, is the output of the second fully connected layer, is the training parameter of the second fully connected layer; S2-2-4, applying an inverse fast Fourier transform to the reconstructed frequency components GF t m Global neighbors of the timestamp t in the time domain: wherein, is the reconstructed value of the global neighbor generated from the frequency domain.
6. The method according to claim 3, wherein, In the step S3, the specific learning process is: S3-1, time domain reconstruction learning, global to local reconstruction learning is performed in the time domain using a two-layer MLP network based on real numbers, which consists of an input layer with (G-L) neurons, a hidden layer with d neurons, and an output layer with L neurons; the input layer and the hidden layer form a first fully connected layer containing (G-L)×d training parameters, and the hidden layer and the output layer form a second fully connected layer containing d×L training parameters; S3-2, frequency domain reconstruction learning, global to local reconstruction learning is performed in the frequency domain using another two-layer MLP network based on complex numbers, which consists of an input layer with (G-L) / 2 neurons, a hidden layer with d neurons, and an output layer with L / 2 neurons; the input layer and the hidden layer form a first fully connected layer containing [(G-L) / 2]×d training parameters, and the hidden layer and the output layer form a second fully connected layer containing d×(L / 2) training parameters; S3-3, loss function, the loss function consists of three parts, including time domain reconstruction loss, frequency domain reconstruction loss and time-frequency domain consistency loss; The time domain reconstruction loss is: where L is the length of the local neighbor, and are the reconstructed and true values of the local neighbor in the time domain; and the frequency domain reconstruction loss is: wherein is the reconstructed value of the local neighbor in the frequency domain; The time-frequency domain consistency loss is:
7. The method according to claim 6, wherein, In the time domain reconstruction learning, the global to local reconstruction process of the mth variable at the tth time stamp is: S3-1-1, the global neighbors of length (G-L) are input into the first fully connected layer of the first real MLP to extract the time domain features: input into the first fully connected layer of the first real MLP to extract the time domain features: wherein, is the output of the first fully connected layer, is the training parameter of the first fully connected layer, ReLU activation function; S3-1-2, GF t Tim input to the second fully connected layer of the first real MLP in the second reconstruction branch to reconstruct the local neighbors of the timestamp t: wherein, denotes the reconstructed value of the local neighbor in the time domain, is the second fully connected layer training parameter.
8. The method according to claim 6, wherein, In the frequency domain reconstruction learning, the global to local reconstruction process of the mth variable at the tth time stamp is: S3-2-1, the global neighbors of length (G-L) are converted into frequency components of length (G-L) / 2 using a fast Fourier transform function S3-2-1, the global neighbors of length (G-L) are converted into frequency components of length (G-L) / 2 using a fast Fourier transform function S3-2-1, the global neighbors of length (G-L) are converted into frequency components of length (G-L) / 2 using a fast Fourier transform function where each frequency component is a complex number, S3-2-2, the obtained frequency component input into a first fully connected layer of a second complex MLP in the second reconstruction branch to generate a frequency domain feature GF t Fre , the formula is: wherein, is the output of the first fully connected layer, is the training parameter of the first fully connected layer; S3-2-3, the frequency domain feature GF t Fre inputting the frequency components LF of the reconstructed local neighbors into a second fully connected layer of a second complex MLP in the second reconstruction branch t m , the formula is: wherein, is the output of the second fully connected layer, is the second fully connected layer training parameter; S3-2-4, inverse fast Fourier transform rFFT is used to generate the reconstructed local neighbor of the tth time stamp: wherein, denotes the reconstructed value of the global neighbor generated from the frequency domain.
9. The method according to claim 8, wherein, In the step S4, the time stamp to be tested is t, and the abnormal score process is: S4-1, the local neighbor and the global neighbor of the tth time stamp are simultaneously input into two reconstruction learning branches and parallel reconstruction learning of M variables is performed; S4-2, each branch will generate three loss values for each timestamp on each variable, the sum of the three loss values is considered as the abnormal score of timestamp t on this variable, S4-1, the score of timestamp t on variable M in the first branch, S4-2, the score of timestamp t on variable M in the second branch; S4-3, compute average abnormal scores of the two branches on M variables and S4-4, the two average anomaly scores are combined with different weights to generate a final anomaly score of the timestamp t, which is: Score t = β x ScoreL t + (1 - β) x ScoreG t (21) wherein β is a preset parameter in [0, 1], when β = 0, only the anomaly score of the second reconstruction branch is considered, and when β = 1, only the anomaly score of the first reconstruction branch is considered; S4-5, after obtaining the anomaly score of the timestamp t, it is determined whether the t-th timestamp is abnormal: wherein λ represents a preset threshold value ranging from 0 to 1; if The tth time stamp is an abnormal time stamp; otherwise, the tth time stamp is a normal time stamp.
Citation Information
Patent Citations
Method for detecting and positioning abnormal behavior of target in key area based on multi-domain information fusion
CN115147921A
Method for detecting abnormal sound of self-supervised machine based on domain transfer
CN115376554A