A method, an electronic device, and a storage medium for detecting anomalies in time series data

Features are extracted through sliding windows and convolution modules, combined with time and space modules to capture the dependence of multi-dimensional time series, and used MLP and AE models to generate detection values, solving the problem of existing methods ignoring time dependence and dynamic dependence, and achieving more accurate timing data anomaly detection.

CN117520982BActive Publication Date: 2025-06-20CIVIL AVIATION UNIV OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311468133.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-06-20
Estimated Expiration
2043-11-06

AI Technical Summary

Technical Problem

The existing multidimensional time series data anomaly detection method based on graph neural network ignores short-term and long-term time dependencies, and the fixed graph structure cannot dynamically capture the potential dependencies between different variables, resulting in the model being unable to fully explore the spatial characteristics of global and local dynamics.

Method used

The time series data is divided by sliding window, and features are extracted through the convolution module, combining the time module and the space module to capture the short-term and long-term time dependencies and dynamic local spatial dependencies of the multi-dimensional time series. Use the MLP prediction model and AE reconstruction model to generate detection values ​​to determine data abnormalities.

Benefits of technology

It realizes a more comprehensive detection of timing anomalies, which can effectively capture the short-term and long-term time dependencies and dynamic spatial dependencies of multi-dimensional time series, and improves the accuracy of anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117520982B_ABST
    Figure CN117520982B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing, and provides a method for detecting anomalies in time series data, an electronic device, and a storage medium, including: obtaining a time series data set D to be detected; obtaining a sliding time series data set DS based on D; inputting DSh into a first convolutional module to obtain a first convolutional feature Fh1<supgt;c< / supgt>; inputting Fh1<supgt;c< / supgt> into a first spatio-temporal module to obtain a first spatio-temporal feature Fh1<supgt;st< / supgt>; obtaining a second spatio-temporal feature Fh2<supgt;st< / supgt> based on Fh1<supgt;st< / supgt> and Fh1<supgt;c< / supgt>; obtaining a second convolutional feature Fh2<supgt;c< / supgt> based on Fh1<supgt;st< / supgt> and Fh2<supgt;st< / supgt>; obtaining a predicted data Dh<supgt;nc< / supgt> and a reconstructed data Dh<supgt;nr< / supgt> for the next moment corresponding to DSh based on Fh2<supgt;c< / supgt>; obtaining a detection value Sh<supgt;next< / supgt> corresponding to the next moment of DSh based on Dh<supgt;next< / supgt>, Dh<supgt;nc< / supgt>, and Dh<supgt;nr< / supgt>. If Sh<supgt;next< / supgt> > S0, then determine that Dh<supgt;next< / supgt> is abnormal data. The present invention can accurately detect anomalies in time series data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a method for detecting anomalies in time series data, an electronic device, and a storage medium. Background Art

[0002] In a cyber-physical system, sensors are usually used to monitor the operating state of the system. With the rapid increase in the number of sensors, a large amount of multivariate time series (MTS) data is generated. For example, in a water treatment system, different sensors are used to collect monitoring data such as water level, flow rate, water pressure, and valve status. By performing efficient and accurate anomaly detection on time series data, it can help staff quickly locate anomalies and promptly implement intervention measures to reduce safety risks and economic losses. For anomaly detection of multivariate time series data, an existing solution is to perform detection based on a graph neural network. However, this detection method has the following defects: (1) It ignores capturing both the short-term and long-term time dependencies of multivariate time series data; (2) It mainly captures the global spatial dependency relationship by constructing a fixed graph structure, ignoring that the potential dependency relationship between different variables in the multivariate time series may change dynamically over time, resulting in the model being unable to fully mine the global and local dynamic spatial features. Summary of the Invention

[0003] In view of the above technical problems, the technical solution adopted by the present invention is as follows:

[0004] An embodiment of the present invention provides a method for detecting anomalies in time series data, the method including the following steps:

[0005] S100, obtain a time series data set D to be detected = {D1, D2,..., D i ,..., D T}, D i is the data set corresponding to the monitoring time i, D i = {d i1 , d i2 ,..., d ij ,..., d in}, d ij is the j-th monitoring data in D i , the value range of i is from 1 to T, T is the length of the time series data, and the value range of j is from 1 to n, n is the number of data monitored at each monitoring time;

[0006] S200, using a sliding window with a sliding step of 1 and a window size of w to divide D, to obtain a sliding time series data set DS = {DS1, DS2, ..., DSh, ..., DSM}, where DSh is the hth sliding time series data set in DS, h ranges from 1 to M, M is the number of sliding time series data sets in DS, and M = T-w+1;

[0007] S300, input DSh into the first convolution module to obtain the corresponding first convolution feature Fh1 c ;

[0008] S400, Fh1 c Input the first spatiotemporal module to obtain the first spatiotemporal feature Fh1 st ; The first spatiotemporal module includes a first time module and a first space module, the first time module includes a first time attention module, a first dilated convolution module and a first channel attention module; the first space module includes a first static graph learning layer, a first dynamic graph learning layer and a first gated graph convolution module, the first gated graph convolution module includes a first static graph convolution module, a first dynamic graph convolution module and a first gated fusion module;

[0009] S500, Fh1 st and Fh1 c The added features are input into the second spatiotemporal module to obtain the second spatiotemporal feature Fh2 st ; The second space-time module has the same structure as the first space-time gate module;

[0010] S600, Fh1 st and Fh2 st The added features are input into the second convolution module to obtain the second convolution feature Fh2 c ;

[0011] S700, Fh2 c Input into the data prediction module and data reconstruction module respectively to obtain the predicted data Dh of the next moment corresponding to DSh nc and reconstruct the data Dh nr ;

[0012] S800, based on Dh next ,Dh nc and Dh nr Get the detection value Sh corresponding to the next moment corresponding to DSh next , if Sh next >S0, then judge Dh next is abnormal data, Dh next is the real data of the next moment corresponding to DSh, and S0 is the set threshold.

[0013] An embodiment of the present invention provides a non-transitory computer-readable storage medium, in which at least one instruction or at least one segment of program is stored, and the at least one instruction or the at least one segment of program is loaded and executed by a processor to implement the foregoing method.

[0014] An embodiment of the present invention further provides an electronic device, including a processor and the foregoing non-transitory computer-readable storage medium.

[0015] The present invention has at least the following beneficial effects:

[0016] For the time series data anomaly detection method provided by the embodiment of the present invention, first, the time module captures the short-term and long-term time dependencies of the multi-dimensional time series, and the space module learns the dynamic local spatial dependencies and global spatial dependencies between different variables. Then, data detection values are obtained based on the prediction model based on MLP and the reconstruction model based on AE. The present invention can detect time series anomalies more comprehensively. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a flowchart of the time series data anomaly detection provided by the embodiment of the present invention. Detailed Embodiments

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0020] An embodiment of the present invention provides a time series data anomaly detection method, as Figure 1 shown, the method may include the following steps:

[0021] S100, obtain a time series data set D = {D1, D2,..., D i ,..., D T}, D i is the data set corresponding to the monitoring moment i, and D i = {d i1 , d i2 ,..., d ij ,..., din}, d ij is D i The j-th monitoring data in, where i ranges from 1 to T, T is the length of the time series data, and j ranges from 1 to n, and n is the number of data monitored at each monitoring moment.

[0022] In the embodiments of the present invention, the time series data to be detected can be data obtained by monitoring multiple monitoring objects within a set time period. For example, data obtained by monitoring a water treatment system, including monitoring data such as water level, flow rate, water pressure, and valve status.

[0023] S200, divide D using a sliding window with a sliding stride of 1 and a window size of w to obtain a sliding time series data set DS = {DS1, DS2,..., DSh,..., DSM}, where DSh is the h-th sliding time series data set in DS, h ranges from 1 to M, M is the number of sliding time series data sets in DS, and M = T - w + 1.

[0024] In the embodiments of the present invention, w can be an empirical value. In one illustrative embodiment, w = 20.

[0025] S300, input DSh into the first convolution module to obtain the corresponding first convolution feature Fh1 c .

[0026] In the embodiments of the present invention, the first convolution module can be a 1×1 convolution layer.

[0027] S400, input Fh1 c into the first spatio-temporal module to obtain the first spatio-temporal feature Fh1 st ; the first spatio-temporal module includes a first time module and a first space module. The first time module includes a first time attention module, a first extended convolution module, and a first channel attention module; the first space module includes a first static graph learning layer, a first dynamic graph learning layer, and a first gated graph convolution module. The first gated graph convolution module includes a first static graph convolution module, a first dynamic graph convolution module, and a first gated fusion module.

[0028] Furthermore, in the embodiments of the present invention, S400 can specifically include:

[0029] S401, input Fh1 c into the first extended convolution module and the first time attention module respectively for feature extraction to obtain the first extended convolution feature Fh1 EC and the first time feature Fh1 t .

[0030] Ordinary convolutional neural networks cannot effectively model temporal dependencies. The present invention utilizes dilated convolutions in the temporal dimension to extract local temporal features of Fh1 c The first dilated convolution module maintains the length of the temporal dimension during the convolution process through a padding strategy. Specifically: φ represents the convolutional kernel, and ε d represents the dilated convolution operation with a dilation factor of d.

[0031] In an embodiment of the present invention, the first temporal attention module uses the multi-head self-attention mechanism to extract features in Fh1 c , including h self-attention heads and a feed-forward neural network. Fh1 can be obtained through the following steps: t :

[0032] (1) Respectively linearly map the input Fh1 i Q through the learnable parameter matrices W i K , W i V into the query matrix Q c , the key matrix K i , and the value matrix V i ; the value of i ranges from 1 to h;

[0033] (2) Calculate the self-attention result O i of the i-th head using scaled dot-product attention: i :

[0034] O i = Attention(Q i , K i , V i ) = softmax((Q T K i i / (d 1 / 2 )) i )V i )

[0035] where softmax() is the normalization function, and d t is the number of channels of the i-th head;

[0036] (3) Concatenate the self-attention results of the h heads and project them back to the original space through a linear mapping matrix to obtain the corresponding calculation result.

[0037] (4) Feed the calculation result into the feed-forward neural network to obtain Fh1 EC .

[0038] S402, take Fh1EC Extract features in the input first-channel attention module to obtain the first attention feature Fh1 a .

[0039] In the embodiment of the present invention, the first-channel attention module may include two fully connected layers with shared weights, namely the first fully connected layer and the second fully connected layer. Further, S402 may specifically include:

[0040] S4021, input Fh1 EC into the first fully connected layer and the second fully connected layer respectively to perform global average pooling and global max pooling processing respectively, and obtain the corresponding average pooling feature Fh1 avg EC and the max pooling feature Fh1 max EC .

[0041] S4022, obtain Fh1 a = M1×Fh1 EC ; M1 is the channel attention weight of the first-channel attention module, M1 = σ(W2(W1(Fh1 avg EC )) + W2(W1(Fh1 max EC ))); where, W1 and W2 are the weight parameters of the first fully connected layer and the second fully connected layer respectively, and σ() represents the sigmoid activation function.

[0042] S403, fuse Fh1 t and Fh1 a to obtain the fused feature Fh m , and input Fh m into the first dynamic graph learning layer and the first gated graph convolutional module respectively.

[0043] In the embodiment of the present invention, Fh m = σ(Fh1 a )°Fh1 t , ° represents the Hadamard product.

[0044] S404, based on the first dynamic graph learning layer, obtain the first dynamic adjacency matrix Ah m corresponding to Fh d , and input it into the first gated graph convolutional module.

[0045] In a multi-dimensional time series, the correlation between different variables is likely to change dynamically over time. Applying only a static graph structure cannot capture this local dynamic spatial dependence. Therefore, the present invention designs a dynamic graph learning layer, the core idea of which is to regard the variables in the multi-dimensional time series as nodes in the graph and use self-attention to calculate the strength of spatial correlation between the nodes.

[0046] Further, Ah d can be obtained based on the following steps:

[0047] S4041, divide Fh m along the time dimension into p segments, and aggregate the features Fh mg of the g-th segment to obtain the corresponding aggregated feature FAh mg ; the value of g ranges from 1 to p.

[0048] In the embodiment of the present invention, the aggregated feature can be implemented through convolution operation. Specifically, FAh mg = AGGREGATE(Fh mg ), where AGGREGATE() represents an aggregation operation implemented through convolution operation, used to reduce the time dimension to 1.

[0049] S4042, construct the dynamic graph structure Gh g corresponding to the g-th segment, and obtain the spatial correlation γ g between any two nodes u and v in Gh g uv = softmax(((FAh mg u W Q )·(FAh mg v W K )) T / (C) 1 / 2 ); where softmax() is a normalization function, FAh mg u is the aggregated feature of node u, FAh mg v is the aggregated feature of node v, the value ranges of u and v are 1 to n respectively; W Q is the query weight, W K is the key weight, ((FAh mg v W K ) T represents the bias matrix of the matrix obtained by multiplying FAh mg v and W K ; C is the channel dimension.

[0050] In the embodiment of the present invention, γ g uv is used to measure the spatial correlation strength between node u and node v in the g-th segment. A large value indicates strong correlation, while a small value indicates weak correlation.

[0051] S4043. Based on γ g uv , obtain the dynamic adjacency matrix Ah g d for the g-th segment. The x-th row of Ah g d includes (γ g x1 , γ g x2 , ……, γ g xz , ……, γ g xn ), where γ g xz is the spatial correlation between the x-th node and the z-th node in the g-th segment, and the values of x and z range from 1 to n respectively.

[0052] S4043. Obtain Ah d =(Ah d 1 , Ah d 2 , ……, Ah g d , ……, Ah p d ).

[0053] It can be known that the learning process of the dynamic graph structure depends on the input feature Fh m . Therefore, as time changes, Ah d will also change.

[0054] S405. Based on the first gated graph convolution module, obtain Fh1 st =α°Fh s +(1 - α)Fh d ; where α is the gating value, ° represents the Hadamard product, Fh s is the first global spatial feature obtained by convolving Fh m and the static adjacency matrix constructed by the first static graph learning layer through the first static graph convolution module, and Fh d is the first local spatial feature obtained by convolving Fh m and Ah d through the first dynamic convolution module.

[0055] To extract the local and global spatial features of time series data, the present invention designs a gated graph convolutional module based on a graph convolutional network.

[0056] In an embodiment of the present invention, the first static graph learning layer aims to learn an adaptive graph adjacency matrix in a data-driven manner without any prior knowledge for modeling the global spatial dependencies between variables. The present invention constructs a static graph structure using two randomly initialized node embedding matrices, and the calculation process is as follows:

[0057] L1 = tanh(E1θ1), L2 = tanh(E2θ2), As = Softx(ReLU(L1×L2 T -L2×L1 T )) where E1 represents

[0058] the source node embedding dictionary, E2 represents the target node embedding dictionary, which are learned by the stochastic gradient descent algorithm during model training, θ1 and θ2 are model parameters respectively, the ReLU activation function is used to eliminate the weak connections between nodes, and the softmax function is used to normalize the adjacency matrix.

[0059] In an embodiment of the present invention, the first static graph convolutional module and the first dynamic graph convolutional module can be existing graph convolutional neural networks. Among them, the convolution process of the first static graph convolutional model is: Fh s = σ(As N ×Fh m ×W), As N is the adjacency matrix obtained by normalizing the static adjacency matrix As constructed by the first static graph learning layer, and W is the weight parameter of the first static graph convolutional module.

[0060] Those skilled in the art know that in the actual prediction stage, the static adjacency matrix constructed by the first static graph learning layer is the trained adjacency matrix.

[0061] Furthermore, Fh d can be obtained through the following steps:

[0062] Step 1: Divide Fh m along the time dimension into p segments, and input the features Fh mg of the g-th segment and the corresponding dynamic adjacency matrix Ah g d into the dynamic graph convolutional layer to obtain the corresponding convolutional features Fh r g ;

[0063] Step 2: Fh r 1 , Fh r 2 , ……, Fhr g , ……, Fh r p are spliced to obtain Fh d .

[0064] S500, the feature obtained by adding Fh1 st and Fh1 c is input into the second spatio-temporal module to obtain the second spatio-temporal feature Fh2 st ; The structures of the second spatio-temporal module and the first spatio-temporal gate module are the same.

[0065] Those skilled in the art know that since the structures of the second spatio-temporal module and the first spatio-temporal module are the same, the specific implementation of the second spatio-temporal module can refer to the specific implementation of the first spatio-temporal module, that is, the acquisition method of Fh2 st can refer to the acquisition method of Fh1 st .

[0066] S600, the feature obtained by adding Fh1 st and Fh2 st is input into the second convolution module to obtain the second convolution feature Fh2 c .

[0067] In the embodiment of the present invention, the second convolution module can be a 1×1 convolution layer.

[0068] In the embodiment of the present invention, the input features of the second spatio-temporal module and the second convolution module consider residual connections, so the problem of gradient disappearance can be avoided.

[0069] S700, Fh2 c is respectively input into the data prediction module and the data reconstruction module to obtain the predicted data Dh nc corresponding to the next moment of DSh and the reconstructed data Dh nr .

[0070] In the embodiment of the present invention, the data prediction module is a multi-layer neural network model, and the data reconstruction module is an autoencoder architecture.

[0071] S800, based on Dh next , Dh nc and Dh nr obtain the detection value Sh next corresponding to the next moment of DSh. If Sh next > S0, then judge that Dh next is abnormal data, Dh next is the real data corresponding to the next moment of DSh, and S0 is a set threshold.

[0072] In an embodiment of the present invention, Sh next = ∑ n q=1 ((dh q next - dh q nc ) 2 +(dh q next - dh q nr ) 2 ) / (1 + β). Wherein, dh q next is the true value of the q-th monitoring data in Dh next , dh q nc is the predicted value corresponding to dh q next , dh q nr is the reconstructed value corresponding to dh q next . β is a hyperparameter used to optimize the combined ratio of the prediction error and the reconstruction error, which can be an empirical value, and preferably β is 0.8.

[0073] The time series data anomaly detection method provided by the embodiment of the present invention can be based on a trained time series data anomaly detection model. The time series data anomaly detection model can include a first convolution module, a first spatio-temporal module, a second spatio-temporal module, a second convolution module, a data prediction module, a data reconstruction module, and an anomaly detection module. The time series data anomaly detection model can be trained by a training set to obtain a trained time series data anomaly detection model.

[0074] Among them, the model training process can be executed with reference to the foregoing steps S100 to S800, and the loss function in the model is the sum of the loss functions of the data prediction module and the data reconstruction module. The model is trained by minimizing the sum of these two loss functions. Those skilled in the art know that the specific training process of the model can be the prior art.

[0075] In an embodiment of the present invention, S0 can be an empirical value.

[0076] In another embodiment of the present invention, S0 can be obtained based on a test set. Specifically, first, the test set is processed according to the above S100 to S800, and the corresponding F1 scores at all prediction times can be obtained. Then, the maximum value among all the F1 scores is taken as S0.

[0077] In summary, for the time series data anomaly detection method provided by the embodiments of the present invention, first, the time module captures the short-term and long-term time dependencies of the multi-dimensional time series, and the space module learns the dynamic local spatial dependencies and global spatial dependencies between different variables. Then, the data detection values are obtained based on the prediction model based on MLP and the reconstruction model based on AE. The present invention can detect time series anomalies more comprehensively.

[0078] An embodiment of the present invention further provides a non-transitory computer-readable storage medium, which can be set in an electronic device to store at least one instruction or at least one program related to a method for implementing a method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided in the above embodiment.

[0079] An embodiment of the present invention further provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.

[0080] An embodiment of the present invention further provides a computer program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps in the method according to various exemplary embodiments of the present invention described above in this specification.

[0081] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention disclosed is defined by the appended claims.

Claims

1. A method for detecting anomalies in time-series data, characterized in that, The method includes the following steps: S100, Obtain the time-series data set D = {D1, D2, ……, D i , ……, D T}, where D i is the data set corresponding to the monitoring time i, and D i = {d i1 , d i2 , ……, d ij , ……, d in}, and d ij is the j-th monitoring data in D i . The value range of i is from 1 to T, where T is the length of the time-series data, and the value range of j is from 1 to n, where n is the number of data monitored at each monitoring time; the time-series data to be detected are the data obtained by monitoring the water treatment system, including water level, flow rate, water pressure, and valve status; S200, divide D using a sliding window with a sliding stride of 1 and a window size of w to obtain a sliding time series data set DS = {DS1, DS2,..., DSh,..., DSM}, where DSh is the h-th sliding time series data set in DS, h ranges from 1 to M, and M is the number of sliding time series data sets in DS, M = T - w + 1; S300, input DSh into the first convolutional module to obtain the corresponding first convolutional feature Fh1 c ; S400, input Fh1 c Input it into the first spatio-temporal module to obtain the first spatio-temporal feature Fh1 st ; The first spatio-temporal module includes a first time module and a first space module. The first time module includes a first time attention module, a first extended convolution module, and a first channel attention module; The first space module includes a first static graph learning layer, a first dynamic graph learning layer, and a first gated graph convolution module. The first gated graph convolution module includes a first static graph convolution module, a first dynamic graph convolution module, and a first gated fusion module; S500 adds Fh1 st and Fh1 c and inputs the obtained feature into the second spatio-temporal module to obtain the second spatio-temporal feature Fh2 st ; the second spatio-temporal module has the same structure as the first spatio-temporal module; S600, add Fh1 st and Fh2 st to obtain the resulting feature, which is input into the second convolutional module to obtain the second convolutional feature Fh2 c ; S700, input Fh2 c into the data prediction module and the data reconstruction module respectively, to obtain the predicted data Dh at the next moment corresponding to DSh nc and the reconstructed data Dh nr ; S800, based on Dh next 、Dh nc and Dh nr Obtain the detection value Sh corresponding to the next moment of DSh next , if Sh next > S0, then determine that Dh next is abnormal data, Dh next is the real data of the next moment corresponding to DSh, and S0 is the set threshold value.

2. The method according to claim 1, characterized in that, Both the first convolutional module and the second convolutional module are 1×1 convolutional layers.

3. The method according to claim 1, characterized in that, S400 specifically includes: S401, input Fh1 c into the first extended convolution module and the first temporal attention module respectively for feature extraction, obtaining the first extended convolution feature Fh1 EC and the first temporal feature Fh1 t ; S402, input Fh1 EC into the first-channel attention module for feature extraction to obtain the first attention feature Fh1 a ; S403, fuse Fh1 t and Fh1 a to obtain the fused feature Fh m , and input Fh m into the first dynamic graph learning layer and the first gated graph convolutional module respectively; S404, obtain Fh based on the first dynamic graph learning layer m The corresponding first dynamic adjacency matrix Ah d , and input it into the first gated graph convolution module; S405, obtain Fh1 based on the first gated graph convolutional module st =α∘Fh s +(1 - α)Fh d ; where α is the gating value, ∘ represents the Hadamard product, and Fh s is the first global spatial feature obtained by convolving the first static graph convolutional module with the static adjacency matrix constructed by Fh m and the first static graph learning layer, and Fh d is the first local spatial feature obtained by convolving the first dynamic graph convolutional module with Fh m and Ah d ​ 4. The method according to claim 3, characterized in that, Ah d obtained based on the following steps: S4041, divide Fh m along the time dimension into p segments, and aggregate the features Fh mg of the g-th segment to obtain the corresponding aggregated feature FAh mg ; g ranges from 1 to p; S4042, construct the dynamic graph structure Gh corresponding to the g-th segment g , and obtain the spatial correlation γ g between any two nodes u and v in Gh g uv = softmax((FAh mg u W Q ) • (FAh mg v W K )) T / (C) 1 / 2 ); where softmax() is the normalization function, FAh mg u is the aggregated feature of node u, FAh mg v is the aggregated feature of node v, and the value ranges of u and v are from 1 to n; W Q is the query weight, W K is the key weight; C is the channel dimension; S4043, based on γ g uv , obtain the dynamic adjacency matrix Ah corresponding to the g-th segment g d , Ah g d 's x-th row includes (γ g x1 , γ g x2 , ……, γ g xz , ……, γ g xn ), where γ g xz is the spatial correlation between the x-th node and the z-th node in the g-th segment, and the values of x and z range from 1 to n respectively; S4043, Obtain Ah d = (Ah d 1 , Ah d 2 , ……, Ah g d , ……, Ah p d ).

5. The method according to claim 3, characterized in that, , φ represents the convolution kernel, represents the dilated convolution operation with a dilation factor of d, and ReLU() is the activation function.

6. The method according to claim 3, characterized in that, S402 specifically includes: S4021, input Fh1 EC into the first fully-connected layer and the second fully-connected layer respectively to perform global average pooling and global max pooling processing respectively, obtaining the corresponding average pooling feature Fh1 avg EC and max pooling feature Fh1 max EC ; S4022, obtain Fh1 a = M1 × Fh1 EC ; M1 is the channel attention weight of the first-channel attention module, M1 = σ(W2(W1(Fh1 avg EC )) + W2(W1(Fh1 max EC )); where W1 and W2 are the weight parameters of the first fully connected layer and the second fully connected layer respectively, and σ() represents the sigmoid activation function.

7. The method according to claim 1, characterized in that, The data prediction module is a multi-layer neural network model, and the data reconstruction module is an autoencoder architecture.

8. The method according to claim 1, characterized in that, Sh next =∑ n q=1 ((dh q next -dh q nc ) 2 +(dh q next -dh q nr ) 2 ) / (1 + β); where, dh q next is the true value of the q-th monitoring data in Dh next , dh q nc is the predicted value corresponding to dh q next , dh q nr is the reconstructed value corresponding to dh q next , and β is a hyperparameter.

9. A non-transitory computer-readable storage medium storing at least one instruction or at least one segment of program, characterized in that, The at least one instruction or the at least one program is loaded and executed by a processor to implement the method according to any one of claims 1-8.

10. An electronic device, characterized in that, It includes a processor and the non-transitory computer-readable storage medium described in claim 9.