A multi-element time sequence anomaly detection method based on a deep learning algorithm

CN122173906BActive Publication Date: 2026-09-29GUIZHOU UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610657098.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-09-29
Estimated Expiration
2046-05-13

AI Technical Summary

Technical Problem

[0004]为此,本发明提供一种基于深度学习算法的多元时间序列异常检测方法,用以克服现有技术中无法有效捕捉多尺度时序依赖和全局时空关联,导致重构误差对异常不敏感,降低了异常检测的准确性的问题

Benefits of technology

[0021]本方案中,通过提供第一预设重构误差计算方式和第二预设重构误差计算方式,可灵活适配不同多元时序数据的特性与异常检测需求,能够更为全面的捕捉各变量维度的重构偏差,兼顾误差计算的精准性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122173906B_ABST
    Figure CN122173906B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of time series data anomaly detection, and particularly relates to a multivariate time series anomaly detection method based on a deep learning algorithm, comprising: S1, screening training normal sample sets according to multivariate time series data; S2, performing feature embedding processing on the preprocessed training normal sample sets; S3, performing down-sampling, time series convolution processing and deep separable convolution processing on the embedded features according to a multistage time series convolution network; S4, performing nonlinear mapping and adaptive weighting on the time series feature representation according to a convolution feedforward network; S5, repeatedly performing S3 and S4 until the feature extraction of a preset number of stages in the multistage time series convolution network is completed, and performing regression reconstruction and inverse normalization processing on the fused feature representation; S6, comparing and analyzing the reconstruction error at each time with a preset standardized error threshold value. The present application can effectively capture multiscale time series dependence and global spatiotemporal correlation, and improve the accuracy of anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of time series data anomaly detection technology, and in particular to a multivariate time series anomaly detection method based on deep learning algorithms. Background Technology

[0002] With the rapid development of industrial automation, intelligent manufacturing, energy systems, and aerospace, large-scale sensor networks are widely deployed in complex systems to continuously collect equipment operating status and environmental information, thus forming high-dimensional, multivariate time series data. However, real-world multivariate time series data typically exhibits high dimensionality, strong noise, complex inter-variable relationships, and frequent data gaps, posing significant challenges to anomaly detection. On the one hand, there are significant temporal dependencies and spatial correlations between different sensors; on the other hand, anomaly patterns often exhibit the coexistence of local mutations and long-term trend anomalies, increasing the difficulty of modeling. To address these issues, existing technologies have proposed various deep learning-based anomaly detection methods. Among them, autoencoder-based reconstruction methods are widely used, constructing an encoder-decoder structure to compress and reconstruct time series features, and utilizing reconstruction errors for anomaly detection. However, these autoencoder-based reconstruction methods typically rely on explicit latent spatial compression processes, which can easily lead to the loss of key information when processing high-dimensional, noisy, and incomplete data, thus introducing reconstruction bias and affecting the stability and accuracy of anomaly detection. In recent years, Temporal Convolutional Networks (TCNs) have been applied to time series modeling to capture local temporal patterns due to their strong parallel computing capabilities and controllable receptive fields. Existing techniques have attempted to expand the model's temporal receptive field by introducing large-kernel convolutions to enhance its ability to model long-term dependencies. However, relying solely on convolutional structures still has limitations in characterizing cross-timescale dependencies and complex interactions between multiple variables. Furthermore, to better model the correlations between different sensors in multivariate time series, existing techniques have introduced spatiotemporal attention mechanisms to characterize long-term dependencies in the time dimension and interactions in the variable dimension. However, existing methods still struggle to simultaneously preserve key information from high-dimensional noisy data, capture cross-timescale temporal dependencies, and model complex multivariate correlations, resulting in anomaly detection accuracy and stability that fail to meet the demands of complex industrial scenarios.

[0003] Chinese Patent Publication No. CN116361635B discloses a method for anomaly detection in multidimensional time-series data, relating to the field of time-series data anomaly detection technology. The method includes: S1, extracting spatial features of multidimensional time-series data using Gaussian Atlas (GAT) and temporal features using a Generic Ruling Unit (GRU); S2, reconstructing the spatial and temporal features of the multidimensional time-series data using a multilayer perceptron; S3, obtaining a training set based on the reconstructed data and establishing an anomaly detection model based on the training set, then obtaining an optimal threshold using the anomaly detection model; S4, inputting the root mean square error of the reconstructed data as anomaly score into the anomaly detection model and comparing it with the optimal threshold to obtain the anomaly detection result. This approach only extracts single-scale features from the time-series data and then directly reconstructs it, failing to effectively capture multi-scale temporal dependencies and global spatiotemporal correlations. This results in the reconstruction error being insensitive to anomalies, reducing the accuracy of anomaly detection. Summary of the Invention

[0004] To address this issue, the present invention provides a multivariate time series anomaly detection method based on deep learning algorithms, which overcomes the problem in existing technologies that cannot effectively capture multi-scale temporal dependencies and global spatiotemporal correlations, resulting in reconstruction errors being insensitive to anomalies and reducing the accuracy of anomaly detection.

[0005] To achieve the above objectives, the present invention provides a multivariate time series anomaly detection method based on deep learning algorithms, comprising: S1. Collect multivariate time series data generated by multi-source sensors from the target monitoring system, and perform correlation analysis on the multivariate time series data to screen and obtain a training normal sample set; S2. Preprocess the normal training sample set to obtain the preprocessed normal training sample set, and perform feature embedding processing on the preprocessed normal training sample set to obtain the embedded features. S3. Input the embedded features into a multi-stage temporal convolutional network. At the beginning of each stage, downsample the embedded features to obtain downsampled features. Then, perform temporal convolution and depthwise separable convolution on the downsampled features to obtain the temporal feature representation. S4. Input the temporal feature representation into the convolutional feedforward network for nonlinear mapping to obtain nonlinear mapping features. Then, fuse the temporal convolutional features and nonlinear mapping features through residuals to obtain enhanced features. Adaptively weight the enhanced features through a spatiotemporal attention mechanism to obtain stage output features, and use the stage output features as input for the next stage. S5. Repeat S3 and S4 until the feature extraction of the preset number of stages in the multi-stage temporal convolutional network is completed, and the fused feature representation is obtained. The fused feature representation is then reconstructed using a regression linear projection method to obtain a reconstructed sequence. The reconstructed sequence is then denormalized to obtain a dimensional recovery sequence. S6. Calculate the time-by-time reconstruction error between the dimensional recovery sequence and the normalized sequence, and standardize the time-by-time reconstruction error to obtain the standardized error. Compare and analyze the standardized error with a preset standardized error anomaly threshold to obtain the anomaly detection result.

[0006] The technical principle of this application is as follows: multi-source time-series data generated by multiple sensors are collected from the target monitoring system. After correlation analysis, a normal sample set is selected for training. After preprocessing, feature embedding is performed to obtain embedded features. The embedded features are input into a multi-stage temporal convolutional network. Each stage first downsamples and then extracts temporal feature representations through temporal convolution and depthwise separable convolution. Nonlinear mapping is achieved through a convolutional feedforward network. The residual is fused with the temporal convolutional features and the nonlinear mapping features to obtain enhanced features. The stage output features are obtained through adaptive weighting by a spatiotemporal attention mechanism. The above steps are repeated to complete the feature extraction of a preset number of stages. The reconstructed sequence is obtained through regression linear projection reconstruction and then inversely normalized. The reconstruction error at each time step is calculated and standardized. The result of the anomaly detection is obtained by comparing it with a preset normalized error anomaly threshold.

[0007] Compared with existing technologies, the advantages of this application are as follows: By performing correlation analysis on multivariate time series data, a normal training sample set is obtained, reducing interference from invalid data; after preprocessing the normal training sample set, feature embedding is performed, effectively improving feature representation capabilities; the embedded features are input into a multi-stage temporal convolutional network, with downsampling at each stage reducing redundant information; the combination of temporal convolution processing and depthwise separable convolution processing accurately extracts the temporal features of multivariate time series; the temporal feature representation is input into a convolutional feedforward network for nonlinear mapping, and the residual is used to fuse the temporal convolution features and nonlinear mapping features to obtain... Enhanced features are then adaptively weighted using a spatiotemporal attention mechanism to highlight key features and improve feature fusion performance. Repeated feature extraction yields a comprehensive fused feature representation, which is then reconstructed and denormalized using a regression-linear projection method to restore the true features of the data. The time-by-time reconstruction error between the dimensional recovery sequence and the normalized sequence is calculated and standardized, then compared with a preset standardized error anomaly threshold to accurately identify anomalies. This effectively addresses the insufficient accuracy of existing technologies, improves the precision and reliability of multivariate time series anomaly detection, and ensures timely and accurate identification of anomalies in the target monitoring system.

[0008] Furthermore, S1 includes the following steps: S11. Construct a multivariate time series matrix corresponding to the multivariate time series data, wherein the multivariate time series matrix... The mathematical expression is: In the formula, Representing a multivariate time series matrix The total number of variables in the middle, Indicates the first Time series of 1 variable Indicates the first Time series of 1 variable Indicates the length of the time series. Indicates the first Time series of 1 variable belong A real space of dimensions; S12. Calculate the correlation coefficients between different variable sequences in the multivariate time series matrix, and construct a correlation coefficient matrix based on the correlation coefficients, wherein the mathematical expression of the correlation coefficient matrix is: In the formula, Represents the first element in the correlation coefficient matrix. Line number Column elements, Represents the first element in a multivariate time series matrix. Time series of 1 variable Represents the first element in a multivariate time series matrix. Time series of 1 variable This indicates the row index of the correlation coefficient matrix. Represents the column index of the correlation coefficient matrix. Represents the first element in a multivariate time series matrix. Time series of 1 variable With the multivariate time series matrix, the first Time series of 1 variable The function for calculating the correlation coefficient between them; S13. Using the correlation coefficient matrix as the correlation benchmark under normal operating conditions, a sliding window is used to divide the multivariate time series matrix into multiple time windows. The window correlation coefficient matrix between variables within each time window is calculated. By calculating the difference between the window correlation coefficient matrix and the correlation benchmark, and combining this with the stability analysis of the statistics of each variable within each time window, it is determined whether each time window is in a stable operating state. The time windows determined to be in a stable operating state are taken as data segments that meet the preset normal mode, and the extracted data segments are determined as the training normal sample set. The mathematical expression is: In the formula, This indicates the total number of normal segments. Indicates the number of times under normal operating conditions A time series segment.

[0009] In this scheme, by constructing a multivariate time series matrix and calculating the correlation coefficients between variable sequences, the relationship between each variable can be accurately characterized. Based on the correlation coefficient matrix, state stability detection can be performed, which can effectively screen out normal pattern data segments and integrate them into a training normal sample set, thereby improving sample quality and ensuring the accuracy and stability of anomaly detection.

[0010] Furthermore, S2 includes the following steps: S21. Perform reversible instance normalization on each variable in the training normal sample set to obtain a normalized sequence, wherein the mathematical expression of the normalized sequence is: In the formula, Indicates the first training sample set Normalized data of invertible instances of each variable Indicates the first training sample set The values ​​of each variable, Indicates the first training sample set The average of the variables, Indicates the first training sample set The standard deviation of each variable ; S22. Patch the normalized sequence based on the time dimension to obtain the partitioned normalized sequence. The mathematical expression of the partitioned normalized sequence is: In the formula, Indicates the number of normalized sequences. The set of time segments obtained by dividing each variable into Patch segments. This represents the total number of time segments obtained after the division. , Indicates the length of the normalized sequence. Indicates the fill length. Indicates the length of a time segment. Indicates the sliding step size. , Indicates the number of normalized sequences. The variable obtained after being partitioned by Patch is the first... A time segment; S23. Embedding features are obtained by performing embedding mapping on the partitioned normalized sequence through one-dimensional convolution. The mathematical expression of the embedding features is as follows: In the formula, Indicates the number of normalized sequences after partitioning. Embedded features obtained by embedding each variable through embedding mapping Indicates the embedding dimension. This represents a one-dimensional convolution weight matrix. Represents a one-dimensional convolution weight matrix Based on the length of the time segment For lines, in embedded dimensions A real matrix of columns, This indicates that based on the one-dimensional convolution weight matrix For input Perform a one-dimensional convolution operation.

[0011] In this scheme, by performing reversible instance normalization on each variable in the training normal sample set, the difference in units can be eliminated and the data distribution characteristics can be preserved; then, based on the time dimension, the data is segmented to adapt to temporal feature extraction; finally, one-dimensional convolution is used to achieve embedding mapping, which effectively extracts local temporal features and improves feature expression capabilities.

[0012] Furthermore, S3 includes the following steps: S31. The embedded features are used as input to a multi-stage temporal convolutional network to obtain input features. Then, the input features are downsampled in the temporal dimension using stride convolution and pooling to obtain downsampled features. The mathematical expression for temporal downsampling is: In the formula, In a multi-stage temporal convolutional network, the first... At the start of the phase, the result is obtained by downsampling the features from the previous phase. This indicates a downsampling operation. In a multi-stage temporal convolutional network, the first... Characteristics of stage output; S32. Perform convolution operations on the downsampled features using a preset temporal convolution kernel to obtain intermediate convolution features. The mathematical expression for the intermediate convolution features is: In the formula, In a multi-stage temporal convolutional network, the first... Intermediate convolutional features of staged temporal convolution. This represents a temporal convolution operation. In a multi-stage temporal convolutional network, the first... The kernel size of the stage; S33. Perform depthwise convolution on each feature channel in the intermediate convolutional features to obtain the temporal feature representation. The mathematical expression for depthwise convolution is: In the formula, Indicates the first intermediate convolutional feature The first feature channel Temporal characteristics at time point This represents the index of a single feature channel in the intermediate convolutional features. Indicates the time step index. This represents the offset index of the depthwise convolution kernel. In a multi-stage temporal convolutional network, the first... Phase, First The depthwise convolution kernel corresponding to the n feature channels is at the nth feature channel. The parameter for each offset position. Indicates intermediate convolutional features The Middle Each feature channel and time position is eigenvalues.

[0013] In this scheme, features are embedded by downsampling at each stage through a multi-stage temporal convolutional network, which can compress redundant temporal information and focus on key features. Temporal dependencies at different stages are extracted by using temporal convolutional kernels, and then deep convolution is performed on each feature channel to finely mine the temporal patterns of single channels, enhance the temporal expression and discriminativeness of features, and improve the model's ability to extract and model features from complex temporal data.

[0014] Furthermore, S4 includes the following steps: S41. Input the temporal feature representation into a convolutional feedforward network, and map the temporal feature representation through a first-layer one-dimensional convolution operation and a non-linear activation function to obtain the first-layer convolutional features. The mathematical expression of the first-layer convolutional features is as follows: In the formula, In a multi-stage temporal convolutional network, the first... The first layer of convolutional features in the stage, In a stage-series temporal convolutional network, the first... The temporal characteristics of the stage are represented. This represents the first layer of one-dimensional convolution operation. Represents a non-linear activation function; S42. The features of the first convolutional layer are processed by the second layer of one-dimensional convolution operation to obtain the nonlinear mapping features, wherein the mathematical expression of the nonlinear mapping features is: In the formula, In a multi-stage temporal convolutional network, the first... Nonlinear mapping characteristics of the stage This represents the second-layer one-dimensional convolution operation; S43. By fusing the time-series feature representation and the nonlinear mapping feature through residuals, the residual fusion feature is obtained. The mathematical expression of the residual fusion feature is as follows: In the formula, In a stage-series temporal convolutional network, the first... Residual fusion characteristics of the stage; S44. The residual fusion features are rearranged to obtain rearranged features. Temporal attention modeling and variable attention modeling are then performed on the rearranged features to obtain temporal attention representation and variable attention representation, respectively. The mathematical expression of the temporal attention representation is as follows: In the formula, In a multi-stage temporal convolutional network, the first... Phase-based time attention representation, This represents the total number of steps in time. Indicates the first Feature vectors at each time step Indicates the first Attention weights at each time step , Indicates the first Attention score at each time step Indicates the first Attention score at each time step Represents an exponential function. , Indicates the first Transpose of the learnable row vectors at each time step This represents the hyperbolic tangent activation function. Indicates the use of the first Feature vectors at each time step The weight matrix to be adjusted Indicates the first Learnable bias vectors at each time step; The mathematical expression for variable attention is: In the formula, In a multi-stage temporal convolutional network, the first... Stage-specific attention representations This represents the total number of feature dimensions. In a multi-stage temporal convolutional network, the first... Phase 1 Each feature dimension Indicates the first Phase 1 Attention weights for each feature dimension. , Indicates the first Attention weights for each feature dimension. Indicates the first Attention weights for each feature dimension. , Indicates the first The transpose of the learnable row vectors used to compute variable attention during the phase. This indicates the term used for the first stage in a multi-stage temporal convolutional network. Phase 1 Each feature dimension The weight matrix to be adjusted In a multi-stage temporal convolutional network, the first... The learnable bias vector for each stage; S45. The residual fusion features, temporal attention representation, and variable attention representation are fused using a residual method to obtain the stage output features, which are then used as the input for the next stage. The mathematical expression for the stage output features is as follows: , In a multi-stage temporal convolutional network, the first... The stage output characteristics of the stage.

[0015] In this scheme, feature representation is enhanced by convolutional feedforward network and residual fusion. After feature rearrangement, temporal attention modeling and variable attention modeling are performed respectively, which can adaptively focus on key time steps and important variable dimensions. Finally, multi-level features are fused by residual fusion and passed level by level, which can fully capture temporal dependencies and variable associations, and improve the model's modeling accuracy for complex multivariate time series data.

[0016] Furthermore, in S5, the reconstructed sequence The mathematical expression is: In the formula, This represents the fusion feature representation. This indicates the representation used to fuse features. The regression weight matrix that linearly maps to the original data space. This represents the regression bias during linear projection. The mathematical expression for the dimensional recovery sequence is: In the formula, Indicates the first reconstructed sequence The first variable Values ​​at each point in time. Indicating the dimensional recovery sequence of the first The first variable Values ​​at each point in time. Indicates the number of normalized sequences. The standard deviation of each variable Indicates the number of normalized sequences. The mean of each variable.

[0017] In this scheme, the reconstructed sequence is obtained by linearly mapping the fusion features, which can accurately restore the features of the original time series data. Then, by combining the mean and standard deviation of the normalized sequence to restore the dimensions of the reconstructed sequence, the data can be restored to the real physical space, ensuring that the reconstruction result is consistent with the actual data and improving the reliability of anomaly detection.

[0018] Furthermore, S6 includes the following steps: S61. Calculate the time-by-time reconstruction error of the dimensional recovery sequence and the normalized sequence at each time point in all variable dimensions, wherein the time-by-time reconstruction error is calculated by any one of the first preset reconstruction error calculation method and the second preset reconstruction error calculation method. S62. Based on the mean and standard deviation of the reconstruction error sequence, each reconstruction error in the reconstruction error sequence is standardized to obtain a standardized error sequence, wherein the mathematical expression of the standardization process is: In the formula, In the reconstruction error sequence, the first... Reconstruction error at each time point In the reconstruction error sequence, the first... The standard error of the reconstruction error at each time point after standardization. This represents the mean of the reconstructed error sequence. This represents the standard deviation of the reconstructed error sequence; S63. Compare each standardized error in the standardized error sequence with a preset standardized error anomaly threshold, and determine the time point corresponding to the standardized error that is greater than the preset standardized error anomaly threshold as an anomaly.

[0019] In this scheme, by calculating the reconstruction error of multiple variables at each time step, subtle deviations in the data can be accurately captured; then, by standardizing the mean and standard deviation of the reconstruction error sequence, the influence of differences in error magnitude can be eliminated; by comparing the standardized error with the preset standardized error anomaly threshold, abnormal time points can be determined efficiently and accurately, thereby improving the sensitivity of anomaly detection.

[0020] Furthermore, in step S61, the mathematical expression for the first preset reconstruction error calculation method is: In the formula, This indicates the first preset reconstruction error calculation method used to calculate the error. The time-by-time reconstruction error at each point in time. Indicates the number of normalized sequences. The first variable Values ​​at each point in time. This represents the total number of variables involved in calculating the time-by-time reconstruction error; The mathematical expression for the second preset reconstruction error calculation method is: In the formula, This indicates the result obtained using the second preset reconstruction error calculation method. The time-by-time reconstruction error at each point in time.

[0021] In this solution, by providing a first preset reconstruction error calculation method and a second preset reconstruction error calculation method, it can flexibly adapt to the characteristics and anomaly detection requirements of different multivariate time series data, and can more comprehensively capture the reconstruction deviation of each variable dimension while taking into account the accuracy of error calculation. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating a multivariate time series anomaly detection method based on deep learning algorithms according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the architecture of a multivariate time series anomaly detection method based on deep learning algorithm according to an embodiment of the present invention; Figure 3 This is a diagram showing the internal structure of a depth-separable convolution in an embodiment of the present invention. Figure 4 This is a flowchart illustrating the spatiotemporal attention mechanism in an embodiment of the present invention; Figure 5 This is a comparison chart of F1 scores for different models in this embodiment of the invention; Figure 6 These are heatmaps of F1 scores for different models in embodiments of the present invention; Figure 7 This is an ablation experiment result of the accuracy of ST-seriesTCN in spatiotemporal attention mechanism in an embodiment of the present invention; Figure 8 This is an ablation experiment result of the recall rate of ST-seriesTCN in the spatiotemporal attention mechanism in this embodiment of the invention; Figure 9 This is an ablation experiment result of the F1 score of ST-seriesTCN in the spatiotemporal attention mechanism in an embodiment of the present invention; Figure 10 This is a diagram showing the ablation experiment results of ST-seriesTCN in this embodiment of the invention, based on the selection of a preset standardized error anomaly threshold. Figure 11 This is an ablation experiment result of the recall rate of ST-seriesTCN in the embodiment of the present invention with the selection of a preset normalized error anomaly threshold; Figure 12 This is a diagram showing the ablation experiment results of ST-seriesTCN with F1 score on the selected preset standardized error anomaly threshold in this embodiment of the invention. Figure 13 This is a graph showing the ablation experiment results of the accuracy of ST-seriesTCN in FFNRatio selection in this embodiment of the invention; Figure 14 This is a diagram showing the ablation experiment results of the recall rate of ST-seriesTCN in the FFNRatio selection in this embodiment of the invention; Figure 15 This is a graph showing the ablation experimental results of ST-seriesTCN with F1 score selected in FFNRatio in an embodiment of the present invention. Detailed Implementation

[0023] The following detailed description illustrates the specific implementation method: like Figure 1 The diagram shown is a flowchart of a multivariate time series anomaly detection method based on a deep learning algorithm according to an embodiment of the present invention, including the following steps: S1. Collect multivariate time series data generated by multi-source sensors from the target monitoring system, and perform correlation analysis on the multivariate time series data to screen and obtain a training normal sample set; S2. Preprocess the normal training sample set to obtain the preprocessed normal training sample set, and perform feature embedding processing on the preprocessed normal training sample set to obtain the embedded features. S3. Input the embedded features into a multi-stage temporal convolutional network. At the beginning of each stage, downsample the embedded features to obtain downsampled features. Then, perform temporal convolution and depthwise separable convolution on the downsampled features to obtain the temporal feature representation. S4. Input the temporal feature representation into the convolutional feedforward network for nonlinear mapping to obtain nonlinear mapping features. Then, fuse the temporal convolutional features and nonlinear mapping features through residuals to obtain enhanced features. Adaptively weight the enhanced features through a spatiotemporal attention mechanism to obtain stage output features, and use the stage output features as input for the next stage. S5. Repeat S3 and S4 until the feature extraction of the preset number of stages in the multi-stage temporal convolutional network is completed, and the fused feature representation is obtained. The fused feature representation is then reconstructed using a regression linear projection method to obtain a reconstructed sequence. The reconstructed sequence is then denormalized to obtain a dimensional recovery sequence. S6. Calculate the time-by-time reconstruction error between the dimensional recovery sequence and the normalized sequence, and standardize the time-by-time reconstruction error to obtain the standardized error. Compare and analyze the standardized error with a preset standardized error anomaly threshold to obtain the anomaly detection result.

[0024] like Figure 2As shown in the diagram, this invention provides an architecture diagram of a multivariate time series anomaly detection method based on deep learning algorithms. First, correlation analysis is performed on multivariate time series data generated by multiple sensors to select normal data as the training normal sample set. Next, reversible instance normalization is performed on each variable in the training normal sample set to eliminate dimensional differences. Then, patching is performed along the time dimension and embedding is done through Conv1D (one-dimensional convolution) to obtain feature tensors. The feature tensors are input into stacked SeriesTCN (Sequence Temporal Convolutional Network) modules for multi-stage modeling: each stage first downsamples to expand the receptive field, then uses large-kernel convolution and depthwise separable convolution to capture long-range temporal dependencies, then uses two convolutional feedforward networks to model channel relationships and variable relationships respectively, and finally uses a spatiotemporal attention mechanism to highlight key time segments and key variables. After multiple stages are repeated, the original sequence is reconstructed through linear projection, and the physical dimensions are restored through inverse normalization. The reconstruction error is calculated and standardized. When the standardized error exceeds a preset standardized error threshold, it is judged as an anomaly. In the training phase, only normal data is trained, while the testing phase uses anomaly data for testing and detection.

[0025] Specifically, S1 includes the following steps: S11. Construct a multivariate time series matrix corresponding to the multivariate time series data, wherein the multivariate time series matrix... The mathematical expression is: In the formula, Representing a multivariate time series matrix The total number of variables in the middle, Indicates the first Time series of 1 variable Indicates the first Time series of 1 variable Indicates the length of the time series. Indicates the first Time series of 1 variable belong A real space of dimensions; S12. Calculate the correlation coefficients between different variable sequences in the multivariate time series matrix, and construct a correlation coefficient matrix based on the correlation coefficients, wherein the mathematical expression of the correlation coefficient matrix is: In the formula, Represents the first element in the correlation coefficient matrix. Line number Column elements, Represents the first element in a multivariate time series matrix. Time series of 1 variable Represents the first element in a multivariate time series matrix. Time series of 1 variable This indicates the row index of the correlation coefficient matrix. Represents the column index of the correlation coefficient matrix. Represents the first element in a multivariate time series matrix. Time series of 1 variable With the multivariate time series matrix, the first Time series of 1 variable The function for calculating the correlation coefficient between them; S13. Using the correlation coefficient matrix as the correlation benchmark under normal operating conditions, a sliding window is used to divide the multivariate time series matrix into multiple time windows. The window correlation coefficient matrix between variables within each time window is calculated. By calculating the difference between the window correlation coefficient matrix and the correlation benchmark, and combining this with the stability analysis of the statistics of each variable within each time window, it is determined whether each time window is in a stable operating state. The time windows determined to be in a stable operating state are taken as data segments that meet the preset normal mode, and the extracted data segments are determined as the training normal sample set. The mathematical expression is: In the formula, This indicates the total number of normal segments. Indicates the number of times under normal operating conditions A time series segment.

[0026] In this embodiment, the target monitoring system includes industrial equipment monitoring systems (such as production line status monitoring systems, CNC machine tool health monitoring systems, and online monitoring systems for chemical / power equipment), server and IT infrastructure monitoring systems (such as server operation and maintenance monitoring platforms, cloud resource monitoring systems, and data center operation and maintenance monitoring systems), and spacecraft / aircraft telemetry monitoring systems (such as satellite telemetry systems, rocket / spacecraft health monitoring systems, and UAV airborne status monitoring systems). The multi-dimensional time-series data includes, but is not limited to, industrial equipment sensor data (such as time-series data from sensors for equipment vibration, temperature, pressure, and rotational speed), server operating indicators (such as CPU utilization, memory usage, disk I / O, network traffic, and service response latency), and spacecraft telemetry data (time-series data from telemetry for attitude, temperature, voltage, fuel, and communication links). Methods for detecting stable operating conditions include sliding window statistics and energy change or distribution consistency checks. The preset normal mode refers to the operating mode of the target monitoring system corresponding to the multivariate time series matrix when it is in a statistically stable state. Specifically, it includes the following indicators: First, the local mean of the time series of each variable in the multivariate time series matrix does not exceed the preset mean stability threshold within the sliding window; Second, the local variance of the time series of each variable in the multivariate time series matrix does not exceed the preset variance stability threshold within the sliding window; Third, the correlation coefficient calculated based on the time series of any two variables in the multivariate time series matrix does not exceed the preset correlation coefficient stability threshold within the sliding window. When the above three indicators are simultaneously within their respective preset stability ranges, the data segment is determined to conform to the preset normal mode.

[0027] In step S13, using the correlation coefficient matrix as the correlation benchmark under normal operating conditions, the multivariate time series matrix is ​​divided into multiple time windows using the sliding window method. A multivariate time series segment is extracted at each window position, and the window correlation coefficient matrix between variables within that segment is calculated. Simultaneously, statistical stability analysis is performed on each variable within the time window (calculating the local mean and local variance of each variable's time series). The length and step size of the sliding window in the sliding window method are set according to the time span and data density of the actual multivariate time series data; for example, the window length is 100 time points, and the step size is 50 time points. The difference between the window correlation coefficient matrix and the correlation benchmark is calculated (specifically, element-by-element comparison). Furthermore, based on the statistical stability analysis results of each variable, the local mean is compared with a preset mean stability threshold, and the preset local variance is compared with a preset variance stability threshold. Among them, the preset mean stabilization threshold, the preset local variance, and the preset variance stabilization threshold are all determined based on the benchmark parameters under normal operating conditions. For example, the mean stabilization threshold is set to 5% of the benchmark mean under normal operating conditions, the variance stabilization threshold is set to 10% of the benchmark variance under normal operating conditions, and the correlation coefficient stabilization threshold is set to 0.1 (used to determine whether the difference between the window correlation coefficient matrix and the correlation benchmark is within the allowable range).

[0028] Based on the above difference calculation and statistical stability analysis, it is determined whether the time window is in a stable operating state: when the absolute value of the difference between all elements in the window correlation coefficient matrix and the corresponding elements of the correlation benchmark is less than the preset correlation coefficient stability threshold (0.1), and the relative deviation of the local mean of all variables from the benchmark mean under normal operating conditions is less than the preset mean stability threshold (5% of the benchmark mean), and the relative deviation of the local variance of all variables from the benchmark variance under normal operating conditions is less than the preset variance stability threshold (10% of the benchmark variance), the time window is determined to be in a stable operating state. The multivariate time series segments corresponding to the time windows in a stable operating state are taken as data segments that meet the preset normal mode. The multivariate time series matrix segments corresponding to the continuous windows that are determined to meet the preset normal mode are extracted, and the overlapping parts are removed to form a training normal sample set.

[0029] Specifically, S2 includes the following steps: S21. Perform reversible instance normalization on each variable in the training normal sample set to obtain a normalized sequence, wherein the mathematical expression of the normalized sequence is: In the formula, Indicates the first training sample set Normalized data of invertible instances of each variable Indicates the first training sample set The values ​​of each variable, Indicates the first training sample set The average of the variables, Indicates the first training sample set The standard deviation of each variable ; S22. Patch the normalized sequence based on the time dimension to obtain the partitioned normalized sequence. The mathematical expression of the partitioned normalized sequence is: In the formula, Indicates the number of normalized sequences. The set of time segments obtained by dividing each variable into Patch segments. This represents the total number of time segments obtained after the division. , Indicates the length of the normalized sequence. Indicates the fill length. Indicates the length of a time segment. Indicates the sliding step size. , Indicates the number of normalized sequences. The variable obtained after being partitioned by Patch is the first... A time segment; S23. Embedding features are obtained by performing embedding mapping on the partitioned normalized sequence through one-dimensional convolution. The mathematical expression of the embedding features is as follows: In the formula, Indicates the number of normalized sequences after partitioning. Embedded features obtained by embedding each variable through embedding mapping Indicates the embedding dimension. This represents a one-dimensional convolution weight matrix. Represents a one-dimensional convolution weight matrix Based on the length of the time segment For lines, in embedded dimensions A real matrix of columns, This indicates that based on the one-dimensional convolution weight matrix For input Perform a one-dimensional convolution operation.

[0030] Specifically, S3 includes the following steps: S31. The embedded features are used as input to a multi-stage temporal convolutional network to obtain input features. Then, the input features are downsampled in the temporal dimension using stride convolution and pooling to obtain downsampled features. The mathematical expression for temporal downsampling is: In the formula, In a multi-stage temporal convolutional network, the first... At the start of the phase, the result is obtained by downsampling the features from the previous phase. This indicates a downsampling operation. In a multi-stage temporal convolutional network, the first... Characteristics of stage output; S32. Perform convolution operations on the downsampled features using a preset temporal convolution kernel to obtain intermediate convolution features. The mathematical expression for the intermediate convolution features is: In the formula, In a multi-stage temporal convolutional network, the first... Intermediate convolutional features of staged temporal convolution. This represents a temporal convolution operation. In a multi-stage temporal convolutional network, the first... The kernel size of the stage; S33. Perform depthwise convolution on each feature channel in the intermediate convolutional features to obtain the temporal feature representation. The mathematical expression for depthwise convolution is: In the formula, Indicates the first intermediate convolutional feature The first feature channel Temporal characteristics at time point This represents the index of a single feature channel in the intermediate convolutional features. Indicates the time step index. This represents the offset index of the depthwise convolution kernel. In a multi-stage temporal convolutional network, the first... Phase, First The depthwise convolution kernel corresponding to the n feature channels is at the nth feature channel. The parameter for each offset position. Indicates intermediate convolutional features The Middle Each feature channel and time position is eigenvalues.

[0031] Furthermore, in step S31, the downsampling operation is completed by a temporal convolutional layer with a stride greater than 1 in conjunction with a pooling layer. The temporal convolutional layer is set with a kernel stride of 2 and combined with a max pooling layer. The pooling window size and stride are both 2. It slides along the time dimension and selects the maximum value within the window as the output each time, thereby reducing the length of the time dimension to half of the original.

[0032] In step S32, the temporal convolution operation uses a preset temporal convolution kernel. The specific value of the preset temporal convolution kernel is determined based on different datasets, or according to the temporal resolution of the downsampled features and the feature extraction requirements. For example, a preset temporal convolution kernel of 13 or 25 indicates the coverage length of the temporal convolution kernel in the time dimension. During convolution, a corresponding number of zero values ​​are padded at both ends of the time dimension according to the size of the temporal convolution kernel to ensure that the length of the convolution output in the time dimension is consistent with the length of the input downsampled features. Maintain consistency; the number of input channels of the temporal convolution kernel and the downsampling features The number of feature channels is matched with the number of output channels, and the number of input channels is kept consistent with the number of output channels. The inner product is calculated and accumulated through a sliding window to generate intermediate convolutional features. .

[0033] In step S33, the depthwise convolution operation treats each feature channel as an independent processing unit. For the c-th feature channel, an independent depthwise convolution is performed. The time length of the depthwise convolution is... Consistent with the temporal convolution kernel size used in step S32, the convolution kernel parameters for each channel are learned independently during training. During depthwise convolution, intermediate convolution features are used. Centered on each time step t, according to the index Extraction length is The time window is defined, and the values ​​within that time window are multiplied element-wise with the corresponding convolutional kernel weights and summed. The result is then assigned to the output temporal convolutional feature Y at the corresponding channel c and time step t. Specifically, the internal structure of depthwise separable convolution is as follows: Figure 3 As shown, after the feature input, it is processed by two parallel branches: one branch first performs one-dimensional small kernel convolution, and then connects to one-dimensional batch normalization; the other branch connects to one-dimensional large kernel convolution. The features processed by the two branches are merged in the feature fusion layer to finally obtain the feature output.

[0034] Specifically, S4 includes the following steps: S41. Input the temporal feature representation into a convolutional feedforward network, and map the temporal feature representation through a first-layer one-dimensional convolution operation and a non-linear activation function to obtain the first-layer convolutional features. The mathematical expression of the first-layer convolutional features is as follows: In the formula, In a multi-stage temporal convolutional network, the first... The first layer of convolutional features in the stage, In a stage-series temporal convolutional network, the first... The temporal characteristics of the stage are represented. This represents the first layer of one-dimensional convolution operation. Represents a non-linear activation function; S42. The features of the first convolutional layer are processed by the second layer of one-dimensional convolution operation to obtain the nonlinear mapping features, wherein the mathematical expression of the nonlinear mapping features is: In the formula, In a multi-stage temporal convolutional network, the first... Nonlinear mapping characteristics of the stage This represents the second-layer one-dimensional convolution operation; S43. By fusing the time-series feature representation and the nonlinear mapping feature through residuals, the residual fusion feature is obtained. The mathematical expression of the residual fusion feature is as follows: In the formula, In a stage-series temporal convolutional network, the first... Residual fusion characteristics of the stage; S44. The residual fusion features are rearranged to obtain rearranged features. Temporal attention modeling and variable attention modeling are then performed on the rearranged features to obtain temporal attention representation and variable attention representation, respectively. The mathematical expression of the temporal attention representation is as follows: In the formula, In a multi-stage temporal convolutional network, the first... Phase-based time attention representation, This represents the total number of steps in time. Indicates the first Feature vectors at each time step Indicates the first Attention weights at each time step , Indicates the first Attention score at each time step Indicates the first Attention score at each time step Represents an exponential function. , Indicates the first Transpose of the learnable row vectors at each time step This represents the hyperbolic tangent activation function. Indicates the use of the first Feature vectors at each time step The weight matrix to be adjusted Indicates the first Learnable bias vectors at each time step; The mathematical expression for variable attention is: In the formula, In a multi-stage temporal convolutional network, the first... Stage-specific attention representations This represents the total number of feature dimensions. In a multi-stage temporal convolutional network, the first... Phase 1 Each feature dimension Indicates the first Phase 1 Attention weights for each feature dimension. , Indicates the first Attention weights for each feature dimension. Indicates the first Attention weights for each feature dimension. , Indicates the first The transpose of the learnable row vectors used to compute variable attention during the phase. This indicates the term used for the first stage in a multi-stage temporal convolutional network. Phase 1 Each feature dimension The weight matrix to be adjusted In a multi-stage temporal convolutional network, the first... The learnable bias vector for each stage; S45. The residual fusion features, temporal attention representation, and variable attention representation are fused using a residual method to obtain the stage output features, which are then used as the input for the next stage. The mathematical expression for the stage output features is as follows: , In a multi-stage temporal convolutional network, the first... The stage output characteristics of the stage.

[0035] Specifically, in S5, the reconstructed sequence The mathematical expression is: In the formula, This represents the fusion feature representation. This indicates the representation used to fuse features. The regression weight matrix that linearly maps to the original data space. This represents the regression bias during linear projection. The mathematical expression for the dimensional recovery sequence is: In the formula, Indicates the first reconstructed sequence The first variable Values ​​at each point in time. Indicating the dimensional recovery sequence of the first The first variable Values ​​at each point in time. Indicates the number of normalized sequences. The standard deviation of each variable Indicates the number of normalized sequences. The mean of each variable.

[0036] In this embodiment, the temporal feature representation is input into a convolutional feedforward network. The number of stages in the preset multi-stage temporal convolutional network is determined based on the temporal complexity and feature depth requirements of different input datasets. For example, the preset number of stages in the multi-stage temporal convolutional network is set to 5. For the s-th stage in the multi-stage temporal convolutional network, the size of the one-dimensional convolutional kernel is determined based on the temporal local dependency feature scale of the input dataset. The number of output channels of the one-dimensional convolutional kernel is set based on the feature dimension of the dataset and the model representation capability requirements. For example, a first-layer one-dimensional convolutional operation with a kernel size of 3 and 64 output channels is used to perform convolution calculation on the temporal feature representation. The convolution stride is determined based on the temporal sampling interval of the input dataset and the feature sliding requirements. Padding... The method is determined based on the matching requirements of the convolution output dimension. For example, the convolution stride of the first-layer one-dimensional convolution operation is set to 1, and the padding method is set to Same to ensure that the output dimension is consistent with the temporal feature representation. The random initialization range of the weight matrix of the first-layer one-dimensional convolution operation is determined based on the network convergence stability and gradient propagation characteristics. The initialization value of the bias vector is determined based on the adaptability of the activation function. For example, the random initialization range of the weight matrix of the first-layer one-dimensional convolution operation is [-0.01, 0.01], and the bias vector of the first-layer one-dimensional convolution operation is initialized to 0. Then, the convolution result is input into the ReLU nonlinear activation function for nonlinear mapping processing, and finally the first-layer convolution feature of the s-th stage in the multi-stage temporal convolutional network is obtained.

[0037] Based on the first-layer convolutional features of the s-th stage in the multi-stage temporal convolutional network, a second-layer one-dimensional convolutional operation with the same parameters as the first-layer one-dimensional convolutional operation is used for feature processing. The kernel size, number of output channels, stride, and padding method of the second-layer one-dimensional convolutional operation are determined synchronously according to the characteristics of the dataset. For example, the kernel size is set to 3, the number of output channels to 64, and the stride to 1. The initialization range of the weight matrix and the initial value of the bias vector for the second-layer one-dimensional convolutional operation are also determined according to the model training characteristics. For example, the padding method is set to Same, the weight matrix is ​​randomly initialized within the range [-0.01, 0.01], and the bias vector is initialized to 0. Convolutional operations are then performed on the first-layer convolutional features to further extract deeper features and complete feature mapping, ultimately obtaining the nonlinear mapping features of the s-th stage in the multi-stage temporal convolutional network.

[0038] When fusing temporal feature representations and nonlinear mapping features using residual fusion, the first step is to ensure that the dimensions of the temporal feature representations and nonlinear mapping features are completely consistent. If there is a dimension mismatch, a 1×1 convolution operation is used to adjust them to a unified number of output channels and a unified number of time steps. The number of output channels and time steps are determined based on the feature dimension and temporal length of the input dataset; for example, a unified dimension of 64 output channels and 100 time steps is achieved. Subsequently, element-wise addition is performed on the temporal feature representations and nonlinear mapping features. During the operation, the feature values ​​at each position of the temporal feature representations and nonlinear mapping features are kept in correspondence, without changing the number of time steps or the feature dimension. This preserves the original information of the temporal feature representations while incorporating the deep features of the nonlinear mapping features, ultimately yielding the residual fused feature of the s-th stage in the multi-stage temporal convolutional network.

[0039] When rearranging the residual fusion features, the dimensions of the residual fusion features are rearranged from [total time steps, total number of feature dimensions] to [total number of feature dimensions, total time steps]. The total number of time steps is determined based on the length of the time series in the input dataset, and the total number of feature dimensions is determined based on the feature dimensions of the dataset variables. For example, the preset total number of time steps is 100, and the total number of feature dimensions is 64. When performing temporal attention modeling on the rearranged features, the dimensions of the feature vector weight matrix are adjusted according to the total number of feature dimensions. For example, a weight matrix with dimensions [64, 64] is used. The initialization range of the feature vector weight matrix is ​​determined according to the feature dimension matching requirements. For example, the initialization range of the feature vector weight matrix is ​​set to [-0.01, 0.01]. The dimension of the feature vector bias vector is determined according to the feature dimension matching requirements. For example, a bias vector with dimensions [64, 1] and initialized to 0 is used. After inputting the hyperbolic tangent activation function, the bias vector is compared with the learnable row vector dimension. The attention score for each time step is obtained by multiplying the transposes of learnable row vectors of dimension [1, 64], determined by the total number of feature dimensions. This score is then normalized using an exponential function to obtain attention weights (the sum of weights is 1). A weighted sum is then used to obtain the temporal attention representation for stage s in the multi-stage temporal convolutional network. Simultaneously, variable attention modeling is performed using the same process. The dimension of the weight matrix for variable attention modeling is determined based on the total number of time steps; for example, the dimension of the weight matrix for variable attention modeling is [100, 100]. Finally, the variable attention representation for stage s in the multi-stage temporal convolutional network is obtained. Specifically, the flowchart of the spatiotemporal attention mechanism is shown below. Figure 4As shown, the input feature shape is [B,M,D,N], which enters two parallel branches: the spatial attention branch reshapes the features to [B×N,M,D] to model the dependencies between different sensors (variables); the temporal attention branch reshapes the features to [B×M,N,D] to capture the importance of key time segments. The features processed by the two branches are then fused to restore the original dimensions [B,M,D,N], and added to the input features via residual connections to obtain the enhanced feature output. This spatiotemporal attention mechanism highlights key variables and key time points by jointly learning the importance weights of the temporal and spatial dimensions.

[0040] When fusing residual fusion features, temporal attention representations, and variable attention representations using a residual approach, a 1×1 convolution operation is first used to adjust the residual fusion features, temporal attention representations, and variable attention representations to a unified number of output channels and a unified number of time steps. The number of output channels and the number of time steps are determined based on the feature dimension and temporal length of the input dataset, for example, a unified dimension of 64 output channels and 100 time steps. Subsequently, element-wise addition is performed on the residual fusion features, temporal attention representations, and variable attention representations to maintain a one-to-one correspondence between feature values ​​at each position. Finally, the stage output feature of the s-th stage in the multi-stage temporal convolutional network is obtained. This stage output feature is directly used as the temporal feature representation of the next stage in the multi-stage temporal convolutional network and enters the next round of feature processing.

[0041] Repeat the relevant feature extraction steps until the preset number of features in the multi-stage temporal convolutional network are extracted. Integrate the output features of all stages to obtain a fused feature representation. Regression-based linear projection is used to reconstruct the fused feature representation. The dimension of the regression weight matrix is ​​determined based on the dimensions of the original data and the fused features; for example, the dimension of the regression weight matrix is ​​set to [original data dimension 64, fused feature dimension 64]. The initial range of the regression weight matrix is ​​determined based on the linear projection fitting effect; for example, the initial range of the regression weight matrix is ​​[-0.01, 0.01]. The initial value of the regression bias is determined based on the data fitting deviation requirements; for example, the regression bias is initialized to 0. A reconstructed sequence is obtained through linear operations. The reconstructed sequence is then denormalized. The standard deviation and mean of each variable in the normalized sequence are determined based on the statistical distribution characteristics of the original data; for example, the standard deviation of each variable in the preset normalized sequence is 0.5 and the mean is 2.0. The value of the c-th variable at time t in the reconstructed sequence is multiplied by the standard deviation of the corresponding variable and added to the mean of the corresponding variable to obtain the dimensionally restored sequence.

[0042] Specifically, S6 includes the following steps: S61. Calculate the time-by-time reconstruction error of the dimensional recovery sequence and the normalized sequence at each time point in all variable dimensions, wherein the time-by-time reconstruction error is calculated by any one of the first preset reconstruction error calculation method and the second preset reconstruction error calculation method. S62. Based on the mean and standard deviation of the reconstruction error sequence, each reconstruction error in the reconstruction error sequence is standardized to obtain a standardized error sequence, wherein the mathematical expression of the standardization process is: In the formula, In the reconstruction error sequence, the first... Reconstruction error at each time point In the reconstruction error sequence, the first... The standard error of the reconstruction error at each time point after standardization. This represents the mean of the reconstructed error sequence. This represents the standard deviation of the reconstructed error sequence; S63. Compare each standardized error in the standardized error sequence with a preset standardized error anomaly threshold, and determine the time point corresponding to the standardized error that is greater than the preset standardized error anomaly threshold as an anomaly.

[0043] Specifically, in S61, the mathematical expression for the first preset reconstruction error calculation method is: In the formula, This indicates the first preset reconstruction error calculation method used to calculate the error. The time-by-time reconstruction error at each point in time. Indicates the number of normalized sequences. The first variable Values ​​at each point in time. This represents the total number of variables involved in calculating the time-by-time reconstruction error; The mathematical expression for the second preset reconstruction error calculation method is: In the formula, This indicates the result obtained using the second preset reconstruction error calculation method. The time-by-time reconstruction error at each point in time.

[0044] In this embodiment, the time-by-time reconstruction error between the dimensional recovery sequence and the normalized sequence is calculated. The total number of variables involved in calculating the time-by-time reconstruction error is determined based on the actual number of monitored variables in the input dataset. For example, if the preset total number of variables involved in calculating the time-by-time reconstruction error is 50, either of the two preset reconstruction error calculation methods can be selected. The first preset reconstruction error calculation method is to sum the absolute differences between the values ​​of the c-th variable at the t-th time point in the dimensional recovery sequence and the values ​​of the c-th variable at the t-th time point in the normalized sequence across all variable dimensions at each time point, thus obtaining the time-by-time reconstruction error at that time point. The second preset reconstruction error calculation method is to square the above absolute differences and then sum them, thus obtaining the time-by-time reconstruction error at that time point. Both methods require error calculation for all time points.

[0045] A reconstruction error sequence is formed based on the time-by-time reconstruction errors at all time points. First, the mean and standard deviation of the reconstruction error sequence are calculated. The mean of the reconstruction error sequence is determined by the statistical average of the reconstruction errors at all time points. For example, the preset mean of the reconstruction error sequence is 0.8. The standard deviation of the reconstruction error sequence is determined by the statistical calculation of the dispersion of the reconstruction errors at all time points. For example, the preset standard deviation of the reconstruction error sequence is 0.2. Then, each time-by-time reconstruction error in the reconstruction error sequence is standardized by subtracting the mean of the reconstruction error sequence from each time-by-time reconstruction error and then dividing by the standard deviation of the reconstruction error sequence to eliminate the influence of the dimension of the error, thus obtaining a standardized error sequence. This ensures that each standardized error has a unified comparison standard and that each standardized error corresponds to the error at a time point in the reconstruction error sequence.

[0046] Based on the above content, conduct experimental analysis: A systematic comparative experiment was conducted using five publicly available multivariate time series anomaly detection datasets (SMD, MSL, SMAP, SWAT, and PSM). Specific information about the multivariate time series anomaly detection datasets is shown in Table 1.

[0047] Table 1. Detailed Information on Public Datasets In this embodiment, the method of this application is set as ST-seriesTCN, such as... Figure 5As shown, ST-SeriesTCN is compared with 12 existing models, including Crossformer, Transformer, Autoformer, and LSTM, on five datasets: SMD (Server Machine Dataset), SMAP (Mars Science Laboratory Dataset), MSL (Soil Moisture Active Passive Satellite Dataset), SWaT (Secure Water Treatment Dataset), and PSM (Pooled Server Metrics Dataset). ST-SeriesTCN achieved F1 scores of 84.0, 83.0, 74.0, 96.0, and 98.0 on the five datasets, respectively, reaching the best or near-best levels. Compared to the near-best models, it improves by 3 and 1 percentage points on SWaT and PSM, respectively, and leads by more than 1 point on SMD and MSL. The ST-seriesTCN performs consistently well across multiple datasets, demonstrating stronger generalization ability and detection accuracy, thus validating the significant advantages of ST-seriesTCN over mainstream methods.

[0048] like Figure 6As shown, the F1 score heatmaps of ST-seriesTCN on five datasets—SMD, MSL, SMAP, SWAT, and PSM—are compared and analyzed. These models include ST-seriesTCN, Crossformer, TimeNet, Autoformer, DLinear, LSTM, Transformer, MICN, LightTS, Stationary, FEDformer, FILM, and Informer (a deep learning model designed for long-sequence time series prediction tasks). The experiments use a unified evaluation framework and compare the performance of each model on multivariate time series anomaly detection tasks. The results show that ST-seriesTCN achieved the highest scores in SMD (84.0%), MSL (82.4%), and SWAT (94.5%), and also maintained a leading position in SMAP (73.0%) and PSM (92.3%), demonstrating significant overall advantages. By fusing temporal convolution and spatial topology, ST-seriesTCN effectively captures long-range dependencies and complex spatial relationships, exhibiting superior detection accuracy and generalization ability in multiple real-world scenarios, validating its outstanding advantages over existing methods.

[0049] like Figure 7 , Figure 8 and Figure 9 As shown, an ablation experiment was conducted on the spatiotemporal attention mechanism module of the ST-seriesTCN model. The experiment mainly set up multiple control schemes for comparative analysis: no spatiotemporal attention mechanism, spatiotemporal attention mechanism before FFN, spatiotemporal attention mechanism after FFN, and spatiotemporal attention mechanism inserted before and after FFN. Comparative experiments on the datasets SMD, MSL, SMAP, SWaT, and PSM showed that the scheme placing the spatiotemporal attention mechanism after FFN performed best. PSM and SWaT maintained high and stable performance in accuracy, recall, and F1 score. SMAP reached its peak in F1 score. MSL's recall significantly improved with the placement of spatiotemporal attention, while only SMD showed relatively small fluctuations. Compared with schemes without spatiotemporal attention, with spatiotemporal attention before, or simultaneously using both, placing spatiotemporal attention after FFN more efficiently captures spatiotemporal features, fully demonstrating the performance advantage of ST-seriesTCN after reasonable configuration of the spatiotemporal attention mechanism.

[0050] like Figure 10 , Figure 11 and Figure 12As shown, the ablation experiments of ST-seriesTCN with preset standardized error outlier thresholds were mainly selected. The ablation experiments were conducted with different preset standardized error outlier thresholds of 0.5, 1.0, 1.5, 2.0, 2.5, and 3.0. Comparative experiments were plotted on the datasets SMD, MSL, SMAP, SWAT, and PSM. Figure 10 , Figure 11 and Figure 12 As can be seen, ST-seriesTCN exhibits robust performance across various preset standardized error anomaly thresholds: the PSM dataset maintains high levels with minimal fluctuations in accuracy, recall, and F1 score, demonstrating the best performance; the SMD dataset achieves a peak F1 score of 84.03 at the preset standardized error anomaly threshold of 2.0, significantly outperforming the others; the MSL dataset shows a steady increase in recall with increasing preset standardized error anomaly threshold, while the F1 score remains relatively stable; the SWaT dataset, although experiencing a slight decrease in accuracy with increasing preset standardized error anomaly threshold, still maintains a high level, with stable recall; the SMAP dataset, while having a slightly lower F1 score, shows a stable trend. This indicates that ST-seriesTCN is highly robust to changes in the preset standardized error anomaly threshold, effectively balancing accuracy and recall.

[0051] like Figure 13 , Figure 14 and Figure 15 As shown, ablation experiments of ST-seriesTCN with varying FFNRatios were conducted, primarily selecting ablation experiments with FFNRatios of 1, 2, 3, and 4. In the ST-seriesTCN model, FFNRatio refers to the ratio of the dimension of the intermediate hidden layers to the dimension of the input features in the feedforward neural network (FFN) module. Comparative experiments were plotted on the datasets SMD, MSL, SMAP, SWAT, and PSM. Figure 13 , Figure 14 and Figure 15 It can be seen that the performance of ST-seriesTCN exhibits a significant regularity with changes in FFNRatio: when FFNRatio is 1, the model demonstrates optimal or near-optimal overall performance across all datasets, with accuracy, recall, and F1 score remaining stable at high levels; when FFNRatio increases to 2-3, some metrics (such as accuracy for SMD and F1 score for PSM) show slight fluctuations or decreases; and when FFNRatio increases to 4, multiple evaluation metrics show a significant decline, indicating reduced model robustness. Comparing different datasets, PSM and SWAT maintain high metric stability across all FFNRatios, while SMD and SMAP are more sensitive to changes in FFNRatio. This suggests that ST-seriesTCN can achieve efficient feature extraction and computational resource balance when FFNRatio=1.

[0052] Table 2 shows a comparison of ST-seriesTCN with existing models Crossformer, TimeNet, Autoformer, DLinear, Lstm, Transformer, MICN, LightTS, stationary, PEDformer, FILM, and Informer on mainstream datasets SMD, MSL, SMAP, SWAT, and PSM in terms of F1 score, P (accuracy), and R (recall).

[0053] Table 2 shows the benchmark results of ST-seriesTCN and existing models on different datasets. As shown in Table 2, ST-SeriesTCN performs exceptionally well in the core evaluation metric of F1 score: it achieves the best results on the SMD (84.03), SWaT (94.49), and PSM (97.35) datasets, significantly outperforming the comparison models; its F1 score (82.37) on the MSL dataset is also among the leading ones. Meanwhile, ST-SeriesTCN also achieves high performance in metrics such as recall (97.50) on the PSM dataset and F1 score on the SWaT dataset, effectively balancing accuracy and recall, demonstrating stronger robustness and generalization ability, and fully verifying the superiority of ST-SeriesTCN in time series anomaly detection tasks.

[0054] The ablation experiment of ST-seriesTCN on the spatiotemporal attention mechanism mainly compared four cases: no spatiotemporal attention mechanism, spatiotemporal attention mechanism before FFN, spatiotemporal attention mechanism after FFN, and spatiotemporal attention mechanism inserted before and after FFN. The comparison experiment on the datasets SMD, MSL, SMAP, SWAT, and PSM is shown in Table 3.

[0055] Table 3. Ablation Experiment Results of ST-seriesTCN on Spatiotemporal Attention Mechanism As shown in Table 3, the scheme that places the spatiotemporal attention mechanism after the FFN exhibits the best overall performance: it achieves an F1 score of 82.37 on the MSL dataset, significantly higher than other configurations; it significantly improves the recall (61.33) and F1 score (72.86) on the SMAP dataset, effectively compensating for its performance limitations on this dataset; the F1 score (94.46) on the SWAT dataset remains high, and the evaluation metrics for the PSM and SMD datasets also remain stable. In contrast, the absence of spatiotemporal attention or the insertion of spatiotemporal attention both before and after the FFN leads to a decline in recall or F1 score on some datasets, indicating that placing the spatiotemporal attention mechanism before the FFN has limited improvement. This demonstrates that ST-seriesTCN, through reasonable configuration of the spatiotemporal attention mechanism's position, can efficiently balance accuracy and recall, fully showcasing the superiority and design rationality of ST-seriesTCN in time series anomaly detection tasks.

[0056] The ablation experiments of ST-seriesTCN with preset normalization error anomaly thresholds were mainly selected. The ablation experiments with different preset normalization error anomaly thresholds of 0.5, 1.0, 1.5, 2.0, 2.5, and 3.0 are shown in Table 4. The detailed comparison experimental data on the datasets SMD, MSL, SMAP, SWAT, and PSM are shown in Table 4.

[0057] Table 4 Ablation Experiments of ST-seriesTCN with Preset Standardized Error Anomaly Threshold Selection As shown in Table 4, ST-seriesTCN maintains robust performance under various preset standardized error anomaly thresholds: the F1 score of the PSM and SWAT datasets is consistently maintained above 94 under multiple preset standardized error anomaly thresholds, demonstrating outstanding robustness; at the same time, it can match the optimal preset standardized error anomaly threshold for different dataset characteristics, efficiently balancing accuracy and recall, and achieving leading F1 scores on multiple datasets, fully verifying the superiority and scene adaptability of ST-seriesTCN in temporal anomaly detection tasks.

[0058] The ablation experiments of ST-seriesTCN with different FFNRatios of 1, 2, 3, and 4 are mainly selected. The detailed comparative experimental data on the datasets SMD, MSL, SMAP, SWAT, and PSM are shown in Table 5.

[0059] Table 5 Ablation Experiment Results of ST-seriesTCN with FFNRatio Selection As shown in Table 5, the performance of ST-seriesTCN varies significantly with changes in FFNRatio: When FFNRatio=1, the model performs best on most datasets: SMD achieves the highest F1 score of 83.58; PSM leads in both F1 score (97.35) and recall (97.50), while maintaining high accuracy, demonstrating efficient feature extraction capabilities under a simplified FFN structure. As FFNRatio increases to 2-4, SMD's accuracy drops sharply from 86.89 to 56.24, and its F1 score continues to decline; MSL's F1 score rises slightly to 81.81; SMAP's recall improves slightly, but its F1 score fluctuates less; SWAT and PSM consistently maintain high and stable evaluation metrics, demonstrating outstanding robustness. This demonstrates that ST-seriesTCN can efficiently balance feature extraction and computation efficiency when FFNRatio=1, and exhibits strong robustness on datasets such as SWaT and PSM, fully verifying the superiority of ST-seriesTCN in temporal anomaly detection tasks and its reasonable hyperparameter design.

[0060] The above are merely embodiments of the present invention. Commonly known structures and characteristics are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A multivariate time series anomaly detection method based on deep learning algorithm, characterized in that: Includes the following steps: S1. Collect multivariate time series data generated by multi-source sensors from the target monitoring system, and perform correlation analysis on the multivariate time series data to screen and obtain a training normal sample set; S2. Preprocess the normal training sample set to obtain the preprocessed normal training sample set, and perform feature embedding processing on the preprocessed normal training sample set to obtain the embedded features. S3. Input the embedded features into a multi-stage temporal convolutional network. At the beginning of each stage, downsample the embedded features to obtain downsampled features. Then, perform temporal convolution and depthwise separable convolution on the downsampled features to obtain a temporal feature representation. S3 includes: using the embedded features as input to the multi-stage temporal convolutional network to obtain input features; performing temporal downsampling on the input features using stride convolution and pooling to obtain downsampled features; performing convolution on the downsampled features using a preset temporal convolution kernel to obtain intermediate convolutional features; and performing depthwise convolution on each feature channel in the intermediate convolutional features to obtain a temporal feature representation. S4. The temporal feature representation is input into a convolutional feedforward network for nonlinear mapping to obtain nonlinear mapped features. The temporal convolutional features and nonlinear mapped features are then fused using residuals to obtain enhanced features. The enhanced features are then adaptively weighted using a spatiotemporal attention mechanism to obtain stage output features, which are then used as input for the next stage. S4 includes: inputting the temporal feature representation into a convolutional feedforward network; mapping the temporal feature representation through a first-layer one-dimensional convolutional operation and a nonlinear activation function to obtain first-layer convolutional features; processing the first-layer convolutional features through a second-layer one-dimensional convolutional operation to obtain nonlinear mapped features; fusing the temporal feature representation and nonlinear mapped features using residuals to obtain residual fused features; rearranging the residual fused features to obtain rearranged features; performing temporal attention modeling and variable attention modeling on the rearranged features to obtain temporal attention representations and variable attention representations; fusing the residual fused features, temporal attention representations, and variable attention representations using residuals to obtain stage output features, which are then used as input for the next stage. S5. Repeat S3 and S4 until the feature extraction of the preset number of stages in the multi-stage temporal convolutional network is completed, obtaining a fused feature representation. Then, the fused feature representation is reconstructed using a regression-based linear projection method to obtain a reconstructed sequence. Finally, the reconstructed sequence is denormalized to obtain a dimensional recovery sequence. In S5, the reconstructed sequence... The mathematical expression is: In the formula, This represents the fusion feature representation. This indicates the representation used to fuse features. The regression weight matrix that linearly maps to the original data space. This represents the regression bias during linear projection. The mathematical expression for the dimensional recovery sequence is: In the formula, Indicates the first reconstructed sequence The first variable Values ​​at each point in time. Indicating the dimensional recovery sequence of the first The first variable Values ​​at each point in time. Indicates the number of normalized sequences. The standard deviation of each variable Indicates the number of normalized sequences. The mean of each variable; S6. Calculate the time-by-time reconstruction error between the dimensional recovery sequence and the normalized sequence, and standardize the time-by-time reconstruction error to obtain the standardized error. Compare and analyze the standardized error with a preset standardized error anomaly threshold to obtain the anomaly detection result.

2. The multivariate time series anomaly detection method based on deep learning algorithm according to claim 1, characterized in that: S1 includes the following steps: S11. Construct a multivariate time series matrix corresponding to the multivariate time series data, wherein the multivariate time series matrix... The mathematical expression is: In the formula, Representing a multivariate time series matrix The total number of variables in the middle, Indicates the first Time series of 1 variable Indicates the first Time series of 1 variable Indicates the length of the time series. Indicates the first Time series of 1 variable belong A real space of dimensions; S12. Calculate the correlation coefficients between different variable sequences in the multivariate time series matrix, and construct a correlation coefficient matrix based on the correlation coefficients, wherein the mathematical expression of the correlation coefficient matrix is: In the formula, Represents the first element in the correlation coefficient matrix. Line 1 Column elements, Represents the first element in a multivariate time series matrix. Time series of 1 variable Represents the first element in a multivariate time series matrix. Time series of 1 variable This indicates the row index of the correlation coefficient matrix. Represents the column index of the correlation coefficient matrix. Represents the first element in a multivariate time series matrix. Time series of 1 variable With the multivariate time series matrix, the first Time series of 1 variable The function for calculating the correlation coefficient between them; S13. Using the correlation coefficient matrix as the correlation benchmark under normal operating conditions, a sliding window is used to divide the multivariate time series matrix into multiple time windows. The window correlation coefficient matrix between variables within each time window is calculated. By calculating the difference between the window correlation coefficient matrix and the correlation benchmark, and combining this with the stability analysis of the statistics of each variable within each time window, it is determined whether each time window is in a stable operating state. The time windows determined to be in a stable operating state are taken as data segments that meet the preset normal mode, and the extracted data segments are determined as the training normal sample set. The mathematical expression is: In the formula, This indicates the total number of normal segments. Indicates the number of times under normal operating conditions A time series segment.

3. The multivariate time series anomaly detection method based on deep learning algorithm according to claim 1, characterized in that: S2 includes the following steps: S21. Perform reversible instance normalization on each variable in the training normal sample set to obtain a normalized sequence, wherein the mathematical expression of the normalized sequence is: In the formula, Indicates the first training sample set Normalized data of invertible instances of each variable Indicates the first training sample set The values ​​of each variable, Indicates the first training sample set The average of the variables, Indicates the first training sample set The standard deviation of each variable ; S22. Patch the normalized sequence based on the time dimension to obtain the partitioned normalized sequence. The mathematical expression of the partitioned normalized sequence is: In the formula, Indicates the number of normalized sequences. The set of time segments obtained by dividing each variable into Patch segments. This represents the total number of time segments obtained after the division. , Indicates the length of the normalized sequence. Indicates the fill length. Indicates the length of a time segment. Indicates the sliding step size. , Indicates the number of normalized sequences. The variable obtained after being partitioned by Patch is the first... A time segment; S23. Embedding features are obtained by performing embedding mapping on the partitioned normalized sequence through one-dimensional convolution. The mathematical expression of the embedding features is as follows: In the formula, Indicates the number of normalized sequences after partitioning. Embedded features obtained by embedding each variable through embedding mapping Indicates the embedding dimension. This represents a one-dimensional convolution weight matrix. Represents a one-dimensional convolution weight matrix Based on the length of the time segment For lines, in embedded dimensions A real matrix of columns, This indicates that based on the one-dimensional convolution weight matrix For input Perform a one-dimensional convolution operation.

4. The multivariate time series anomaly detection method based on deep learning algorithm according to claim 1, characterized in that: S3 includes the following steps: S31. The embedded features are used as input to a multi-stage temporal convolutional network to obtain input features. Then, the input features are downsampled in the temporal dimension using stride convolution and pooling to obtain downsampled features. The mathematical expression for temporal downsampling is: In the formula, In a multi-stage temporal convolutional network, the first... At the start of the phase, the result is obtained by downsampling the features from the previous phase. This indicates a downsampling operation. In a multi-stage temporal convolutional network, the first... Characteristics of stage output; S32. Perform convolution operations on the downsampled features using a preset temporal convolution kernel to obtain intermediate convolution features. The mathematical expression for the intermediate convolution features is: In the formula, In a multi-stage temporal convolutional network, the first... Intermediate convolutional features of staged temporal convolution. This represents a temporal convolution operation. In a multi-stage temporal convolutional network, the first... The kernel size of the stage; S33. Perform depthwise convolution on each feature channel in the intermediate convolutional features to obtain the temporal feature representation. The mathematical expression for depthwise convolution is: In the formula, Indicates the first intermediate convolutional feature The first feature channel Temporal characteristics at time point This represents the index of a single feature channel in the intermediate convolutional features. Indicates the time step index. This represents the offset index of the depthwise convolution kernel. In a multi-stage temporal convolutional network, the first... Phase, First The depthwise convolution kernel corresponding to the n feature channels is at the nth feature channel. The parameter for each offset position. Indicates intermediate convolutional features The Middle Each feature channel and time position is eigenvalues.

5. The multivariate time series anomaly detection method based on deep learning algorithm according to claim 1, characterized in that: S4 includes the following steps: S41. Input the temporal feature representation into a convolutional feedforward network, and map the temporal feature representation through a first-layer one-dimensional convolution operation and a non-linear activation function to obtain the first-layer convolutional features. The mathematical expression of the first-layer convolutional features is as follows: In the formula, In a multi-stage temporal convolutional network, the first... The first layer of convolutional features in the stage, In a stage-series temporal convolutional network, the first... The temporal characteristics of the stage are represented. This represents the first layer of one-dimensional convolution operation. Represents a nonlinear activation function; S42. The first-layer convolutional features are processed by a second-layer one-dimensional convolutional operation to obtain nonlinear mapping features, wherein the mathematical expression of the nonlinear mapping features is: In the formula, In a multi-stage temporal convolutional network, the first... Nonlinear mapping characteristics of the stage This represents the second-layer one-dimensional convolution operation; S43. By fusing the time-series feature representation and the nonlinear mapping feature through residuals, the residual fusion feature is obtained. The mathematical expression of the residual fusion feature is as follows: In the formula, In a stage-series temporal convolutional network, the first... Residual fusion characteristics of the stage; S44. The residual fusion features are rearranged to obtain rearranged features. Temporal attention modeling and variable attention modeling are then performed on the rearranged features to obtain temporal attention representation and variable attention representation, respectively. The mathematical expression for the temporal attention representation is: In the formula, In a multi-stage temporal convolutional network, the first... Phase-based time attention representation, This represents the total number of steps in time. Indicates the first Feature vectors at each time step Indicates the first Attention weights at each time step , Indicates the first Attention score at each time step Indicates the first Attention score at each time step Represents an exponential function. , Indicates the first Transpose of the learnable row vectors at each time step This represents the hyperbolic tangent activation function. Indicates the use of the first Feature vectors at each time step The weight matrix to be adjusted Indicates the first Learnable bias vectors at each time step; The mathematical expression for variable attention is: In the formula, In a multi-stage temporal convolutional network, the first... Stage-specific attention representation, This represents the total number of feature dimensions. In a multi-stage temporal convolutional network, the first... Phase 1 Each feature dimension Indicates the first Phase 1 Attention weights for each feature dimension. , Indicates the first Attention weights for each feature dimension. Indicates the first Attention weights for each feature dimension. , Indicates the first The transpose of the learnable row vectors used to compute variable attention during the phase. This indicates the term used for the first stage in a multi-stage temporal convolutional network. Phase 1 Each feature dimension The weight matrix to be adjusted In a multi-stage temporal convolutional network, the first... The learnable bias vector for each stage; S45. The residual fusion features, temporal attention representation, and variable attention representation are fused using a residual method to obtain the stage output features, which are then used as the input for the next stage. The mathematical expression for the stage output features is as follows: , In a multi-stage temporal convolutional network, the first... The stage output characteristics of the stage.

6. The multivariate time series anomaly detection method based on deep learning algorithm according to claim 1, characterized in that: S6 includes the following steps: S61. Calculate the time-by-time reconstruction error of the dimensional recovery sequence and the normalized sequence at each time point in all variable dimensions, wherein the time-by-time reconstruction error is calculated by any one of the first preset reconstruction error calculation method and the second preset reconstruction error calculation method. S62. Based on the mean and standard deviation of the reconstruction error sequence, each reconstruction error in the reconstruction error sequence is standardized to obtain a standardized error sequence, wherein the mathematical expression of the standardization process is: In the formula, In the reconstruction error sequence, the first... Reconstruction error at each time point In the reconstruction error sequence, the first... The standard error of the reconstruction error at each time point after standardization. This represents the mean of the reconstructed error sequence. This represents the standard deviation of the reconstructed error sequence; S63. Compare each standardized error in the standardized error sequence with a preset standardized error anomaly threshold, and determine the time point corresponding to the standardized error that is greater than the preset standardized error anomaly threshold as an anomaly.

7. The multivariate time series anomaly detection method based on deep learning algorithm according to claim 6, characterized in that: In step S61, the mathematical expression for the first preset reconstruction error calculation method is: In the formula, This indicates the first preset reconstruction error calculation method used to calculate the error. The time-by-time reconstruction error at each point in time. Indicates the number of normalized sequences. The first variable Values ​​at each point in time. This represents the total number of variables involved in calculating the time-by-time reconstruction error; The mathematical expression for the second preset reconstruction error calculation method is: In the formula, This indicates the result obtained using the second preset reconstruction error calculation method. The time-by-time reconstruction error at each time point.

Citation Information

Patent Citations

  • A method for anomaly detection in multidimensional time series data

    CN116361635B

  • Command and control system resource trend prediction method based on fusion of long and short time sequence characteristics

    CN121412907A

  • Time sequence anomaly detection method based on PatchTST dual reconstruction consistency constraint

    CN121637381A