Multi-sensor data anomaly detection method and device based on spatio-temporal information fusion
Patent Information
- Application Number
- CN202311078583.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-24
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-08-24
AI Technical Summary
然而上述两类方法均未主动构建时空交互特征,忽略了时空交互特征在时序数据中起到的重要作用
[0052] First, the multi-sensor data anomaly detection method and apparatus based on spatiotemporal information fusion of the present invention fully considers the interaction between time and space information, constructs spatiotemporal interaction features through two sets of parallel encoders and decoders, captures spatiotemporal dependencies through an interactive attention mechanism, and realizes spatiotemporal fusion with cross-feature interaction.
Smart Images

Figure CN117540333B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and computer technology, and specifically to a method and apparatus for detecting anomalies in multi-sensor data based on spatiotemporal information fusion. Background Technology
[0002] Time series refers to a series of data that contains results that change over time. Time series anomalies are subsequences that differ from the normal pattern within the time series context. The task of time series anomaly detection is to identify abnormal events or behaviors from a normal time series. Currently, most mainstream anomaly detection methods are based on deep learning. In real-world production environments—such as water treatment data collection and soil sample collection—multiple sensors are involved, and time series data contains features across multiple dimensions. Therefore, current deep learning methods mainly focus on multivariate time series anomaly detection.
[0003] Due to the high cost of acquiring anomaly labels, anomaly detection mostly employs unsupervised learning methods. Unsupervised deep anomaly detection techniques detect outliers by learning the inherent characteristics of the data. These methods can be divided into prediction-based and reconstruction-based anomaly detection methods. Prediction-based anomaly detection methods obtain prediction errors by comparing predicted and true values. When the prediction error exceeds a selected threshold, it is considered an anomaly; otherwise, it is considered normal. Most common RNN-based anomaly detection methods predict anomalies in this way. However, due to the complex periodicity of multidimensional time series and their susceptibility to perturbations, prediction-based anomaly detection methods have a high false positive rate. Reconstruction-based anomaly detection algorithms determine anomalies by using reconstruction errors. Autoencoders (AEs) are the most common reconstruction models in anomaly detection. An autoencoder is a neural network model that encodes and decodes data. Through an encoding-decoding reconstruction operation, it learns the feature distribution of normal data to detect anomalies. Another commonly used reconstruction model is based on Generative Adversarial Networks (GANs). A GAN contains a discriminator and a generator. The generator produces reconstructed data that closely resembles the original data, while the discriminator needs to distinguish between the original and generated data as much as possible. The two networks learn from each other. Anomaly detection methods based on reconstruction learn the latent distribution of normal time series and determine anomalies by calculating the error between the reconstructed value and the true value of the test data. Therefore, this invention chooses a reconstruction-based method for anomaly detection.
[0004] Early multivariate time series anomaly detection methods focused on feature extraction from temporal dependencies. In recent years, with the advent of novel deep neural networks, research on anomaly detection based on spatiotemporal feature dependencies has been launched, primarily focusing on cascaded and parallel approaches for spatiotemporal feature extraction. Parallel extraction methods typically use two feature extraction modules to capture temporal and spatial dependencies separately, and then use fully connected layers to concatenate the spatiotemporal features. Cascaded extraction methods capture both types of dependencies by stacking spatiotemporal feature extraction modules. In real-world production environments, sensor monitoring indicators are often interconnected—for example, soil moisture content in neighboring areas is highly correlated—and the corresponding sensor data also exhibits correlated change patterns over time within a specific spatiotemporal context. Therefore, in addition to temporal and spatial correlations, multi-sensor time series data also possess spatiotemporal interaction features, representing the correspondence between time and space. Capturing these dependencies is crucial for learning time series data. However, neither of the aforementioned methods actively constructs spatiotemporal interaction features, neglecting their important role in time series data. Summary of the Invention
[0005] The purpose of this invention is to propose a multi-sensor data anomaly detection method and device based on spatiotemporal information fusion. It fully considers the interaction between time and space information, constructs spatiotemporal interactive features through two sets of parallel encoders and decoders, and captures spatiotemporal dependencies through an interactive attention mechanism, thereby realizing spatiotemporal fusion with cross-feature interaction and improving the accuracy of sensor data anomaly detection.
[0006] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows:
[0007] A method for detecting anomalies in multi-sensor data based on spatiotemporal information fusion, the method comprising the following steps:
[0008] Step S1: Use multi-scale convolutional attention to extract spatiotemporal features from multi-sensor time-series data;
[0009] Step S2: Use cross-attention to associate complementary feature matrices to perform deep fusion of spatiotemporal features of multi-sensor time series data;
[0010] Step S3: Aggregate spatiotemporal interaction features from the sensor perspective using channel attention, and construct attention weights for the spatiotemporal interaction features through a fully connected layer;
[0011] Step S4: Reconstruct multi-sensor time-series data. By calculating the error between the reconstructed multi-sensor time-series data and the real multi-sensor time-series data, determine whether the multi-sensor time-series data contains abnormal data.
[0012] Furthermore, in step S1, the process of extracting spatiotemporal features from multi-sensor time-series data using multi-scale convolutional attention includes the following steps: using convolutional kernels of different scales to capture the global and local temporal correlations of multi-sensor data, and obtaining the correlations between different sensors through standard pointwise attention.
[0013] Furthermore, in step S1, the process of extracting spatiotemporal features from multi-sensor time-series data using multi-scale convolutional attention includes the following steps:
[0014] Step S11: Represent the time series data of multiple sensors, which consists of multiple univariate time series within the same entity, as X = {x1, x2, ..., x...} n}, where n represents the maximum length of the timestamp, x t ∈R d Let X be a d-dimensional vector representing the value at time t in the time series data, where d represents the sensor dimension size and t = 1, 2, ..., n; and let X be divided into multiple sliding windows using a sliding window method.
[0015] Step S12: The temporal encoder uses causal convolution to calculate the query matrix, key matrix, and value matrix, sets convolution kernels of different sizes to learn the local context of the sensor data, and captures the dependencies of features at the sequence level through efficient attention.
[0016]
[0017] V t =FFN(V′) t )
[0018] Among them, V′ t The expression represents the computational result of convolutional attention. `Conv_Attention` represents the convolutional attention operation, which includes causal convolution and effective attention. `F` represents the result of the convolutional attention computation. o This indicates a fully connected layer; Concat represents the concatenation operation. This represents the attention outputs of different time-encoded groups H. This yields the h-th attention result in the time-encoded sequence, where h = 1, 2, ..., H. It is the time-encoded query matrix, key matrix, and value matrix obtained through weight matrix calculation, ρ q and ρ k All are softmax normalization functions; X t V represents the value of time series data X in the t-th time window. t This represents the calculation result of the time encoder, and FFN represents the feedforward network;
[0019] Step S13: The spatial encoder sets the convolution kernel size to 1 and uses a standard linear mapping to learn the correlation between sensors point by point.
[0020]
[0021] V s =FFN(V′) s )
[0022] Among them, V′ s This represents the result of the convolutional attention calculation, where Conv_Attention represents the convolutional attention operation. This indicates the attention outputs of H different groups that obtain spatial encoding. This yields the h-th attention result for spatial encoding. It is the query matrix, key matrix, and value matrix of the spatial encoding obtained by calculating the weight matrix; V represents the transpose of the time series data X at time t, used to calculate the correlation between sensors. s This represents the computational result of the spatial encoder, and FFN represents the feedforward network.
[0023] Furthermore, in step S2, a temporal decoder and a spatial decoder are used, and the feature matrix of complementary features associated with cross-attention is used to perform deep fusion of the spatiotemporal information of multi-sensor time series data. The fusion process includes two stages. The first stage includes two operations: convolutional attention and cross-attention. Convolutional attention extracts the spatiotemporal features of the previous time window, and cross-attention is used to initially fuse the spatiotemporal features. The second stage includes two cross-attention operations, which respectively aggregate the spatiotemporal encoded features of the previous stage and complementaryly decode the spatiotemporal features to further fuse the spatiotemporal information of the time series data.
[0024] Further, in step S2, the process of using cross-attention to associate complementary feature matrices to deeply fuse the spatiotemporal features of multi-sensor time-series data includes the following steps: The input to the time decoder includes two parts: the previous time window and the time-encoded features obtained in step S1. In the first stage, spatiotemporal information of the previous timestamp is captured through convolutional attention and cross-attention, and spatiotemporal interaction features are initially extracted. In the second stage, cross-attention is used to fuse the time-encoded features obtained in step S1 and the initially extracted spatiotemporal interaction features to further construct the spatiotemporal interaction features. The calculation formula for the two-stage fusion of the time decoder is as follows:
[0025]
[0026]
[0027] in, X represents the first-stage output of the time decoder. t-1 This represents the time series data from the previous time window. This represents the output of the first stage of convolutional attention. This represents the cross-attention output of the first stage, Concat represents the concatenation operation, and F... o Indicates a fully connected network. This represents the second-stage output of the time decoder. and These represent the two cross-attention outputs of the second stage;
[0028] The input to the spatial decoder consists of two parts: the spatial coded features obtained from the previous time window and the spatial coded features obtained in step S1. While independently decoding the spatial features, the temporal coded features are complementaryly acquired. The calculation formula for the two-stage fusion of the spatial decoder is as follows:
[0029]
[0030]
[0031] in, This represents the first-stage output of the time decoder. This represents the transpose of the time series data from the previous time window. This represents the output of the first stage of convolutional attention. This represents the cross-attention output of the first stage, Concat represents the concatenation operation, and F... o Indicates a fully connected network. This represents the second-stage output of the time decoder. and These represent the two cross-attention outputs of the second stage.
[0032] Further, in step S3, channel attention is used to aggregate spatiotemporal interaction features from the sensor perspective. Constructing attention weights for the spatiotemporal interaction features through a fully connected layer refers to performing deep fusion between spatiotemporal interaction features by assigning weights based on the sensor dimension to the merged spatiotemporal decoded features, including the following steps:
[0033] S31: Compress global information into a single channel descriptor over time. By outputting the signal of each channel in the data, the network can focus more on extracting feature dependencies.
[0034] z j =F sq (x j )
[0035] Among them, z j Let F represent the j-th element of the compressed embedded representation, where j∈{1,...,d}.sq Indicates compression operation, x j V represents the merged spatiotemporal decoding feature I The vector of the j-th feature;
[0036] S32: Attention weights for different sensor features are obtained through two fully connected layers, and the transformation of the feature vector is adjusted based on these weights to obtain weighted features:
[0037] s = F ex (z,F o )
[0038] (f′) j =s j ·x j
[0039] Where s represents the attention vector with different feature weights, F ex Let z represent the excitation operation, z represent the compressed embedding representation, and F represent the compression operation. o It includes two fully connected layer operations, (f′). j It is the j-th vector of the weighted features after attention weight transformation, j∈{1,...,d}, s j It is the j-th element of the attention vector, x j V represents the merged spatiotemporal decoding feature I The vector of the j-th feature;
[0040] S33: Initial spatiotemporal decoding features V after splicing and merging using residual networks I The output of the multi-sensor time-series data feature fusion stage based on channel attention is obtained.
[0041] Further, in step S4, the process of reconstructing the multi-sensor time-series data and determining whether the multi-sensor time-series data contains abnormal data by calculating the error between the reconstructed multi-sensor time-series data and the real multi-sensor time-series data includes the following steps:
[0042] The following formula is used to calculate the error between the reconstructed value and the actual value of each spatiotemporal interaction feature in the multi-sensor time series data, and the average value of d spatiotemporal interaction features is obtained and used as the anomaly score for the current timestamp:
[0043]
[0044] Where score represents the calculated anomaly score, d represents the number of spatiotemporal interaction features, and x j This represents the actual value of the j-th spatiotemporal interaction feature. This represents the reconstructed value of the j-th spatiotemporal interaction feature. It is the actual value x j and reconstructed values The root mean square error between the actual value and the reconstructed value of the j-th spatiotemporal interaction feature represents the degree of deviation between the actual value and the reconstructed value. This invention also discloses a multi-sensor data anomaly detection device based on spatiotemporal information fusion, which includes a multi-scale convolutional attention spatiotemporal feature extraction module, a cross-attention cross-spatiotemporal feature decoding module, a channel attention spatiotemporal feature fusion module, and a time-series data anomaly detection module.
[0045] The multi-scale convolutional attention spatiotemporal feature extraction module is used to extract spatiotemporal features from multi-sensor time-series data using multi-scale convolutional attention.
[0046] The cross-attention spatiotemporal feature decoding module is used to use a feature matrix of complementary features associated with cross-attention to achieve deep fusion of spatiotemporal features of multi-sensor time series data;
[0047] The channel attention spatiotemporal feature fusion module is used to aggregate spatiotemporal interaction features from the sensor perspective using channel attention, and to construct attention weights for the spatiotemporal interaction features through a fully connected layer.
[0048] The anomaly detection module uses a model to reconstruct multi-sensor time-series data. By calculating the error between the reconstructed multi-sensor time-series data and the real multi-sensor time-series data, it determines whether the multi-sensor time-series data contains abnormal data.
[0049] The present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method described above.
[0050] The present invention also discloses an electronic device, which includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the steps of the method described above.
[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0052] First, the multi-sensor data anomaly detection method and apparatus based on spatiotemporal information fusion of the present invention fully considers the interaction between time and space information, constructs spatiotemporal interaction features through two sets of parallel encoders and decoders, captures spatiotemporal dependencies through an interactive attention mechanism, and realizes spatiotemporal fusion with cross-feature interaction.
[0053] Second, the multi-sensor data anomaly detection method and device based on spatiotemporal information fusion of the present invention adopts a multi-scale convolutional attention mechanism to effectively capture local contextual information while retaining the ability to extract global features, thereby achieving more effective extraction of temporal features.
[0054] Third, the multi-sensor data anomaly detection method and device based on spatiotemporal information fusion of the present invention adopts a channel attention mechanism to allocate attention weights from the feature dimension, enhances the fusion effect of spatiotemporal features, and improves the noise resistance of the model by retaining low-frequency information. Attached Figure Description
[0055] Figure 1 This is a flowchart of the multi-sensor data anomaly detection method based on spatiotemporal information fusion according to the present invention;
[0056] Figure 2 This is a framework diagram of the multi-sensor data anomaly detection method based on spatiotemporal information fusion of the present invention;
[0057] Figure 3 This is a detailed structural diagram of the convolutional attention mechanism of the present invention;
[0058] Figure 4 This is a detailed structural diagram of the channel attention mechanism of the present invention;
[0059] Figure 5 The figure shows the ablation experiment results of this invention;
[0060] Figure 6 This is a visualization of the experimental results of the present invention;
[0061] Figure 7 This is an experimental diagram illustrating the robustness analysis of the present invention. Detailed Implementation
[0062] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0063] This invention proposes a multi-sensor data anomaly detection method based on spatiotemporal information fusion, the multi-sensor data anomaly detection method comprising the following steps:
[0064] Step S1: Use multi-scale convolutional attention to extract spatiotemporal features from multi-sensor time-series data;
[0065] Step S2: Use cross-attention to associate complementary feature matrices to perform deep fusion of spatiotemporal features of multi-sensor time series data;
[0066] Step S3: Aggregate spatiotemporal interaction features from the sensor perspective using channel attention, and construct attention weights for the spatiotemporal interaction features through a fully connected layer;
[0067] Step S4: Reconstruct multi-sensor time-series data. By calculating the error between the reconstructed multi-sensor time-series data and the real multi-sensor time-series data, determine whether the multi-sensor time-series data contains abnormal data.
[0068] This invention first captures the temporal and spatial correlations of multi-sensor data at different scales through convolutional attention during the encoding stage. Then, during the decoding stage, it uses cross-attention to associate complementary feature matrices for deep fusion, constructing spatiotemporal interactive features. Finally, based on channel attention, it aggregates features from different sensors from a spatial perspective, further enhancing the fusion effect of spatiotemporal information, thereby effectively improving the accuracy of anomaly detection in sensor data.
[0069] See Figure 1 The multi-sensor data anomaly detection method based on spatiotemporal information fusion of the present invention specifically includes the following steps:
[0070] S1, Spatiotemporal feature extraction of multi-sensor time-series data based on convolutional attention.
[0071] In step S1, the process of extracting spatiotemporal features from multi-sensor time-series data using multi-scale convolutional attention includes the following steps: using convolutional kernels of different scales to capture the global and local temporal correlations of multi-sensor data, and obtaining the correlations between different sensors through standard pointwise attention.
[0072] Multi-sensor data consists of multiple univariate time series within the same entity. This invention represents the multivariate time series as X = {x1, x2, ..., x...} n}, where n represents the maximum length of the timestamp, and each x t ∈R d Both are d-dimensional vectors representing the value at time t in the time series data, where d represents the feature dimension. To facilitate model processing, a sliding window is used to divide the time series, and the multidimensional time series X is divided into multiple sliding windows as model input, represented as X. t ={x t-l+1 , ..., x t-1 x t}, where X t This represents a sliding window ending at the t-th timestamp, x t Let y represent the vector at the t-th timestamp, and l represent the length of the sliding window. The task of multivariate time series anomaly detection is to output the vector y∈R. n y represents the time length. i ∈{0,1} indicates whether the i-th timestamp is abnormal.
[0073] To enhance the robustness of the model, eliminate the adverse effects between feature dimensions of time series data, and shorten the training time of the model, we normalize each time series:
[0074]
[0075] in, This represents the preprocessed output, x t Let X be a vector representing the t-th timestamp, min(X) represent the minimum value of time series X, and max(X) represent the maximum value of time series X.
[0076] The temporal encoder uses causal convolution to compute the query matrix, key matrix, and value matrix. It sets convolutional kernels of different sizes to learn the local context of multi-sensor data across different ranges and captures feature dependencies at the sequence level through efficient attention.
[0077]
[0078] in, X represents t The result after causal convolution operation This represents a causal convolution operation, where m is the size of the convolution kernel, and X... t This represents the value of time series data X in the t-th time window;
[0079]
[0080]
[0081]
[0082] V t =FFN(V′) t )
[0083] in, It is a time-encoded query matrix, key matrix, and value matrix obtained through weight matrix calculation. X represents t The result after causal convolution operation It is a learnable weight matrix. This yields the h-th attention result in the time-encoded sequence, ρ. q and ρ k Both are softmax normalization functions, V′ t F represents the embedding representation of aggregating multiple attention outputs. o This indicates a fully connected layer; Concat represents the concatenation operation. Let V' represent the different attention outputs of group H, and denote the calculation of convolutional attention in the temporal encoding part as V′. t =Conv_Attention(X) t V t This represents the calculation result of the time encoder, and FFN represents the feedforward network.
[0084] The spatial encoder sets the kernel size to 1 to learn spatial relevance point-by-point using a standard linear mapping:
[0085]
[0086] in, express The result after causal convolution operation This represents a causal convolution operation, where m is the size of the convolution kernel, which is set to 1 here. This represents the transpose of time series data X at the t-th time window;
[0087]
[0088]
[0089]
[0090] V s =FFN(V′) s )
[0091] in, It is a spatially encoded query matrix, key matrix, and value matrix obtained through weight matrix calculation. express The result after causal convolution operation It is a learnable weight matrix. This yields the h-th attention result for spatial encoding, ρ q and ρ k Both are softmax normalization functions, V′ s F represents the embedding representation of aggregating multiple attention outputs. o This indicates a fully connected layer; Concat represents the concatenation operation. Representing the H different attention outputs, the computation of the convolutional attention in the spatial encoding part is denoted as... V s This represents the computational result of the spatial encoder, and FFN represents the feedforward network.
[0092] S2, cross-temporal feature decoding based on cross-attention.
[0093] In step S2, a temporal decoder and a spatial decoder are used, and a feature matrix of complementary features associated with cross-attention is used to perform deep fusion of the spatiotemporal information of multi-sensor time series data. The fusion process includes two stages. The first stage includes two operations: convolutional attention and cross-attention. Convolutional attention extracts the spatiotemporal features of the previous time window, and cross-attention is used to initially fuse the spatiotemporal features. The second stage includes two cross-attention operations, which respectively aggregate the spatiotemporal encoded features of the previous stage and complementaryly decode the spatiotemporal features to further fuse the spatiotemporal information of the time series data.
[0094] See Figure 3 This invention, to fuse spatiotemporal features, designs two parallel Transformer decoders to process spatiotemporal features separately, and uses interactive attention to associate the Q, K, and V of complementary features for deep fusion. The input of the temporal decoder includes two parts: the previous time window and the temporal encoded features obtained in the previous step. In the first stage, spatiotemporal information of the previous timestamp is captured through convolutional attention and cross-attention, and spatiotemporal interactive features are initially extracted.
[0095]
[0096] in, This represents the output of the temporal decoding convolutional attention operation, where Conv_Attention represents the convolutional attention operation, and X... t-1 This represents the data from the previous time window.
[0097] The cross-attention calculation in the first stage of temporal decoding is as follows:
[0098]
[0099]
[0100]
[0101]
[0102] in, It is the time-decoded query matrix, key matrix, and value matrix obtained through weight matrix calculation, X t-1 This represents the time window preceding the input for time decoding. This represents the time window preceding the spatial decoding input. It is a learnable weight matrix. This is the h-th attention result obtained from the temporal encoded cross-attention, ρ q and ρ k All are softmax normalization functions. F represents the embedding representation of aggregating multiple attention outputs. o This indicates a fully connected layer; Concat represents the concatenation operation. Representing the different attention outputs of group H, the cross-attention calculation in the first stage of temporal decoding is denoted as... This represents the computation result of the time decoder, and FFN represents the feedforward network.
[0103] In the second stage, cross-attention is used to fuse encoded features with information from the previous stage to further construct spatiotemporal interaction features.
[0104]
[0105]
[0106]
[0107] in, This represents the cross-attention output of the aggregated temporally encoded features, where Cross_Attention represents the cross-attention operation. V represents the first-stage output of the time decoder. t This represents the time-coded features of the previous stage. This represents the cross-attention output of the interaction space decoder. This represents the output of the spatial decoder in the first stage. F represents the second-stage output of the time decoder. o This indicates a fully connected network, and Concat represents a concatenation operation.
[0108] The spatial decoder's input includes two parts: the spatial encoding features from the previous time window and the spatial features obtained in the previous step. While independently decoding the spatial features, it complementaryly acquires the temporal decoding features. In the first stage, convolutional attention and cross-attention are used to capture the spatiotemporal information of the previous timestamp, and initial spatiotemporal interaction features are extracted.
[0109]
[0110] in, This represents the output of the spatial decoding convolutional attention operation, where conv_Attention represents the convolutional attention operation. This represents the data from the previous time window.
[0111] The cross-attention calculation in the first stage of spatial decoding is as follows:
[0112]
[0113]
[0114]
[0115]
[0116] in, It is the query matrix, key matrix, and value matrix obtained through spatial decoding by calculating the weight matrix. X represents the previous time window of the spatial decoding input. t-1 This represents the time window preceding the input for time decoding. It is a learnable weight matrix. This is the h-th attention result obtained from spatial encoding cross-attention, ρ q and ρ k All are softmax normalization functions. F represents the embedding representation of aggregating multiple attention outputs. o This indicates a fully connected layer; Concat represents the concatenation operation. Representing the different attention outputs of group H, the cross-attention calculation in the first stage of spatial decoding is denoted as... This represents the computation result of the spatial decoder, and FFN represents the feedforward network.
[0117] In the second stage, cross-attention is used to fuse encoded features with information from the previous stage to further construct spatiotemporal interaction features.
[0118]
[0119]
[0120]
[0121] in, Cross_Attention represents the cross-attention output of the aggregated spatial encoded features. V represents the first-stage output of the spatial decoder. s This represents the spatial coding features of the previous stage. This represents the cross-attention output of the interaction space decoder. This represents the output of the spatial decoder in the first stage. F represents the second-stage output of the time decoder. o This indicates a fully connected network, and Concat represents a concatenation operation.
[0122] S3, multi-sensor time-series data feature fusion based on channel attention.
[0123] See Figure 4 , with the merged feature data V I As input, the channel attention mechanism first performs a compression operation, compressing global information into a single channel descriptor. By outputting the signal of each channel in the data, the network focuses more on extracting feature dependencies. Formally, the compressed generation z is obtained by shrinking over the time dimension, where the j-th element of z is calculated using the following formula:
[0124]
[0125] Among them, z jLet F represent the j-th element of the compressed embedded representation, where j∈{1,...,d}. sq Indicates compression operation, x j V represents the merged spatiotemporal decoding feature I The vector of the j-th feature, where n represents the length of the time window, x j (i) represents the value of the j-th feature at the i-th timestamp.
[0126] The activation operation first obtains attention weights for different features through two fully connected layers, and then adjusts the transformation of the feature vector based on these weights to obtain weighted features.
[0127]
[0128] (f′) j =s j ·x j
[0129] Where s represents the attention vector with different feature weights, F ex Let z represent the excitation operation, z represent the compressed embedding representation, and F represent the compression operation. o This represents a fully connected layer, where σ represents the Sigmoid function. Represents the ReLU function. and There are two fully connected layers. (f′) j It is the j-th vector of the weighted features after attention weight transformation, j∈{1,...,d}, s j It is the j-th element of the attention vector, x j V represents the merged spatiotemporal decoding feature I The vector of the j-th feature.
[0130] The initial spatiotemporal decoding feature V is obtained by splicing and merging residual networks. I The output of the multi-sensor time-series data feature fusion stage based on channel attention is obtained.
[0131] f = F o (Concat(V I ,f′))
[0132] Where f represents the output of the spatiotemporal feature fusion stage, F o This indicates a fully connected layer; Concat represents the concatenation operation; V I f' represents the initial spatiotemporal decoding features, f' represents the weighted features after attention weight transformation, and the final output of the interactive decoder is generated by the feedforward network.
[0133]
[0134] in, denoted as the final output of the interactive decoder, FFN represents the feedforward network, and f represents the output of the spatiotemporal feature fusion stage.
[0135] S4, anomaly detection based on reconstruction error of multi-sensor time series data.
[0136] Anomaly detection based on reconstruction error determines the likelihood of an anomaly at a given point by calculating the error between the reconstructed data and the actual data. The model first calculates the error between the reconstructed value and the actual value for each feature, and then takes the average of d features.
[0137]
[0138] Where score represents the calculated anomaly score, d represents the number of features, and x j Indicates actual value and Indicates the reconstructed value. It is the actual value x j and reconstructed values The root mean square error between the actual value and the reconstructed value of the j-th feature represents the degree of deviation between them. An inference score for each timestamp is obtained by calculating the anomaly score for each feature and averaging these scores. This invention uses a point adjustment metric; if any point within a segment is detected, the segment is considered an anomalous. We automatically select a threshold using the Peak Exceeds Threshold (POT) calculation method; if the score calculated above is greater than the threshold, we mark the timestamp as an anomaly.
[0139] experiment
[0140] The experiment consisted of two steps. The first step was to train a multi-sensor data anomaly detection model based on spatiotemporal information fusion. The second step was to use the learned model to perform anomaly detection on a test set.
[0141] During training, Adam was chosen as the training optimization algorithm, with an initial learning rate of 0.001, a sliding window length of 100, 30 training iterations, and a batch size of 32. First, the multidimensional time series X was divided into sliding windows as model input. Then, the sliding windows were input into a multi-scale convolutional attention spatiotemporal feature extraction module to obtain multi-scale temporal and spatial features of the time series. Cross-attention was used for cross-spatiotemporal feature fusion, and channel attention was used to further fuse spatiotemporal features. The model's reconstruction error was calculated as the model loss function, and the overall network was trained by backpropagation by minimizing the overall loss function.
[0142] During the detection phase, a trained spatiotemporal information fusion multi-sensor data anomaly detection model is used to calculate the anomaly score for each timestamp in the test set, thereby performing anomaly detection.
[0143] At this point, the multi-sensor data anomaly detection of this invention has been calculated. All experiments were conducted on a server running Windows 10 (64-bit), equipped with an NVIDIA GeForce GTX 1060 graphics processing unit (GPU) and 16GB of memory. PyTorch and Python were used for implementation. To evaluate the effectiveness of this invention on multi-sensor data, six publicly available datasets—PSM, SWAT, WADI, SMAP, MSL, and EPSY—were used for testing. The performance of this invention was compared with benchmark methods LSTM-AD, DAGMM, OmniAnomaly, MSCRED, MAD-GAN, USAD, MTAD-GAT, GDN, TranAD, and DTAAD.
[0144] Performance evaluation is mainly based on three metrics: precision, recall, and F1 score.
[0145] (1) Precision. This indicates the proportion of true anomalies among the detected anomalies.
[0146] (2) Recall. Represents the proportion of all true anomalies that are marked as anomalies by the model.
[0147] (3) F1-score. F1-score is a performance metric that takes into account both precision and recall.
[0148]
[0149] The primary objective is to verify whether the anomaly detection of multi-sensor data extracted in this invention is related to functional and modular independence. The evaluation metrics for the test results are mainly precision, recall, and F1-score. The experimental performance comparison with other methods is shown in Table 2. This invention achieves very promising results on all datasets, validating its advantages on multi-sensor datasets.
[0150] Table 1 Comparison of experimental results
[0151]
[0152]
[0153] Table 1 shows the experimental data of the method of this invention compared with ten other methods. It can be seen that the method of this invention achieved the highest F1 score on the four multi-sensor datasets PSM, WADI, SMAP, and EPSY, especially improving the F1 score by 17.32% on the high-dimensional time-series dataset WADI, demonstrating the effectiveness of the method. LSTM-AD fully utilizes the long-term memory capability of LSTM for data prediction, while DAGMM combines a deep autoencoder and a Gaussian mixture model to generate low-dimensional features. These two methods have relatively simple feature extraction methods and achieved poor results on multiple datasets. However, both methods achieved excellent results on a small-scale dataset like EPSY, due to overfitting. In contrast, other more complex models have not yet fully learned the features of the data. With the deep fusion of spatiotemporal features achieved through the interactive decoding module, this invention can more quickly capture the inherent spatiotemporal relationships of the data and also performs well in reconstruction tasks on small-scale datasets. MSCRED is a method that uses reconstruction to distinguish anomalies, which leads to inaccurate identification of subtle anomalies in time-series data, achieving relatively average results on multiple datasets. MAD-GAN uses an LSTM-RNN model simultaneously in the generator and discriminator to capture temporal relationships. However, its performance on multiple datasets is mediocre because it relies heavily on the sliding window size, easily overlooking contextual dependencies. This invention avoids these problems by using multi-scale convolutional attention. By setting a larger sliding window and multiple sets of different convolutional kernels and strides, the model takes into account both point information and contextual information extraction, maximizing the fit to the training data while reducing sensitivity to the sliding window size. Regarding USAD, MTAD-GAT, and GDN, USAD is based on an autoencoder and amplifies small anomalies through adversarial training; MTAD-GAT introduces graph attention mechanisms in the temporal and spatial dimensions to learn more complex dependencies; and GDN is a method that uses graph structures to learn multivariate time series, explicitly modeling the correlations between time series features. All three methods have achieved good experimental results due to their respective advantages. However, USAD neglects spatial correlation, a point that has proven quite important in several works; conversely, GDN overemphasizes spatial features, neglecting temporal correlation; MTAD-GAT learns spatiotemporal features simultaneously, but only performs simple concatenation, ignoring the extraction of potential spatiotemporal dependencies. This invention interactively fuses spatiotemporal information during the decoding stage, relying on Q, K, and V matrices to construct spatiotemporal relationships in greater detail. This interaction between the two stages allows the decoder to simultaneously focus on the spatiotemporal dependencies of the current and previous time windows, enhancing the model's ability to learn normal data distributions and thus improving anomaly detection performance. The introduction of channel attention further improves the model's fusion efficiency, overcoming the strong randomness inherent in general fully connected layers.TranAD amplifies errors through adversarial training, enabling it to detect subtle anomalies. DTAAD introduces global and local TCNs to address the issue of diverse dependencies in temporal data. Both methods demonstrate superior performance on multiple datasets, illustrating the effectiveness of Transformer and global-local information extraction in learning temporal data. Overall, this invention still outperforms TranAD and DTAAD on eight datasets. Compared to ten typical models, this invention exhibits a more comprehensive learning ability for temporal data in various aspects, and its superior performance on multiple datasets proves the effectiveness of this method.
[0154] To verify the effectiveness of the key modules of STTD, this section will conduct ablation experiments on three multi-sensor datasets: SMAP, MSL, and PSM. This invention designs three variants of STTD, removing multi-scale convolutional attention, channel attention, and interactive attention respectively. The three variant models of STTD are described below:
[0155] (1) Model 1 - STTD model without interactive attention: The model ignores the interaction of spatiotemporal information in the decoding stage and relies only on two parallel decoders.
[0156] (2) Model 2 - STTD model without convolutional attention: Ignoring the extraction of multi-scale features, the model only relies on standard point-by-point attention to learn global features.
[0157] (3) Model 3 - STTD model without channel attention: The model ignores the extraction of frequency domain information and uses basic splicing operation in the fusion part.
[0158] The performance of the three models on the three datasets is as follows: Figure 5 As shown. Since SMAP and MSL have similar anomaly patterns, the models show similar trends on these two datasets. The interactive decoding mechanism has the greatest impact on model performance. Interactive attention decomposes spatiotemporal information into three matrices (Q, K, V) and performs interactive aggregation. Model 1, ignoring this module, saw a 4.89% decrease in performance on the MSL dataset, demonstrating the importance of joint interactive spatiotemporal information for time-series data. Convolutional attention uses convolutional kernels of different sizes to focus on feature information at different scales. Model 2, ignoring convolutional attention, saw a 2.49% decrease in performance on the SMAP dataset, indicating that aggregating local contextual information helps reconstruct performance. Channel attention can enhance feature semantic information; Model 3, ignoring this module, also experienced a performance decrease. Analysis shows that the computational complexity of the interactive attention mechanism is O(nd). 2 The computational complexity of the convolutional attention mechanism is O(nk). max d+nd 2 ), k maxThe maximum convolution kernel is represented by , and the computational complexity of the channel attention mechanism is O(nd). Overall, compared to STTD, the F1 scores of the three variant models decreased by 3.55%, 2.14%, and 1.28%, respectively, indicating that the three modules optimized the model performance to varying degrees with relatively low overhead.
[0159] To evaluate the anomaly detection performance of different methods, a comparative visualization experiment was designed. Figure 6 The reconstruction performance of the three methods on the typical server dataset PSM is compared. Figure 6 The True_value column displays the true values of a feature at different timestamps on the PSM dataset, with the background representing the true outlier segments of the data. The Rec_value column contains three images, corresponding to the reconstructed data of the true sequence by STTD, TranAD, and DTAAD, respectively. The background represents the predicted outlier segments of the corresponding method, and the different shades of the region below represent the outlier scores.
[0160] To verify the model's detection performance in noisy environments, robustness experiments were conducted. This invention simulates real-world environments by injecting Gaussian noise. We set the input noise ratios to 0.05, 0.1, and 0.25, respectively, and the experimental results are as follows: Figure 7 As shown, the F1 scores of all methods decrease with increasing signal-to-noise ratio. USAD and MTAD-GAT show more significant variations across the three datasets. In contrast, TranAD and DTAAD exhibit smaller fluctuations on the SMAP and MSL datasets, demonstrating strong noise resistance. However, on SWAT, DTAAD is significantly affected by noise. The channel attention compression operation focuses on low-frequency information in the frequency domain, which helps the model learn robust temporal features, resulting in small fluctuations across the three datasets and demonstrating strong robustness against interference.
[0161] like Figure 2 The diagram shown is a structural diagram of the device of the present invention. Figure 2 As can be seen above, the overall network architecture of this invention is divided into two stages: an encoding stage and a decoding stage. The encoder and decoder in both stages process temporal and spatial features respectively. However, in the decoding stage, the model deeply fuses spatiotemporal information based on cross-attention to capture the spatiotemporal interaction features of the data, and further enhances the fusion effect through channel attention. Specifically, this invention also discloses a multi-sensor data anomaly detection device based on spatiotemporal information fusion. This device includes a multi-scale convolutional attention spatiotemporal feature extraction module, a cross-attention cross-spatiotemporal feature decoding module, a channel attention spatiotemporal feature fusion module, and a time-series data anomaly detection module.
[0162] The multi-scale convolutional attention spatiotemporal feature extraction module is used to obtain global and local temporal correlations in the time dimension, as well as the correlations between sensors in the spatial dimension. Specifically, different fields of view are obtained by using convolutional kernels of different sizes, and then an attention matrix is calculated using an efficient attention mechanism. The spatiotemporal feature representation of the encoding stage is obtained through a feedforward network.
[0163] The cross-attention spatiotemporal feature decoding module is used to fuse the spatiotemporal information of time series data to construct spatiotemporal interactive features. The fusion process includes two stages. The first stage includes two operations: convolutional attention and cross-attention. Convolutional attention extracts the spatiotemporal features of the previous time window, and cross-attention initially fuses the spatiotemporal features. The second stage includes two cross-attention operations, which respectively aggregate the spatiotemporal encoded features of the previous stage and complementaryly decode the spatiotemporal features to further fuse the spatiotemporal information of the time series data.
[0164] The channel attention spatiotemporal feature fusion module is used to enhance the spatiotemporal feature fusion effect. Specifically, the data is compressed into channel descriptors through compression operations, then attention vectors are calculated through activation operations, and the attention score is multiplied by each feature vector. Finally, the initial features are spliced through a residual network, and the final result is obtained through a feedforward network.
[0165] The anomaly detection module uses a model to detect anomalies in the test data. It calculates the reconstruction error of the test data to determine the probability that a certain point in the time series is an anomaly, thereby completing the anomaly detection of multi-sensor data based on spatiotemporal information fusion.
[0166] In addition, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method steps of the above-described multi-sensor data anomaly detection method.
[0167] In addition, the present invention also provides an electronic device, which includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the method steps of the above-described multi-sensor data anomaly detection method.
[0168] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0169] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0170] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0171] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment, causing a series of operational steps to be executed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that run on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0172] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0173] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A multi-sensor data anomaly detection method based on spatiotemporal information fusion, characterized in that, The multi-sensor data anomaly detection method includes the following steps: Step S1: Use multi-scale convolutional attention to extract spatiotemporal features from multi-sensor time-series data; Step S2: Using a temporal decoder and a spatial decoder, the feature matrix of complementary features associated with cross-attention is used to perform deep fusion of the spatiotemporal information of multi-sensor time series data. The fusion process includes two stages. The first stage includes two operations: convolutional attention and cross-attention. Convolutional attention extracts the spatiotemporal features of the previous time window, and cross-attention is used to initially fuse the spatiotemporal features. The second stage includes two cross-attention operations, which respectively aggregate the spatiotemporal encoded features of the previous stage and complementaryly decode the spatiotemporal features to further fuse the spatiotemporal information of the time series data. Step S3: Aggregate spatiotemporal interaction features from the sensor perspective using channel attention, and construct attention weights for the spatiotemporal interaction features through a fully connected layer; Step S4: Reconstruct multi-sensor time-series data. By calculating the error between the reconstructed multi-sensor time-series data and the real multi-sensor time-series data, determine whether the multi-sensor time-series data contains abnormal data.
2. The multi-sensor data anomaly detection method based on spatiotemporal information fusion according to claim 1, characterized in that, In step S1, the process of extracting spatiotemporal features from multi-sensor time-series data using multi-scale convolutional attention includes the following steps: using convolutional kernels of different scales to capture the global and local temporal correlations of multi-sensor data, and obtaining the correlations between different sensors through standard pointwise attention.
3. The multi-sensor data anomaly detection method based on spatiotemporal information fusion according to claim 2, characterized in that, Step S1, the process of extracting spatiotemporal features from multi-sensor time-series data using multi-scale convolutional attention, includes the following steps: Step S11: Represent the time series data of multiple sensors, which consists of multiple univariate time series within the same entity, as follows: ,in, Indicates the maximum length of a timestamp. Represent a A dimensional vector representing time series data. The value at time, Indicates the sensor dimension size. Using a sliding window to divide the time series, multidimensional time series... Divided into multiple sliding windows; Step S12: The temporal encoder uses causal convolution to calculate the query matrix, key matrix, and value matrix, sets convolution kernels of different sizes to learn the local context of the sensor data, and captures the dependencies of features at the sequence level through efficient attention. in, This represents the result of the convolutional attention calculation. This represents the convolutional attention operation, which includes two parts: causal convolution and effective attention. Indicates a fully connected layer. This indicates a splicing operation. express Different groups of attention outputs with temporal encoding were obtained. , It is the first time code obtained One attention result, , It is a time-encoded query matrix, key matrix, and value matrix obtained through weight matrix calculation. and All Normalization function; Representing time series data In the The value of a time window, This represents the calculation result of the time encoder. Indicates a feedforward network; Step S13: The spatial encoder sets the convolution kernel size to 1 and uses a standard linear mapping to learn the correlation between sensors point by point. in, This represents the result of the convolutional attention calculation. This indicates the convolutional attention operation. express Different groups of attention outputs are obtained by encoding spatial codes. , It is the first spatial code obtained One attention result, It is the query matrix, key matrix, and value matrix of the spatial encoding obtained by calculating the weight matrix; Representing time series data exist The transpose of the time values is used to calculate the correlation between sensors. This represents the calculation results of the spatial encoder. This represents a feedforward network.
4. The multi-sensor data anomaly detection method based on spatiotemporal information fusion according to claim 1, characterized in that, In step S2, a feature matrix of complementary features associated with cross-attention is used to perform deep fusion of the spatiotemporal features of multi-sensor time-series data. The process includes the following steps: The input to the temporal decoder consists of two parts: the previous time window and the temporal encoding features obtained in step S1. In the first stage, spatiotemporal information of the previous timestamp is captured through convolutional attention and cross-attention, and spatiotemporal interaction features are initially extracted. In the second stage, cross-attention is used to fuse the temporal encoding features obtained in step S1 and the initially extracted spatiotemporal interaction features to further construct the spatiotemporal interaction features. The calculation formula for the two-stage fusion of the temporal decoder is as follows: in, This represents the first-stage output of the time decoder. This represents the time series data from the previous time window. This represents the output of the first stage of convolutional attention. This represents the output of the cross-attention process in the first stage. This indicates a splicing operation. Indicates a fully connected network. This represents the second-stage output of the time decoder. and These represent the two cross-attention outputs of the second stage; The input to the spatial decoder consists of two parts: the spatial coded features obtained from the previous time window and the spatial coded features obtained in step S1. While independently decoding the spatial features, the temporal coded features are acquired complementaryly. The calculation formula for the two-stage fusion of the spatial decoder is as follows: in, This represents the first-stage output of the time decoder. This represents the transpose of the time series data from the previous time window. This represents the output of the first stage of convolutional attention. This represents the output of the cross-attention process in the first stage. This indicates a splicing operation. Indicates a fully connected network. This represents the second-stage output of the time decoder. and These represent the two cross-attention outputs of the second stage.
5. The multi-sensor data anomaly detection method based on spatiotemporal information fusion according to claim 1, characterized in that, In step S3, channel attention is used to aggregate spatiotemporal interaction features from the sensor perspective. Constructing attention weights for the spatiotemporal interaction features through a fully connected layer refers to performing deep fusion between spatiotemporal interaction features by assigning weights based on the sensor dimension to the merged spatiotemporal decoded features. This includes the following steps: S31: Compress global information into a single channel descriptor over time. By outputting the signal of each channel in the data, the network can focus more on extracting feature dependencies. in, The first embedded representation after compression One element, , This indicates a compression operation. Indicates the spatiotemporal decoding features of merging The A vector of features; S32: Attention weights for different sensor features are obtained through two fully connected layers, and the transformation of the feature vector is adjusted based on these attention weights to obtain weighted features: in, Attention weights representing different sensor features Indicates an incentive operation. This represents the compressed embedding representation. It includes two fully connected layer operations. It is the first weighted feature after attention weight transformation. A vector, , It is the first of the attention weights One element, Indicates the spatiotemporal decoding features of merging The A vector of features; S33: Initial spatiotemporal decoding features obtained by splicing and merging residual networks The output of the multi-sensor time-series data feature fusion stage based on channel attention is obtained.
6. The multi-sensor data anomaly detection method based on spatiotemporal information fusion according to claim 1, characterized in that, In step S4, the process of reconstructing multi-sensor time-series data and determining whether the multi-sensor time-series data contains abnormal data by calculating the error between the reconstructed multi-sensor time-series data and the actual multi-sensor time-series data includes the following steps: The following formula is used to calculate the error between the reconstructed and actual values of each spatiotemporal interaction feature in the multi-sensor time series data, and to obtain... The average of the spatiotemporal interaction features is used as the anomaly score for the current timestamp: in, This represents the calculated anomaly score. Indicates the number of spatiotemporal interaction features. This represents the actual value of the j-th spatiotemporal interaction feature. This represents the reconstructed value of the j-th spatiotemporal interaction feature. This is the actual value. and reconstructed values The root mean square error between them represents the first... The degree of deviation between the actual value and the reconstructed value of a spatiotemporal interaction feature.
7. A multi-sensor data anomaly detection device based on spatiotemporal information fusion, characterized in that, The multi-sensor data anomaly detection device includes a multi-scale convolutional attention spatiotemporal feature extraction module, a cross-attention cross-spatiotemporal feature decoding module, a channel attention spatiotemporal feature fusion module, and a time-series data anomaly detection module; The multi-scale convolutional attention spatiotemporal feature extraction module is used to extract spatiotemporal features from multi-sensor time-series data using multi-scale convolutional attention. The cross-attention spatiotemporal feature decoding module is used to perform deep fusion of spatiotemporal information of multi-sensor time series data by employing a temporal decoder and a spatial decoder and using a feature matrix of complementary features associated by cross-attention. The fusion process includes two stages. The first stage includes two operations: convolutional attention and cross-attention. Convolutional attention extracts the spatiotemporal features of the previous time window and cross-attention initially fuses the spatiotemporal features. The second stage includes two cross-attention operations, which respectively aggregate the spatiotemporal encoded features of the previous stage and complementaryly decode the spatiotemporal features to further fuse the spatiotemporal information of the time series data. The channel attention spatiotemporal feature fusion module is used to aggregate spatiotemporal interaction features from the sensor perspective using channel attention, and to construct attention weights for the spatiotemporal interaction features through a fully connected layer. The anomaly detection module uses a model to reconstruct multi-sensor time-series data. By calculating the error between the reconstructed multi-sensor time-series data and the real multi-sensor time-series data, it determines whether the multi-sensor time-series data contains abnormal data.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method described in any one of claims 1-6.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing a computer program, which, when executed by the processor, implements the steps of the method described in any one of claims 1-6.