Communication equipment anomaly detection method based on channel attention and frequency domain analysis
By introducing channel attention and frequency domain analysis methods in timing data abnormal detection, a timing data reconstruction model was designed, which solved the problem of difficulty in capturing long-term changes and multi-dimensional correlation in the prior art, and achieved more efficient anomaly detection performance.
Patent Information
- Application Number
- CN202510150464.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-11
AI Technical Summary
When handling multi-dimensional and long-sequence time series anomaly detection, it is difficult to capture long-term changes and multi-dimensional correlations, and the lack of effective anomaly label data, resulting in poor detection results.
A time series data reconstruction model based on channel attention and frequency domain analysis was designed. Through timing decomposition, stacked convolution and channel attention mechanisms, the changes in the period and period of time series data of multi-index timing data, as well as the interdependence between different indicators.
This model can improve data reconstruction effect and abnormal detection performance in multi-dimensional and long-sequence abnormal detection scenarios, and significantly improve the abnormal detection ability of communication equipment monitoring data.
Smart Images

Figure CN120067805A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence time series anomaly detection and analysis, and particularly relates to an anomaly detection method for network communication devices applied in environments such as the Internet of Things and industrial automation. Background Art
[0002] Anomaly detection of network communication devices refers to a technology that monitors and records device performance indicators (such as latency, packet loss rate, throughput, etc.) and operating states (such as port state, connection state, etc.) in devices such as routers and switches, and identifies abnormal behaviors or anomalies from the recorded multi-index time series data. Abnormalities in communication devices may lead to problems such as performance degradation, data loss, and connection interruption. Therefore, by monitoring changes in various indicators, system failures or abnormal services can be processed in a timely manner. For communication network access providers, ensuring the availability and performance of services is crucial, and anomaly detection of communication devices is the key technology to solve these problems. Therefore, anomaly detection of network communication devices has become an important research direction in the field of information technology. Its main goal is to identify abnormal time points that are significantly different from the normal pattern of time series data and issue early warnings to management personnel in a timely manner.
[0003] Compared with other types of data, the characteristic of time series lies in its continuity. Only some scalars are saved at each time point. Therefore, a single time point usually does not provide enough semantic information for analysis. Therefore, the research focus is usually on its characteristics changing over time, which can reflect important attributes such as the continuity, periodicity, and trend of the time series, providing richer information. In the past few years, researchers have proposed many methods based on statistics and traditional machine learning. The methods based on statistics identify outliers by analyzing the statistical characteristics of time series data, such as analyzing statistical quantities such as the distribution characteristics, mean, and standard deviation of the data, as well as the threshold assumptions related to these statistical quantities for data distribution, to determine whether a data point is abnormal, such as the autoregressive integrated moving average model ARIMA (AutoRegressive Integrated Moving Average), and the local outlier factor detection method (Local Outlier Factor, LOF). The methods based on traditional machine learning such as: K-Nearest Neighbors (KNN), One-Class SVM, Extreme Learning Machine (ELM), etc., make assumptions about the potential model in the data and detect by setting data boundaries and outlier thresholds. However, the changes in time series in the real world usually have complex patterns, and their high-dimensionality and complexity make the algorithms based on mathematical statistical models and traditional machine learning no longer applicable. Therefore, new methods need to be introduced to improve the accuracy and efficiency of time series. And due to the scarcity of abnormal data, it is too expensive to obtain abnormal label data, resulting in poor performance of traditional supervised methods.
[0004] Deep learning methods benefit from their non-linear modeling ability and network structure for processing high-dimensional data, and can learn the complex dynamic changes in the data without making assumptions about the potential model in the data. Therefore, they have achieved better results compared with traditional methods. For example: models such as LSTM, Gated Recurrent Unit (GRU), Deep Auto-Encoder Gaussian Mixture Model (DAGMM), Variational Auto-Encoder (VAE), Graph Neural Network (GNN) are the key technologies of unsupervised models. In general, these unsupervised time series anomaly detection models often use prediction or reconstruction methods to identify anomalies in the data. And recent research has shown that the reconstruction-based method performs better than the prediction-based method in the field of anomaly detection.
[0005] Although these deep learning-based models have achieved excellent performance in the field of time series prediction, they usually have difficulty capturing long-term changes and multi-dimensional correlations in time series data. For example, in the method based on the graph convolutional network (GCN), its topological structure is restricted to a small number of adjacent nodes to maintain the limitation of computational complexity, making it difficult to access global information and complete detection tasks at larger time steps. Deep models based on Transformer perform poorly in scenarios of long sequence analysis, with high computational costs for long time series and it is difficult to directly find reliable dependencies from scattered time points, and their performance in common prediction tasks is even inferior to that of the single-layer linear model DLinear.
[0006] The invention patent with the publication number CN117556311A discloses an unsupervised time series anomaly detection method based on multi-dimensional feature fusion, which is applied to server anomaly detection. It focuses on multi-dimensional feature fusion, including the extraction and fusion of one-dimensional and two-dimensional features; analyzes a single cycle; uses a stacked separable convolutional network and a fully connected layer for feature fusion, parallelly uses a fully connected layer and a convolutional network, and finally performs feature fusion.
[0007] The invention patent with the publication number CN117688496A discloses an anomaly diagnosis method, system and device for satellite telemetry multi-dimensional time series data, which realizes anomaly detection and diagnosis through an improved transfer entropy method and a graph neural network, and tends to be able to diagnose the root cause of anomalies. Adopts a graph neural network, combines causal relationships and similarity relationships to construct the graph structure of the graph neural network to realize anomaly detection and diagnosis of satellite telemetry multi-dimensional time series data.
[0008] The publication number CN117851920A discloses a power Internet of Things data anomaly detection method and system, which uses discrete wavelet transform and a spatio-temporal network model to mine time series features and complex correlations between sequences. Combines a graph convolutional layer and a Transformer model, and determines the anomaly score through a multivariate Gaussian distribution and Mahalanobis distance. Summary of the Invention
[0009] The core of the present invention lies in designing a time series reconstruction model for anomaly detection, aiming to overcome the limitations of the existing technology. In the scenario of multi-dimensional and long-sequence anomaly detection, it can find reliable dependencies from scattered time points, capture the normal data patterns of communication device monitoring data, improve the data reconstruction effect, and further improve the anomaly detection performance.
[0010] A communication device anomaly detection method based on channel attention and frequency domain analysis includes the following steps:
[0011] Step S1: Extract original data from the database, and after normalization processing, obtain subsequences;
[0012] Step S2: Perform temporal decomposition on the subsequence of Step S1 to obtain a trend sequence component and a periodic sequence component;
[0013] Step S3: Perform FFT transformation and analysis on the periodic sequence component to obtain an average amplitude set and establish an input feature map;
[0014] Step S4: Apply squeeze-and-excitation to the input feature map to generate a channel attention map for adjusting the weights of each index in the input feature map;
[0015] After weighting the feature map, perform a convolution operation to obtain which can simultaneously capture intra-period and inter-period changes of multi-index time series data, as well as the interdependencies between different indices;
[0016] Step S6: Recombine and output the convolution result of Step S5 into a one-dimensional tensor output to obtain the output sequence of the stacked convolution module;
[0017] Step S7: For the output sequence of the convolution module in Step S6, learn the deep abstract temporal data features of the stacked convolution results under different periods through a multi-layer perceptron, and output the result;
[0018] Step S8: Aggregate the initial trend sequence component and weight it, output a residual sequence, and use it as the input for the next layer;
[0019] Step S9: Use a loss function to update the parameters of each layer layer by layer;
[0020] Step S10: Perform anomaly detection, reconstruct the result based on the training set and calculate the reconstruction error, and find the anomaly threshold according to the pre-input hyperparameter anomaly rate; calculate the anomaly score, and compare it with the anomaly threshold for anomaly judgment; if the anomaly score is greater than the anomaly threshold, it is determined that there is an anomaly at this time point.
[0021] Beneficial effects:
[0022] Existing time series data anomaly detection technologies face the challenges of high dimensionality and complexity of data and lack of labeled data when dealing with large-scale data. The anomaly detection model based on channel attention and frequency domain analysis proposed by the present invention has the following advantages compared with the prior art:
[0023] 1. The temporal decomposition block enhances the performance of the model in mining periodic relationships in long sequences, and improves the adaptability of the model to perform frequency domain analysis and feature extraction on long sequences under different seasonal changes.
[0024] 2. By introducing a channel attention mechanism, corresponding weights can be assigned according to the importance of each index of the time-series data, enabling the model to focus on the important features of the communication device monitoring data. This effectively suppresses the influence of noise sequences on data reconstruction, thereby improving the model's ability to reconstruct key sequences, and ultimately significantly enhancing the model's anomaly detection ability.
[0025] 3. The present invention does not rely on the statistical characteristics of the data or make assumptions about the potential model in the time-series data. Instead, it adopts a reconstruction-based anomaly detection method, maps the time-series data into a high-dimensional feature map, and models the time-series data in a deep neural network
[0026] Disperse the non-linear relationships between data points, and perform reconstruction based on high-dimensional features to capture complex characteristics such as long-term dependencies, periodic changes, and multi-dimensional correlations of time-series data. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a schematic flowchart of the communication device anomaly detection method of the present invention.
[0028] Figure 2 It is a schematic diagram of the overall architecture of the data reconstruction model of the present invention.
[0029] Figure 3 It is a schematic diagram of slicing and stacking of the present invention.
[0030] Figure 4 It is a schematic diagram of the weight assignment of the convolutional channel attention module CAM of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0031] The present invention proposes a communication device anomaly detection method based on channel attention and frequency domain analysis. This anomaly detection method is implemented based on a reconstruction model. First, the detection data generated by the network communication device is preprocessed to generate a time-series data dataset. In the model training stage, a time-series data reconstruction model is used to capture the features and distribution of normal data; in the anomaly detection stage, the data to be detected is reconstructed, the reconstructed data is compared with the input data, and the time points where the data to be detected is abnormal are found according to the reconstruction error. The specific process is as Figure 1 shown: First, divide the dataset based on historical data, train the reconstruction model to capture the features of the normal mode of the data. In the anomaly detection stage, obtain the time-series data to be detected, input it into the time-series data reconstruction model to obtain the reconstructed sequence, compare the time-series data to be detected and the reconstructed data, calculate the anomaly scores at each time point, and finally find the abnormal time points by comparing the anomaly scores of the time-series data with a preset anomaly threshold.
[0032] The present invention is directed to anomaly detection of communication devices based on reconstruction, and designs a data reconstruction model based on CNN. By performing trend cycle decomposition on the original data, slicing and stacking based on cycle characteristics, and performing two-dimensional convolution after using channel attention weight assignment, reliable dependencies based on different cycles can be found among scattered time points. Among them, the time series decomposition block, the stacking and convolution module are the main parts of the reconstruction module, constituting a block. The various blocks of the model are connected by residual connections to construct a data reconstruction model in the anomaly detection process. Figure 2 It is the overall architecture of the data reconstruction model.
[0033] The core of the present invention lies in designing a time series data reconstruction model for anomaly detection, aiming to overcome the limitations of the existing technologies. In the anomaly detection scenario of multi-dimensional and long sequences, reliable dependencies can be found among scattered time points, the normal data patterns of the communication device monitoring data can be captured, the data reconstruction effect can be improved, and thus the anomaly detection performance can be enhanced. Its core lies in: First, a stacking and convolution module is proposed in the reconstruction model, which slices and stacks the time series data based on cycle characteristics and performs convolution to capture the intra-cycle and inter-cycle changes of multi-index time series data, as well as the mutual dependencies among different indexes. Second, on this basis, in order to focus on the important indexes of the time series data, a channel attention mechanism is introduced in the stacking and convolution module to suppress the responses of irrelevant noise sequences, which is beneficial to the feature extraction of the convolution module. Third, in order to learn complex time patterns in the context of long-term time series data, a time series decomposition block is introduced, and the idea of decomposition is adopted to extract the long-term stable trend sequence, reducing the impact on the reconstruction effect. Compared with the original deep models (such as GRU, GNN, DLinear, Transformer series models, TCN, etc.), this model is easier to capture the long-term changes and multi-dimensional correlations of time series data, and can take into account the relative importance of each index sequence. Therefore, it has better feature extraction ability and time series data reconstruction effect, and can achieve higher precision and lower false negative rate in the anomaly detection task.
[0034] Glossary:
[0035] Time series data: Time series data is a string of data indexed by the time dimension, where each data point has a timestamp. This type of data describes the measured values of a certain measured subject at each time point within a time range. This type of data is very common in various fields, such as finance (stock prices), meteorology (daily temperatures), industry (sensor data), etc. Its main feature is that its index is time, and there is continuity and correlation in time among data points.
[0036] Time domain: The time domain is a representation that describes how a signal changes over time. In the time domain, a signal is represented as a function of amplitude or intensity changing over time. A time-domain signal can be continuous or discrete. For a continuous-time domain signal, we use a continuous time variable to represent how the signal changes over time. For a discrete-time domain signal, we use a series of discrete time points to represent the value of the signal at specific moments. Time-domain analysis can reveal some important characteristics of a signal, such as waveform, period, amplitude, etc. However, analyzing a signal only in the time domain may not fully reveal its frequency characteristics, and in this case, frequency-domain analysis methods such as the Fourier transform are needed.
[0037] Fast Fourier Transform (FFT): The Fast Fourier Transform (FFT) is an efficient algorithm for calculating the Discrete Fourier Transform (DFT). It can transform a discrete-time signal from the time domain to the frequency domain. The FFT greatly reduces the complexity of calculating the DFT and significantly improves the computational efficiency. Through the FFT, we can quickly understand the intensity of different frequency components contained in the signal, which helps to achieve operations such as filtering, compression, and feature extraction. In practical applications, the FFT is widely used in various scenarios. For example, in audio processing, it can be used to analyze the frequency components of sound, helping to achieve audio compression, noise reduction, and sound effect optimization. In the field of communication, the FFT is used for signal modulation and demodulation, spectrum analysis, and channel estimation, etc.
[0038] ShuffleNet Convolution: ShuffleNet is a convolutional neural network (CNN) architecture with extremely high computational efficiency. This architecture utilizes two new operations: pointwise grouped convolution and channel shuffle. Through grouped convolution, each convolution operates only on its corresponding input channel group, which effectively reduces the network's capacity and thus greatly reduces the computational cost of traditional convolution. And through the channel shuffle operation to scramble the groups, the information flow between channel groups can be achieved, enabling the acquisition of information from other groups. These two operations enable ShuffleNet convolution to maintain high accuracy while reducing the computational cost.
[0039] Multilayer Perceptron (MLP): MLP, that is, Multilayer Perceptron, is a common feedforward artificial neural network model. It consists of an input layer, multiple hidden layers, and an output layer. Each layer contains multiple neurons, and the neurons are connected by weights. The input signal is processed through weighted summation and activation functions in each layer, gradually extracting and transforming features, and finally generating a prediction or classification result at the output layer. In terms of applications, the MLP can be used to predict stock prices, sales volumes, etc.; in pattern classification tasks, it can classify images, texts, etc.
[0040] Residual connection: A residual connection refers to directly adding the output of a certain previous layer to the input of a certain subsequent layer in a neural network. The purpose of doing this is to solve the problems of vanishing gradients and network performance degradation that may occur when the network depth increases. Through residual connections, information can be transmitted more smoothly in the network, making the network easier to train and enabling the construction of deeper and more powerful models.
[0041] Channel attention mechanism: Channel attention is a mechanism used in deep learning to evaluate the importance and assign weights to the channel dimension of feature maps. It assigns different weights to each channel by analyzing the importance of the information contained in each channel. This enables the model to pay more attention to the channel features that are more valuable for the task, suppress less important channel features, and thus improve the performance and expressive ability of the model. It is usually applied in the field of computer vision. For example, in image recognition tasks, channel attention helps the model to more accurately extract key image features and improve the recognition accuracy.
[0042] F1 score: The F1 score is a comprehensive metric used to evaluate the performance of a classification model. It comprehensively considers the precision and recall of the model. Precision measures the proportion of samples predicted as positive by the model that are actually positive, and recall represents the proportion of samples that are actually positive and are correctly predicted as positive by the model. The F1 score is the harmonic mean of precision and recall, and its calculation formula is: F1 = 2 * (precision * recall) / (precision + recall). In practical applications, such as in information retrieval, anomaly detection, etc., when we hope that the model can accurately identify positive examples and also minimize the omission of positive examples, the F1 score can provide a more comprehensive and balanced evaluation to help us better judge the quality of the model.
[0043] To illustrate the technical solution of the present invention in more detail, the present invention will be described in detail below in conjunction with embodiments:
[0044] An anomaly detection method for communication devices based on channel attention and frequency domain analysis includes the following steps:
[0045] Step S1: Extract original data from the database, and after normalization processing, obtain subsequences;
[0046] Step S2: Decompose the subsequences in Step S1 in time series to obtain a trend sequence component and a periodic sequence component;
[0047] Step S3: Perform FFT transformation and analysis on the periodic sequence component to obtain an average amplitude set and establish an input feature map;
[0048] Step S4: Compress and stimulate the input feature map to generate a channel attention map for adjusting the weights of each index in the input feature map;
[0049] Step S5: After weighting the feature map, perform a convolution operation to obtain which is used to simultaneously capture the intra-period and inter-period changes of multi-index time series data, as well as the interdependencies between different indexes;
[0050] Step S6: Output and reorganize the convolution result of Step S5 into a one-dimensional tensor output to obtain the output sequence of the stacked convolution module;
[0051] Step S7: For the convolution module output sequence in Step S6, learn the deep abstract time series data features of the stacked convolution results under different periods through a multi-layer perceptron, and output the result;
[0052] Step S8: Weight and aggregate the initial trend sequence component, output the residual sequence, and use it as the input of the next layer;
[0053] Step S9: Adopt a loss function to update the parameters of each layer layer by layer;
[0054] Step S10: Perform anomaly detection, calculate the reconstruction error based on the reconstruction result of the training set, and find the anomaly threshold according to the pre-input hyperparameter anomaly rate; calculate the anomaly score, and compare it with the anomaly threshold for anomaly judgment; if the anomaly score is greater than the anomaly threshold, it is determined that there is an anomaly at this time point.
[0055] Specifically:
[0056] Step S1.1: Extract the multi-monitoring index information of a certain network communication device from the time series data of the full-professional equipment operation in the network domain continuously accessed from the Jiangsu Telecom data sharing platform as the original data (in this example, the packet forwarding rate, packet loss rate, broadband utilization rate, etc. of the first eight ports are extracted, a total of 8 * 12 index sequences), and convert them into a multi-index time series data sequence to obtain a multi-dimensional time series dataset S of C×L, where C is the number of data indexes and L is the total time step length. Perform min-max normalization processing to obtain the normalized dataset S′.
[0057] Step S1.2: Slide and segment the data with a sliding window of length T at a step size of T / 2, and divide the above normalized time series S′ into X 1 , X 2 , …, X n , a total of n subsequences of length T.
[0058] Step S2: Input the subsequence into the time series decomposition block of the reconstruction model. The time series decomposition block divides the subsequence X ∈ {X 1 , X 2 , …, X n} Decompose into a periodic sequence component X s and a trend sequence component X t , respectively reflecting the trend and periodicity of the time series, where is multi-dimensional time series data with a time step of T and an index number of C. Equations (1) and (2) are the specific decomposition processes. Padding(·) means padding zeros at the beginning of the original sequence to an integer multiple of the pooling window size, and AvgPool(·) means average pooling in the time dimension.
[0059] X t = AvgPool(Padding(X)) (1)
[0060] X s = X - X t (2)
[0061] When the pooling window size is h and the step size is s, the average pooling AvgPool(·) is as follows. X[i] represents the i-th data point of the vector X:
[0062]
[0063] Through the above transformation, the original time series data is decomposed into a trend sequence component and a periodic sequence component Use X t , X s = Decomp(X) to summarize the above process.
[0064] Step S3.1: Input the periodic sequence X s into the stacked convolution module. For the periodic sequence X s of multi-index data, first analyze its frequency domain characteristics through the fast Fourier transform, find the first K largest amplitudes and their corresponding periods, and split and stack the original input data based on different period lengths so that the subsequent convolution processing can capture the inter-period and intra-period correlations of multi-index data simultaneously. After the sequence X s is transformed and analyzed by FFT, the following results can be obtained:
[0065] A = Avg(Amplitude(FFT(X s ))) (4)
[0066] A i = TopK(A), i ∈ {1, …, k} (5)
[0067]
[0068] Among them, FFT(·) represents performing fast Fourier transform on the time-series data of C different indicators respectively, Amplitude(·) represents taking the amplitude part of the transformation result, Avg(·) represents calculating the average amplitude in the C-index dimension, and finally obtaining the set of average amplitudes of the multi-dimensional time-series data varying with the frequency f
[0069] To avoid the influence of high-frequency noise, we only select the first K amplitude values and their corresponding frequencies at. Here, the frequency corresponding to the j-th amplitude is f j , representing the number of cycles at time T, and the corresponding cycle length is TopK(·) represents taking the top K largest numbers in the set, A i is the top K largest amplitudes of. argTopK(·) represents taking the indices corresponding to the top K maximum values in the set. By argTopK(A), the frequency f corresponding to the amplitude A can be obtained. Ceil(·) is the ceiling function, and p i represents the corresponding cycle lengths of the K most significant frequencies.
[0070] Step S3.2: Through the processing of Step S3.1, K different cycle lengths p i , i ∈ {1, …, k} are obtained. Based on this, the original time-series data is sliced and stacked to obtain K multi-dimensional tensors with different shapes. The specific process is as Figure 3 shown. Formula (7) outlines the specific process of this slicing and stacking. To ensure that the sliced sequence is an integer multiple of p i , before slicing and stacking, 0 is padded to the end of X s to an integer multiple of p i . This operation is represented by Padding(·). SplitRecomb(p, x) represents slicing x in the time dimension with p as the period and stacking it vertically in sequence. represents the multi-dimensional tensor obtained after slicing and stacking based on the i-th cycle p i .
[0071]
[0072] Taking as the input feature map of the convolutional neural network to capture the trend changes in adjacent time periods, the long-term dependencies between different cycles, and the dependencies between multiple indicators.
[0073] Step S4: Compress and excite the input feature map to generate the channel attention map S i, focus on the important features of multi-index data and suppress the response of noise sequences. During the training process, the model will learn the weight allocation method according to the contribution degree of each index to the reconstruction effect, and generate the channel attention map S i , used to adjust the input feature map The weights of each index in are adjusted. This adaptive weight adjustment mechanism enables the model to focus more on learning the channel information that contributes greatly to the reconstruction task. Among the monitoring indicators of communication devices, important indicators such as throughput, packet loss rate, and broadband utilization rate will be assigned higher weights, while indicators with relatively lower importance such as the status of non-critical links will be assigned lower weights to reduce their impact on the reconstruction of other sequences during the convolution stage. The calculation and allocation of the weights of each index sequence are as Figure 4 shown
[0074] For the input feature map Average pooling and max pooling are respectively performed on each index sequence to aggregate the time and period dimension information, obtaining two C-dimensional pooled feature maps to complete the compression of the feature map. Then, the two feature maps are fed into a shared multi-layer perceptron with one hidden layer to obtain two 1x1xC channel attention maps, completing the excitation of the feature map and generating weight information. The two channel attention maps obtained through the multi-layer perceptron are added and activated to finally generate the channel attention map S of the input feature map i . Take S i as the weight information of each channel, that is, the weight information of each index. To reduce the number of parameters, the number of neurons in the hidden layer is C / r, where r is the dimensionality reduction coefficient. The input feature map The generation of the corresponding channel attention map S i is shown in formula (8).
[0075]
[0076] Among them, AvgPool(·) and MaxPool(·) represent average pooling and max pooling respectively, and MLP(·) represents the multi-layer perceptron. Through the compression and excitation of the convolutional attention module, the abstract features of each channel of the input feature map are extracted. σ(·) represents the ReLU activation function, and maps the addition result to the output end through a non-linear transformation. Generate the channel attention map S i as the weight matrix, and assign different weights to each channel of the input .
[0077]
[0078] Among them, F scale (u, s) represents the element-wise multiplication of u and s
[0079] Step S5: For the feature map After weighting, we import it into the efficient computer vision network, ShuffleNet, to perform convolutional operations, aiming to capture the intra-period and inter-period variations of multi-metric time series data, as well as the interdependencies between different metrics. ShuffleNet employs a grouped convolution strategy, ensuring that each convolutional operation focuses only on its corresponding input channel group, enabling sparse channel connections and significantly reducing the computational cost. Additionally, through channel shuffling, ShuffleNet allows the convolutional layer to receive feature information from different groups, maintaining high accuracy while reducing the computational cost. Thanks to the efficient architecture of ShuffleNet and its optimized utilization of computational resources, our model is applicable to larger time steps and larger hyperparameter K, thus facilitating the capture of long-term features of time series.
[0080]
[0081] Step S6.1: Recombine the two-dimensional convolution result output of into one dimension and restore the original length L to obtain K one-dimensional tensor outputs
[0082]
[0083] where SplitRecomb -1 (·) represents the inverse transformation of splitting and recombination, which splits and recombines the original two-dimensional output into a one-dimensional output, and SubSeq(·) represents restoring the time series data with a length that is an integer multiple of p i to its original time step L.
[0084] Step S6.2: The amplitude A obtained through Step S3.1 can reflect the relative importance of the selected frequency and period, thus corresponding to the importance of each transformed two-dimensional tensor. Normalize the amplitude calculated based on frequency domain analysis to obtain which is used as the weight of the output sequence to obtain the output S i of the stacked convolution module. Where Softmax(A 1 ,…,A k ) represents the normalization operation.
[0085]
[0086] Step S7: For the output sequence S i of the convolution module considering K different periodic features, i ∈ {1,…,k}, learn the deep abstract time series data features of the stacked convolution results under different periods through a multi-layer perceptron, and output the result For \(i\in\{1,\ldots,k\}\), it has the same shape as the original input. The multi-layer perceptron separately considers the convolution results under \(K\) different periods and deeply reconstructs the results according to the abstract data features of the learned hidden layer.
[0087]
[0088] Step S8: Finally, the initial trend sequence component \(X\) s is weighted and aggregated to output the residual sequence \(O\), which is used as the input for the next layer. Residual connections are used between multiple layers, where \(i\in\{1,2\}\) is a learnable parameter.
[0089]
[0090] Step S9: In the training phase, the loss function used is the mean square error (MSE), and its calculation formula is:
[0091]
[0092] where \(Y\) ij is the predicted value, is the reference value. According to the chain rule, the gradient of the loss function \(L\) with respect to each parameter of each layer is calculated layer by layer forward, and the parameters are updated along the opposite direction of the gradient. For example, to update the parameter \(W\) of a certain linear layer in the MLP, first calculate the gradient of the loss function \(L\) with respect to \(Y\), that is, the partial derivative, and then through backpropagation, calculate the gradient of \(Y\) with respect to the output \(S\) of the MLP i of calculate the gradient of \(S\) i with respect to \(W\) Finally, according to the preset learning rate \(\alpha\), update \(W\) along the opposite direction of the gradient to In this way of backpropagation, the parameters of each layer are updated layer by layer.
[0093] Step S10.1: In the anomaly detection phase, first reconstruct the results based on the training set and calculate the reconstruction error, and find the anomaly threshold according to the pre-input hyperparameter anomaly rate.
[0094]
[0095] where \(E\) is the set composed of the reconstruction errors of each time point of all samples, \(n\) is the number of samples, \(L\) is the time step of the samples, and \(Y\) is the reconstruction result. \(E'\) is the result of sorting \(E\), \(a\) is the pre-input anomaly rate, \(|E'|\) represents the number of elements in the set \(E'\), and \(THR\) represents the calculated anomaly threshold.
[0096] Step S10.2: Perform anomaly judgment by calculating the anomaly score and comparing it with the threshold. Since this time-series data reconstruction model captures features such as periodic changes and trend changes in the time-series dataset through training, and can infer the reconstruction result of the normal mode from the input data, the actual time-series data to be detected is reconstructed using the model. According to the comparison between the reconstruction result and the actual data, calculate the anomaly score at each time point, compare the anomaly score with the anomaly threshold, and determine whether there is an anomaly at a specific time point.
[0097]
[0098] Where E t is the anomaly score of the data to be detected at time t. By comparing this anomaly score with the threshold THR, if the anomaly score is greater than the threshold THR, it is determined that there is an anomaly at this time point.
[0099] The present invention selects multiple monitoring index data of a network device within 2 months from February 2024 to April 2024 obtained from the time-series data of the continuous access network domain full-professional equipment operation of the Jiangsu Telecom data sharing platform, collected once every 20 minutes, and 24 hours of data are collected every day. Select 13-dimensional indicators from April 1, 2024 to April 14, 2024, a total of 13,104 indicator data as the verification data of the model in this article. In order to compare the performance of this model and other baseline methods, the same data preprocessing methods and evaluation indicators used in previous time-series data anomaly detection work are used, including:
[0100]
[0101] Where F1 is the F1 score, P is the precision, R is the recall rate, TP is the number of true positives, FP is the number of false positives, and FN is the number of false negatives. Do not pay attention to the anomaly threshold selection strategy. If a model does not give a specific threshold selection strategy, obtain the F1 score of each model by brute-force searching for the threshold at an anomaly rate of 2%. The purpose of doing this is to unify the threshold search strategy. In addition, in order to maintain the consistency of the experimental settings, the output adjustment strategy used by the previous model, point-adjust, is adopted: for a segment anomaly, if the model triggers any subset of anomalies, this is acceptable. Therefore, if any observation value in the true anomaly segment is detected as an anomaly by the model, it can be considered that the anomaly detection of this segment is correct. For general point anomalies, no adjustment is made.
[0102] For the comparison models, a single-layer linear model DLinear, a CNN-based TCN, as well as Transformer, TimesNet, and Autoformer are selected as baselines. The parameters of each model are shown in Table 1. In the table, d_model is the dimension of the neural network in the model, and d_ff is the dimension of the feed-forward neural network in the model. seq_len represents the length of the input data in the model, num_kernels represents the number of convolutional kernels, and top_k represents the size of the hyperparameter K in the frequency-domain analysis. moving_avg is the window size of the moving average algorithm, and moving_step is the moving step of the moving average algorithm. anomaly_ratio represents the pre-input anomaly rate. Most of the existing parameters of the original model remain the same configuration as in the previous TimesNet architecture model paper. For the Transformer series models, label_len represents the length of the starting token, enc_in represents the encoder input size, and e_layers represents the number of encoder layers. Our model uses the ADAM optimizer and adopts an Early Stop mechanism, which allows us to stop the model training based on the validation set to prevent overfitting and enhance the generalization performance. When patience is 3, that is, if the performance does not improve in three consecutive epochs, the training process will stop. All experimental results are represented as the average of 8 consecutive repeated experiments.
[0103] Table 1 Baseline Models and Corresponding Parameters
[0104]
[0105] In the experiment, this model is compared with five advanced baseline methods, namely DLinea, TCN, Transformer, TimesNet, and Autoformer. Table 2 shows the detailed performance comparison results using the multivariate time series data extracted from the real telecom network device traffic data as the experimental dataset.
[0106] Table 2 Performance of This Model and Baseline Models on the Telecom Dataset
[0107]
[0108]
[0109] In the F1 score, the main reference evaluation metric for the anomaly detection task, we achieved the best performance, exceeding the best baseline score by 2.23%. Among these models, TCN obtained the worst results. TCN constructs long-term dependencies through one-dimensional convolution and dilated convolution. However, one-dimensional convolution has the problem of limited receptive field, which restricts the ability of long-term prediction. For complex multi-dimensional variables, in addition, TCN is difficult to capture the complex interactions and dependencies among multiple metrics. When modeling complex data distributions, the two-dimensional convolution we used significantly achieved better results.
[0110] Transformer adopts the attention mechanism to adaptively capture long-term time dependencies and has good global modeling ability. However, in long sequence analysis, when the time step is increased to 192 or higher, its performance drops significantly. This is because Transformer has a high computational cost when facing long time series and it is difficult to directly find reliable dependencies from scattered time points. In contrast, DLinear achieved better results. DLinear models the trend and residual sequences using two single-layer linear networks, which can simply but effectively extract the time relationships in time series. This direct multi-step prediction strategy avoids the error accumulation effect in autoregressive prediction. However, the DLinear model is based on a linear assumption. For time series data in the real world that contains complex non-linear relationships, DLinear cannot effectively capture its inherent non-linear characteristics, thus affecting the prediction accuracy. At the same time, DLinear cannot consider the interactions among variables, resulting in inaccurate reconstruction results. Autoformer adopts a phased prediction strategy based on Transformer. First, it generates a rough global trend prediction, and then gradually refines it to each time point. This strategy can more accurately capture short-term and long-term dependencies, thereby improving the prediction accuracy and reducing the error accumulation effect to a certain extent. Therefore, a higher F1 score can be obtained in the anomaly detection task.
[0111] Combined with the above discussion, compared with the prior art with the publication number CN117556311A, in the application field, the present invention is mainly applied to communication devices; in terms of technical means, the present invention focuses on frequency domain analysis and channel attention mechanism; in the stacked convolution part, the present invention analyzes based on K different periods. Additionally, in terms of the model architecture: the present invention uses the ShuffleNet convolutional network and the channel attention mechanism for serial processing. Compared with the prior art with the publication number CN117688496A, the present invention focuses on frequency domain analysis and channel attention mechanism to improve the accuracy of anomaly detection. In terms of the technical means adopted, the present invention uses a CNN-based method to extract the data changes within and between periods. Compared with the prior art with the publication number CN117851920A, the present invention focuses on frequency domain analysis and channel attention mechanism to improve the accuracy of anomaly detection; in terms of the technical architecture: the present invention uses the ShuffleNet convolutional network and the channel attention mechanism; in terms of the anomaly detection strategy: the present invention determines anomalies through reconstruction error. Therefore, the technical solution of the present invention is significantly different from the prior art in terms of technical architecture and technical solution.
[0112] The present invention adopts the idea of decomposition to learn complex time patterns in the long-term scenario, and decomposes the time series into a trend part and a period part. By analyzing the deep abstract features of the period part, the ability to discover period-based dependencies in frequency domain analysis is improved. The results of DLinear and Autoformer show that decomposing the original sequence into a trend and a period part for separate processing can effectively improve the prediction effect of long-term time series. To capture the complex interactions and dependencies between multiple metrics and complete the modeling of multi-dimensional complex data distributions, we use sliced stacked fast processing followed by two-dimensional convolution to capture its periodic characteristics, and introduce a channel attention mechanism to focus on the important features of time series data, suppress the response of irrelevant noise information, and be able to capture global key information more stably and accurately, thereby effectively reducing the false alarm rate and further improving the F1 score.
[0113] The present invention proposes a method for anomaly detection of communication devices based on channel attention and frequency domain analysis. The specific steps include: in the training stage, using a time series reconstruction model to capture the normal mode of the data; in the anomaly detection stage, reconstruct the data to be detected, compare the reconstructed data with the input data, and find the time points where the data to be detected is abnormal according to the reconstruction error.
[0114] Compared with the prior art, the following key points are included:
[0115] 1. A communication device anomaly detection model based on channel attention and frequency domain analysis is proposed, where the reconstruction module includes a time series decomposition block and a stacked convolution module, and residual connections are used between layers.
[0116] 2. A design is proposed to decompose the periodic component from the original sequence before processing by the stacked convolutional module and use it as the input of the stacked convolutional module, so as to explore the potential laws of time series data from the frequency perspective.
[0117] 3. A stacked convolutional module is designed by embedding a channel attention mechanism in the stacking and convolutional processing, which can focus on the important sequences of multi-index data to improve the reconstruction accuracy.
Claims
1. A communication equipment anomaly detection method based on channel attention and frequency domain analysis, characterized in that The steps include: Step S1: extract the original data from the database, perform normalization processing, and obtain a subsequence; Step S2: Decompose the subsequence of step S1 into time series to obtain trend sequence components and periodic sequence components; Step S3: Perform FFT transformation and analysis on the periodic sequence components to obtain an average amplitude set and establish an input feature map; Step S4: compress and excite the input feature map to generate a channel attention map to adjust the weights of each indicator in the input feature map; Step S5: After weighting the feature map, perform a convolution operation to obtain It is used to simultaneously capture the intra-cycle and inter-cycle changes of multi-indicator time series data, as well as the interdependencies between different indicators; Step S6: reorganize the convolution result output of step S5 into a one-dimensional tensor output to obtain an output sequence of the stacked convolution module; Step S7: for the convolution module output sequence of step S6, the deep abstract time series data features of the stacked convolution results at different periods are learned by a multi-layer perceptron, and the results are output; Step S8: Aggregate the initial trend sequence components with weights, output the residual sequence, and use it as the input of the next layer; Step S9: using the loss function to update the parameters of each layer layer by layer; Step S10: perform anomaly detection, reconstruct the results based on the training set and calculate the reconstruction error, and find the anomaly threshold according to the pre-input hyperparameter anomaly rate; Calculate the anomaly score and compare it with the anomaly threshold to make an anomaly judgment; If the anomaly score is greater than the anomaly threshold, it is determined that an anomaly exists at this time point.
2. The communication device anomaly detection method based on channel attention and frequency domain analysis according to claim 1 is characterized in that The specific process of the above step S1 is: Step S1.1: extract multiple monitoring indicator information of a certain network communication device from the database as raw data, and convert it into a multi-indicator time series data sequence to obtain a C×L multidimensional time series data set S, where C is the number of data indicators and L is the total time step. Use min-max normalization processing to obtain a normalized data set S′; Step S1.2: Sliding the data with a sliding window of time step T and step size T / 2, and divide the normalized S′ into X1, X2, …, X n , a total of n subsequences with a time step length of T.
3. The communication device anomaly detection method based on channel attention and frequency domain analysis according to claim 2 is characterized in that The specific process of the above step S2 is: The subsequence is input into the time series decomposition block of the reconstruction model; the time series decomposition block converts the subsequence X∈{X1,X2,…,X n } is decomposed into a periodic series component X that reflects the periodicity of the time series s and the trend series component X that reflects the trend of the time series t ,in For multidimensional time series data with a time step of T and a number of indicators of C, the specific decomposition process is as follows: X t =AvgPool(Padding(X)) (1) X s =X-X t (2) Where Padding(·) means padding the original sequence with zeros to an integer multiple of the pooling window size, and AvgPool(·) means average pooling in the time dimension. When the pooling window size is h and the step size is s, the average pooling AvgPool(·) is as follows, where X[i] represents the i-th data point of vector X: Through the above transformation, the original time series data is decomposed into trend sequence components and periodic series components Use X t ,X s =Decomp(X) summarizes the above process.
4. The communication device anomaly detection method based on channel attention and frequency domain analysis according to claim 3 is characterized in that The specific process of the above step S3 is: Step S3.1: Convert the periodic sequence component X s Input stacked convolution module; for the periodic series component X of multi-index data s First, the frequency domain characteristics are analyzed by fast Fourier transform to find the first K maximum amplitudes and corresponding periods, which are used to segment and stack the original input data based on different period lengths, so that the subsequent convolution processing can simultaneously capture the inter-period and intra-period correlations of multiple index data; the sequence X s After FFT transformation and analysis, the following results are obtained: A=Avg(Amplitude(FFT(X s ))) (4) AND i =TopK(A),i∈{1,…,k} (5) Among them, FFT(·) means to perform fast Fourier transform on the time series data of C different data indicators respectively, Amplitude(·) means to take the amplitude part of the transformation result, and Avg(·) means to calculate the average amplitude in the dimensions of C data indicators, and finally obtain the average amplitude set of multi-dimensional time series data changing with frequency f To avoid the influence of high frequency noise, select The first K amplitude values and their corresponding frequencies at the time; here the frequency corresponding to the jth amplitude is f j , represents the number of cycles under the time step T, and the corresponding cycle length is TopK(·) means taking the largest top K numbers in the set, A i for The first K maximum amplitudes of ; argTopK(·) means taking the index corresponding to the first K maximum values in the set, and obtaining the frequency f corresponding to the amplitude A through argTopK(A), Ceil(·) is the upper integer function, p i represents the corresponding period lengths of the K most significant frequencies; Step S3.2: Through the processing of step S3.1, K different period lengths p are obtained i ,i∈{1,…,k}, based on which the original time series data is cut and stacked to obtain K multi-dimensional tensors with different shapes; Formula (7) summarizes the specific process of the split stacking. In order for the split sequence to be p i Integer multiple length, before splitting and stacking, s Fill the tail with 0 to p i The operation is represented by Padding(·). SplitRecomb(p,x) means to split x into periods of p in the time dimension and stack them up and down in sequence. Represents the period p based on the i-th period i The multidimensional tensor obtained after splitting and stacking; by As the input feature map of the neural network, it is used to capture the trend changes in adjacent time periods, the long-term dependencies between different cycles, and the dependencies between multiple indicators.
5. The communication device anomaly detection method based on channel attention and frequency domain analysis according to claim 4 is characterized in that The specific process of the above step S4 is: S4: Input feature map Compress excitation and generate channel attention map S i , focus on the important features of multi-index data and suppress the response of noise sequences; During the training process, the reconstruction model learns the weight distribution method according to the contribution of each indicator to the reconstruction effect and generates a channel attention map S i , used to adjust the input feature map The weight of each indicator in the convolution phase is determined; among the monitoring indicators of communication equipment, important indicators will be given higher weights, while indicators with relatively lower importance will be given lower weights, reducing their impact on the reconstruction of other sequences in the convolution phase. The calculation and allocation of the weights of each indicator sequence are as follows: For the input feature map Perform average pooling and maximum pooling on each indicator sequence, aggregate the time and period dimension information, obtain two C-dimensional pooling feature maps, and complete the compression of the feature map; then send the two feature maps to a shared multi-layer perceptron containing a hidden layer to obtain two 1x1xC channel attention maps, complete the excitation of the feature map, and generate weight information; The two channel attention maps obtained by the multi-layer perceptron are added and activated to finally generate the input feature map The channel attention map S i ; S i As the weight of each channel, that is, the weight of each indicator; in order to reduce the number of parameters, the number of neurons in the hidden layer is C / r, where r is the dimensionality reduction coefficient; the input feature map Corresponding channel attention map S i The generation of is shown in formula (8): Among them, AvgPool(·) and MaxPool(·) represent average pooling and maximum pooling respectively, and MLP(·) represents multi-layer perceptron. Through the compression and excitation of the convolutional attention module, the abstract features of each channel of the input feature map are extracted. σ(·) represents the ReLU activation function, which maps the addition result to the output through nonlinear transformation. The channel attention map S is generated. i As the weight matrix, the input Different weights are assigned to each channel; where F scale (u,s) represents the element-wise multiplication of u and s.
6. The communication device anomaly detection method based on channel attention and frequency domain analysis according to claim 5 is characterized in that The specific process of the above step S6 is: S6.1: The convolution result output is reorganized into one dimension, and the original time step L is restored to obtain K one-dimensional tensor outputs Where SplitRecomb -1 (·) represents the inverse transformation of splitting and reorganizing, splitting and reorganizing the original two-dimensional output into one-dimensional output, and SubSeq(·) represents the transformation of p i The time series data with integer multiple lengths is restored to its original time step L; S6.2: The amplitude A obtained in step S3.1 is normalized to obtain This is the output sequence The weights of the stacked convolution module are obtained. i ; Among them, Softmax(A1,…,A k ) represents a normalization operation.
7. The communication device anomaly detection method based on channel attention and frequency domain analysis according to claim 6 is characterized in that The specific process of the above step S7 is: For the convolution module output sequence S considering K different periodic features i , i∈{1,…,k}, the deep abstract time series data features of the stacked convolution results at different periods are learned by multi-layer perceptron, and the output is the result Same shape as the original input; The multi-layer perceptron considers the convolution results under K different cycles respectively, and deeply reconstructs the results according to the abstract data features of the learned hidden layer.
8. The communication device anomaly detection method based on channel attention and frequency domain analysis according to claim 7 is characterized in that The specific process of the above step S8 is: Finally, the initial trend sequence component X s With weighted aggregation, the residual sequence O is output and used as the input of the next layer; residual connections are used between multiple layers; in as a learnable parameter.
9. The communication device anomaly detection method based on channel attention and frequency domain analysis according to claim 8 is characterized in that The specific process of the above step S9 is: In the training phase, the loss function used is the mean square error, and the calculation formula is: where Y ij is the predicted value, is the reference value. According to the chain rule, the gradient of the loss function L for each parameter of each layer is calculated forward layer by layer, and the parameters are updated in the opposite direction of the gradient. To update the parameters W of a linear layer in the MLP, the gradient of the loss function L to Y, that is, the partial derivative, is first calculated, and then Y is calculated for the MLP output S through back propagation. i Gradient Calculate S i The gradient of W Finally, according to the preset learning rate α, W is updated in the opposite direction of the gradient: Through this back-propagation method, the parameters of each layer are updated layer by layer.
10. The communication device anomaly detection method based on channel attention and frequency domain analysis according to claim 9 is characterized in that The specific process of the above step S10 is as follows: S10.1: In the anomaly detection stage, firstly, the reconstruction result and reconstruction error are calculated based on the training set, and the anomaly threshold is found according to the pre-input hyperparameter anomaly rate; Where E is the set of reconstruction errors of all samples at each time point, n is the number of samples, L is the time step of the sample, and Y is the reconstruction result; E′ is the result after sorting E, a is the pre-input anomaly rate, |E′| represents the number of elements in the E′ set, and THR represents the calculated anomaly threshold; S10.2: Calculate the anomaly score and compare it with the threshold to make an anomaly judgment; since the time series data reconstruction model captures the characteristics of periodic changes and trend changes of the time series data set through training, it can infer the reconstruction result of the normal mode through the input data, reconstruct the actual time series data to be detected using the model, and calculate the anomaly score of each time point based on the comparison of the reconstruction result with the actual data, and compare the anomaly score with the anomaly threshold to determine whether there is an anomaly at a specific time point; Where E t is the anomaly score of the data to be detected at time t. By comparing the anomaly score with the threshold THR, if the anomaly score is greater than the threshold THR, it is determined that an anomaly exists at this time point.
Citation Information
Patent Citations
Unsupervised time sequence anomaly detection method based on multi-dimensional feature fusion
CN117556311A
Abnormity diagnosis method, system and equipment for satellite telemetering multi-dimensional time series data
CN117688496A
Electric power internet of things data anomaly detection method and system
CN117851920A
Multi-dimensional time series data anomaly detection method and device based on adversarial training and frequency domain improved self-attention mechanism, and medium
CN116502164A
Time series data anomaly detection method based on multi-head attention model
CN117076936A
Cited By
Communication dynamic analysis method for virtual network security
CN120415917A
Abnormal feature analysis method for complex metering sensor based on long-term accumulated data
CN120832626A
Passive optical fiber multi-parameter digital twin drive abnormal root cause positioning method and system
CN121188579A
Engine part production tracing method and system based on barcode data
CN122114759A