Generation method of network sequential sequence anomaly detection model and related equipment
Through Fourier transform and contrast learning algorithm, the problem of high annotation cost of abnormal data in the network timing sequence anomaly detection model is solved, and the ability to efficiently identify abnormal patterns in multiple perspectives is realized.
Patent Information
- Application Number
- CN202510189078.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-07-11
AI Technical Summary
The existing network timing sequence anomaly detection model requires a large number of negative sample annotations during the construction process, which makes it costly and difficult to effectively identify multi-dimensional anomaly data.
By obtaining network timing sequence data, multi-dimensional data filling and conversion processing are performed, the data is converted from the time domain to the frequency domain using Fourier transform, the time domain and frequency domain characteristics are enhanced, and the abnormal detection model is trained through a comparative learning algorithm to avoid labeling of abnormal data.
Without the need for anomaly data annotation, timing features can be extracted from multiple perspectives, reducing costs and improving the ability to capture abnormal patterns, ensuring accurate identification in data without labels or complex features.
Smart Images

Figure CN120296613A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technologies, and in particular, to a method for generating a network time series anomaly detection model and related devices. Background Art
[0002] Network time series anomaly detection aims to identify abnormal behaviors in network time series data that are significantly different from normal patterns. When constructing a network time series anomaly detection model, current contrastive learning algorithms usually rely on positive and negative sample pairs to learn feature representations by maximizing the similarity between positive samples and minimizing the similarity between negative samples. However, current contrastive learning algorithms require a large number of negative samples and need to perform comprehensive anomaly judgments on multiple dimensions of network time series, resulting in a high cost of abnormal data annotation.
[0003] In view of this, how to avoid the high cost caused by annotating abnormal data when constructing an anomaly sequence detection model has become an urgent technical problem to be solved. Summary of the Invention
[0004] In view of this, an object of the present disclosure is to propose a method for generating a network time series anomaly detection model and related devices to solve or partially solve the above technical problems.
[0005] Based on the above object, a first aspect of the present disclosure proposes a method for generating a network time series anomaly detection model, the method comprising:
[0006] Obtain a network time series, extract multi-dimensional data from the network time series, and perform padding processing and conversion processing on the multi-dimensional data to obtain standard data;
[0007] Convert the standard data from the time domain to the frequency domain through Fourier transform to obtain time domain features and frequency domain features;
[0008] Perform enhancement processing on the time domain features to obtain time domain samples, and perform enhancement processing on the frequency domain features to obtain frequency domain samples;
[0009] Train an initial anomaly detection model according to the time domain samples and the frequency domain samples through a contrastive learning algorithm to obtain a network time series anomaly detection model.
[0010] Based on the same inventive concept, a second aspect of the present disclosure proposes a device for generating a network time series anomaly detection model, comprising:
[0011] An extraction module, configured to obtain a network time series, extract multi-dimensional data from the network time series, and perform padding processing and conversion processing on the multi-dimensional data to obtain standard data;
[0012] A conversion module, configured to convert the standard data from the time domain to the frequency domain through Fourier transform to obtain time-domain features and frequency-domain features;
[0013] An enhancement module, configured to perform enhancement processing on the time-domain features to obtain time-domain samples, and perform enhancement processing on the frequency-domain features to obtain frequency-domain samples;
[0014] A training module, configured to train an initial anomaly detection model according to the time-domain samples and the frequency-domain samples through a contrastive learning algorithm to obtain a network time-series anomaly detection model.
[0015] Based on the same inventive concept, a third aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor. When the processor executes the computer program, the above-mentioned method is implemented.
[0016] Based on the same inventive concept, a fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the above-mentioned method.
[0017] As can be seen from the above, the present disclosure provides a method for generating a network time-series anomaly detection model and related devices. Obtain a network time series, extract multi-dimensional data from the network time series, and perform padding processing and conversion processing on the multi-dimensional data to obtain standard data. Convert the standard data from the time domain to the frequency domain through Fourier transform to obtain time-domain features and frequency-domain features. Perform enhancement processing on the time-domain features to obtain time-domain samples, and perform enhancement processing on the frequency-domain features to obtain frequency-domain samples. Train an initial anomaly detection model according to the time-domain samples and the frequency-domain samples through a contrastive learning algorithm to obtain a network time-series anomaly detection model. In this way, by converting the standard data of the network time series from the time domain to the frequency domain, time-series features can be extracted from multiple perspectives, enhancing the ability to capture abnormal patterns. When constructing an abnormal sequence detection model, it is not necessary to label abnormal data, thereby reducing costs, and it can ensure accurate identification of abnormal patterns in the case of unlabeled data or complex features. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following descriptions are only embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1Flowchart of the method for generating a network time series anomaly detection model according to an embodiment of the present disclosure;
[0020] Figure 2 Structural schematic diagram of a time series anomaly detection system according to an embodiment of the present disclosure;
[0021] Figure 3 Flowchart of the time series preprocessing method according to an embodiment of the present disclosure;
[0022] Figure 4 Flowchart of the model training method for a network time series anomaly detection model according to an embodiment of the present disclosure;
[0023] Figure 5 Structural schematic diagram of a Transformer encoder according to an embodiment of the present disclosure;
[0024] Figure 6 Structural schematic diagram of a mapper according to an embodiment of the present disclosure;
[0025] Figure 7 Structural schematic diagram of a generating device for a network time series anomaly detection model according to an embodiment of the present disclosure;
[0026] Figure 8 Structural schematic diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0027] To make the objectives, technical solutions and advantages of the present disclosure more clear and understandable, the following further describes the present disclosure in detail with reference to specific embodiments and the accompanying drawings.
[0028] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0029] Based on the description of the background technology, network time series anomaly detection aims to identify abnormal behaviors in network time series data that are significantly different from the normal patterns. However, due to the high-dimensionality and complexity of network data, traditional network time series anomaly detection methods often have limited effectiveness. As a self-supervised learning method, contrastive learning obtains useful feature representations by learning the similarities between samples. In time series anomaly detection, applying contrastive learning can, in the absence of labeled data, learn the latent structure of network time series by performing different perspective transformations or data augmentations on the same data. In this way, the model can more effectively capture the differences between normal and abnormal patterns, and detect abnormal behaviors in network time series by setting thresholds based on the trained model, improving the accuracy and robustness of anomaly detection.
[0030] Traditional contrastive learning methods usually rely on positive and negative sample pairs to learn feature representations by maximizing the similarity between positive samples and minimizing the similarity between negative samples. However, this method has some problems in practical applications. For example, it requires a large number of negative samples, which may lead to sampling bias, and it is difficult to clearly define positive and negative samples in some cases. In network time series anomaly detection, these problems are particularly prominent because network time series often contain multi-dimensional data in multiple dimensions such as CPU, bandwidth, and memory. The annotation of abnormal data cannot only consider the abnormality in a single dimension, but often requires comprehensive judgment of multiple dimensions, which brings a high cost to the annotation of abnormal data.
[0031] The following are the glossary explanations related to this disclosure:
[0032] Contrastive learning is a self-supervised learning method that aims to learn useful feature representations by comparing the similarities between samples. Some contrastive learning methods (such as SimCLR, BYOL) no longer rely on positive and negative sample pairs, but learn representations by performing different data augmentations on the same data. This method uses a large amount of unlabeled data to train the model, reducing the dependence on manual annotation, and has achieved remarkable results in fields such as computer vision and natural language processing, improving the generalization ability and robustness of the model.
[0033] Network time series anomaly detection aims to identify abnormal patterns or behaviors in network time series data. Common detection methods are mainly divided into two categories: predictive and reconstructive. Predictive methods use historical data to predict future values and compare the prediction results with the actual values to detect anomalies; reconstructive methods reconstruct the input data by building a model and use the reconstruction error to judge anomalies. In addition, there are also methods based on statistical features and density estimation.
[0034] The Fourier transform is a mathematical tool used to convert a time-domain signal into a frequency-domain representation. By decomposing the original signal into a combination of sine waves and cosine waves of different frequencies, it reveals the frequency components and periodic characteristics of the signal. In the fields of signal processing, communication, and physics, the Fourier transform is widely used to analyze and process continuous or discrete time-series data. Using the Fourier transform, one can gain a deeper understanding of the spectral characteristics of the signal, providing an important theoretical basis for data analysis and processing.
[0035] As described above, how to avoid the high cost caused by annotating abnormal data when constructing an abnormal sequence detection model has become an important research issue.
[0036] Based on the above description, as Figure 1 shown, the method for generating a network time-series abnormal detection model proposed in this embodiment includes:
[0037] Step 101, obtain a network time series, extract multi-dimensional data from the network time series, and perform padding processing and conversion processing on the multi-dimensional data to obtain standard data.
[0038] Specifically, Figure 2 is a schematic structural diagram of the time-series abnormal detection system according to an embodiment of the present disclosure. As Figure 2 shown,
[0039] Figure 3 is a flowchart of the time-series preprocessing method according to an embodiment of the present disclosure. As Figure 3 shown, obtain multi-dimensional data such as CPU, bandwidth, and memory of the network time series and corresponding labels from the original data set, perform corresponding missing value and outlier processing, convert the data into a unified tensor format, and adjust the dimensions to ensure that the channel positions meet the input requirements of the model.
[0040] Use most of the data in the training data as the training set and the remaining part as the validation set. After the model reads in the data set and the validation set, obtain the sample data of the network time series and the corresponding labels from the data set, and randomly shuffle the network time series to ensure the randomness of data input during model training and reduce the bias caused by the data order.
[0041] Align the lengths of the sample data of the network time series to unify the lengths of all input network time series. By setting a parameter with a fixed length, only retain the data of the preset number of network time points in the front of each sample data to ensure the consistency of the input data in the time dimension.
[0042] Step 102, convert the standard data from the time domain to the frequency domain through the Fourier transform to obtain time-domain features and frequency-domain features.
[0043] In specific implementation, the Fourier transform is used to transform multi-dimensional data such as CPU, bandwidth, and memory of the network time series from the time domain to the frequency domain, obtaining time-domain features and frequency-domain features. In this way, by extracting the frequency-domain features, a dual perspective required for contrastive learning can be constructed. Among them, x(n) represents the time-domain signal, X(f) represents the corresponding frequency-domain signal, N is the total number of samples of the signal, and j is the imaginary unit. This process decomposes the time-domain signal into a superposition of different frequency components, revealing the frequency characteristics of the signal.
[0044] Step 103: Perform enhancement processing on the time-domain features to obtain time-domain samples, and perform enhancement processing on the frequency-domain features to obtain frequency-domain samples.
[0045] In specific implementation, enhancement processing is respectively performed on the time-domain features and the frequency-domain features to obtain time-domain samples and frequency-domain samples.
[0046] For the time-domain samples, the time-domain features are enhanced by adding Gaussian white noise to the time-domain features to obtain time-domain samples.
[0047] For the frequency-domain samples, the frequency part in the frequency-domain features is removed, and frequency noise is added to the frequency-domain features to perform enhancement processing on the frequency-domain features to obtain frequency-domain samples.
[0048] Step 104: Through a contrastive learning algorithm, train the initial anomaly detection model according to the time-domain samples and the frequency-domain samples to obtain a network time series anomaly detection model.
[0049] In specific implementation, the time-domain samples are input into the time-domain encoder to obtain time-domain feature representations, and the frequency-domain samples are input into the frequency-domain encoder to obtain frequency-domain feature representations. The time-domain feature representations are input into the mapper to obtain time-domain latent space representations, and the frequency-domain feature representations are input into the mapper to obtain frequency-domain latent space representations. Based on the time-domain feature representations, frequency-domain feature representations, time-domain latent space representations, and frequency-domain latent space representations, a weighted loss function for contrastive learning is obtained. According to the weighted loss function, the model parameters of the initial anomaly detection model are updated through backpropagation to obtain an updated anomaly detection model, and a validation set is used to determine whether the updated anomaly detection model is overfitting. When the updated anomaly detection model is overfitting, the updated anomaly detection model is used as the network time series anomaly detection model. When the updated anomaly detection model is not overfitting but the number of iterations reaches a preset number threshold, the updated anomaly detection model is used as the network time series anomaly detection model.
[0050] Through the above embodiments, a network time series is obtained, multidimensional data is extracted from the network time series, and the multidimensional data is subjected to filling processing and conversion processing to obtain standard data. The standard data is transformed from the time domain to the frequency domain through Fourier transform to obtain time domain features and frequency domain features. The time domain features are enhanced to obtain time domain samples, and the frequency domain features are enhanced to obtain frequency domain samples. Through a contrast learning algorithm, an initial anomaly detection model is trained according to the time domain samples and frequency domain samples to obtain a network time series anomaly detection model. In this way, by transforming the standard data of the network time series from the time domain to the frequency domain, time series features can be extracted from multiple perspectives, the ability to capture abnormal patterns can be enhanced, and abnormal data does not need to be labeled when constructing an abnormal sequence detection model, thereby reducing high costs and ensuring accurate identification of abnormal patterns in the case of unlabeled data or complex features.
[0051] In some embodiments, step 103 includes:
[0052] Step 1031, the time domain features are enhanced by adding Gaussian white noise to the time domain features to obtain time domain samples.
[0053] Step 1032, the frequency domain features are enhanced by moving out the frequency part in the frequency domain features and adding frequency noise to the frequency domain features to obtain frequency domain samples.
[0054] Specifically, in the training stage, the time domain features and frequency domain features are respectively subjected to data augmentation to obtain time domain samples and frequency domain samples. In this way, variant samples from multiple perspectives can be generated.
[0055] In the time domain part, the time domain features are enhanced by adding Gaussian white noise to the time domain features to obtain time domain samples. Adding Gaussian white noise is a commonly used data augmentation method. Gaussian white noise is a random noise with a Gaussian distribution, whose mean is 0 and the variance can be adjusted according to the actual situation. The specific formula is: μ is the mean, which determines the central position of the data distribution; σ is the standard deviation, which measures the degree of dispersion of the data. By adding Gaussian white noise, the noise interference in the actual environment can be simulated, enabling the model to learn more robust feature representations and improving the model's adaptability to data noise.
[0056] In the frequency domain, the frequency domain features are enhanced by removing the frequency components in the frequency domain features and adding frequency noise to obtain frequency domain samples. The enhanced time domain samples and frequency domain samples are used for contrast learning. By comparing the data features from different perspectives, the model can better learn the internal laws of the data, thereby enhancing the detection ability and robustness for abnormal situations.
[0057] Through the above solution, by adding Gaussian white noise to the time-domain features, it is possible to simulate the noise interference in the actual environment, thereby realizing the enhancement processing of the time-domain features, enabling the model to learn more robust feature representations, and improving the adaptability of the model to data noise. By removing the frequency part from the frequency-domain features and adding frequency noise to the frequency-domain features, it is possible to realize the enhancement processing of the frequency-domain features. In this way, the enhanced time-domain samples and frequency-domain samples are used for contrastive learning. By comparing the data features from different perspectives, the model can better learn the internal laws of the data, thereby improving the detection ability and robustness to abnormal situations.
[0058] In some embodiments, step 104 includes:
[0059] Step 1041, inputting the time-domain samples into a time-domain encoder to obtain a time-domain feature representation, and inputting the frequency-domain samples into a frequency-domain encoder to obtain a frequency-domain feature representation.
[0060] Step 1042, inputting the time-domain feature representation into a mapper to obtain a time-domain latent space representation, and inputting the frequency-domain feature representation into a mapper to obtain a frequency-domain latent space representation.
[0061] Step 1043, obtaining a weighted loss function for contrastive learning based on the time-domain feature representation and the frequency-domain feature representation.
[0062] Step 1044, according to the weighted loss function, updating the model parameters of the initial anomaly detection model through backpropagation to obtain an updated anomaly detection model, and using a validation set to determine whether the updated anomaly detection model is overfitting.
[0063] Step 1045, in response to determining that the updated anomaly detection model is overfitting, using the updated anomaly detection model as the network time-series anomaly detection model.
[0064] Step 1046, in response to determining that the updated anomaly detection model is not overfitting but the number of iterations reaches a preset number threshold, using the updated anomaly detection model as the network time-series anomaly detection model.
[0065] Specifically, when implemented, Figure 4 is a flowchart of the model training method for the network time-series anomaly detection model of the present disclosure embodiment. As Figure 4As shown, time-domain samples and frequency-domain samples are sent to a data loader, which forwards the time-domain samples to a Transformer time-domain encoder and the frequency-domain samples to a Transformer frequency-domain encoder. The Transformer time-domain encoder encodes the time-domain samples to obtain a time-domain feature representation and sends it to the corresponding mapper. The Transformer frequency-domain encoder encodes the frequency-domain samples to obtain a frequency-domain feature representation and sends it to the corresponding mapper. The mapper maps the time-domain feature representation to obtain a time-domain latent space representation, and the mapper maps the frequency-domain feature representation to obtain a frequency-domain latent space representation. A weighted contrast loss is calculated based on the time-domain latent space representation and the frequency-domain latent space representation, and the initial anomaly detection model is updated according to the weighted contrast loss. The updated anomaly detection model is verified according to a validation set, and it is determined whether the updated anomaly detection model is overfitting. When the updated anomaly detection model is overfitting, the training is completed, and the updated anomaly detection model is used as the network time series anomaly detection model. When the updated anomaly detection model is not overfitting, it is determined whether the number of iterations reaches a preset number threshold. When the number of iterations reaches the preset number threshold, the training is completed, and the updated anomaly detection model is used as the network time series anomaly detection model. After the model training is completed, the network time series anomaly detection model is used to detect abnormal data in a test set. When the number of iterations does not reach the preset number threshold, the step of sending the time-domain samples and the frequency-domain samples to the data loader is returned, and the updated anomaly detection model is continuously trained.
[0066] Through the above solution, the Transformer time-domain encoder encodes the time-domain samples to obtain a time-domain feature representation, and the Transformer frequency-domain encoder encodes the frequency-domain samples to obtain a frequency-domain feature representation, which can simultaneously focus on multi-dimensional information of the network time series, so as to better identify abnormal data. In this way, it is not necessary to label the network time series, which can reduce costs, and can also capture multi-scale features of the data, improve the sensitivity of the model to abnormal data, and enhance the accuracy and robustness of anomaly detection.
[0067] In some embodiments, step 1041 includes:
[0068] Step 1041A, inputting the time-domain feature and the time-domain sample into a time-domain encoder, and processing the time-domain sample through a self-attention mechanism to obtain a time-domain vector.
[0069] Step 1041B, performing normalization processing on the time-domain vector to obtain a normalized time-domain vector;
[0070] Step 1041C: Map the normalized time-domain vector through a feed-forward neural network module to obtain a high-dimensional time-domain feature representation and an enhanced time-domain feature representation.
[0071] Step 1041D: Input the frequency-domain features and the frequency-domain samples into a frequency-domain encoder, and process the frequency-domain samples through a self-attention mechanism to obtain a frequency-domain vector.
[0072] Step 1041E: Perform normalization processing on the frequency-domain vector to obtain a normalized frequency-domain vector;
[0073] Step 1041F: Map the normalized frequency-domain vector through a feed-forward neural network module to obtain a high-dimensional frequency-domain feature representation and an enhanced frequency-domain feature representation.
[0074] During specific implementation, Figure 5 is a schematic structural diagram of the Transformer encoder according to an embodiment of the present disclosure. As Figure 5 shown, the input network time series first enters the self-attention mechanism module (Self Attention). For each element in the network time series, a query vector (Query, Q), a key vector (Key, K), and a value vector (Value, V) are respectively generated through a linear transformation. Then, through the self-attention formula the output of the self-attention mechanism is obtained, where d k is the dimension of the Key vector, which is used to scale the attention scores. In the multi-head self-attention mechanism, multiple different attention heads will execute the above calculation process in parallel, and the formula is MultiHead(Q, K, V) = Concat(head i , …, head h ), and the final output of the multi-head self-attention mechanism is obtained after combination.
[0075] The output of the multi-head attention mechanism is connected with the original input through a residual connection, which can retain the original input information and avoid problems such as gradient disappearance. Then, it enters the normalization layer (Layer Norm) for normalization operation, and the feature dimensions of each sample are normalized to adjust the data distribution to a range with a mean of 0 and a variance of 1, which helps to accelerate model training and improve stability.
[0076] The output after passing through the normalization layer enters the feed-forward neural network module (Feed Forward). The feed-forward neural network consists of two linear transformations, and the ReLU activation function is usually used in the middle. The input data is further transformed in terms of features, and the expression ability of the model is increased through the non-linear activation function, and the input is mapped to a new feature space to generate a high-dimensional time-domain feature representation h of the original dataT , high-dimensional frequency domain feature representation h F , enhanced time domain feature representation of the enhanced data and enhanced frequency domain feature representation The four embeddings output enter the mapper as shown in Figure 6 shown and are finally mapped to the same latent space.
[0077] Through the above scheme,
[0078] In some embodiments, step 1042 includes:
[0079] Step 1042A, input the time domain feature representation into the mapper, and map the time domain feature representation to the latent space through the first linear layer to obtain the mapped time domain feature representation.
[0080] Step 1042B, perform normalization processing on the mapped time domain feature representation through the normalization layer to obtain the normalized time domain feature representation.
[0081] Step 1042C, set the time domain feature representation with a negative value to zero through the activation function layer, and keep the time domain feature representation with a positive value to obtain the set time domain feature representation.
[0082] Step 1042D, perform linear transformation processing on the set time domain feature representation through the second linear layer to obtain the original time domain latent space representation and the enhanced time domain latent space representation.
[0083] Step 1042E, input the frequency domain feature representation into the mapper, and map the frequency domain feature representation to the latent space through the first linear layer to obtain the mapped frequency domain feature representation.
[0084] Step 1042F, perform normalization processing on the mapped frequency domain feature representation through the normalization layer to obtain the normalized frequency domain feature representation.
[0085] Step 1042G, set the frequency domain feature representation with a negative value to zero through the activation function layer, and keep the frequency domain feature representation with a positive value to obtain the set frequency domain feature representation.
[0086] Step 1042H, perform linear transformation processing on the set frequency domain feature representation through the second linear layer to obtain the original frequency domain latent space representation and the enhanced frequency domain latent space representation.
[0087] Specifically, when implemented, Figure 6 is a schematic structural diagram of the mapper of the embodiment of the present disclosure. As shown in Figure 6As shown in the figure, the time-domain features and frequency-domain features are input into the mapper. The embeddings (time-domain feature representation and frequency-domain feature representation) of the input mapper first enter the first linear layer (Linear). The first linear layer performs a linear transformation on the input vectors (time-domain feature representation and frequency-domain feature representation) through matrix multiplication, maps the input vectors to a new feature space, changes the dimension of the vectors or the form of feature representation, and obtains the mapped time-domain feature representation and the mapped frequency-domain feature representation.
[0088] The mapped time-domain feature representation and the mapped frequency-domain feature representation then enter the batch normalization layer (BatchNorm), and each feature dimension is normalized on the batch data to obtain the normalized time-domain feature representation and the normalized frequency-domain feature representation. In this way, the data distribution is stabilized, the mean is 0, and the variance is 1, which helps to accelerate training and prevent the problems of gradient vanishing or explosion.
[0089] The data after batch normalization enters the ReLU activation function layer. The ReLU function sets all negative input values to 0 and keeps positive input values unchanged, obtaining the set time-domain feature representation and the set frequency-domain feature representation. In this way, non-linear factors are introduced to enhance the expressive power of the model.
[0090] Finally, a linear transformation is performed again through the second linear layer (Linear) to further adjust the feature representation, and finally the four embedding vectors are mapped to the same latent space to generate the original time-domain latent space representation z T and the original frequency-domain latent space representation z F and the enhanced time-domain latent space representation and the enhanced frequency-domain latent space representation
[0091] Through the above solution, the time-domain feature representation and the frequency-domain feature representation are respectively mapped to the latent space through the first linear layer, and the dimension or the form of feature representation of the time-domain feature representation can be changed. By normalizing the mapped time-domain feature representation and the mapped frequency-domain feature representation through the normalization layer, the distributions of the normalized time-domain feature representation and the normalized frequency-domain feature representation can be stabilized, which helps to accelerate training and prevent the problems of gradient vanishing or explosion. By setting the time-domain feature representation or the frequency-domain feature representation with negative values to zero through the activation function layer and keeping the time-domain feature representation or the frequency-domain feature representation with positive values, non-linear factors can be introduced to enhance the expressive power of the model. By performing a linear transformation on the set time-domain feature representation or the set frequency-domain feature representation through the second linear layer, the time-domain feature representation and the frequency-domain feature representation can be further adjusted.
[0092] In some embodiments, step 1043 includes:
[0093] Step 1043A: Obtain the time-domain loss function of the time-domain features based on the similarity between the high-dimensional time-domain feature representation and the enhanced time-domain feature representation in the time-domain feature representation.
[0094] Step 1043B: Obtain the frequency-domain loss function of the frequency-domain features based on the similarity between the high-dimensional frequency-domain feature representation and the enhanced frequency-domain feature representation in the frequency-domain feature representation.
[0095] Step 1043C: Obtain the cross-domain contrast loss function based on the similarity between the time-domain latent space representation and the frequency-domain latent space representation.
[0096] Step 1043D: Obtain the weighted loss function according to the time-domain loss function, the frequency-domain loss function, and the cross-domain contrast loss function.
[0097] In specific implementation, the model calculates the contrastive learning loss through the normalized temperature-scaled cross-entropy loss (NTXentLoss). The specific formula is where the similarity function is the cosine similarity, which is used to measure the angular similarity between two vectors u and v. The value of the similarity function ranges between -1 and 1. 1 indicates complete similarity, and -1 indicates complete dissimilarity; represents a binary function. When i ≠ j, has a value of 1, indicating that the sample itself is excluded during calculation, and only the comparison with other samples is considered; the temperature parameter τ is a scaling factor used to control the scaling of the similarity. A smaller τ will amplify the differences between similarities and affect the learning dynamics of the model. x j ∈D ta represents other samples or their enhanced samples used in contrastive learning.
[0098] Calculate different loss functions loss through the loss function. The specific selection of positive and negative samples is as Figure 2 shown. The red dashed line represents the positive sample pair, and the blue dashed line represents the negative sample pair. Among them, the time-domain loss function loss_t is the similarity between the high-dimensional time-domain feature representation h T of the original data and the enhanced time-domain feature representation of the enhanced data. The frequency-domain loss function loss_f is the contrast loss between the high-dimensional frequency-domain feature representation h F of the original data and the enhanced frequency-domain feature representation of the enhanced data. The cross-domain contrast loss function I_TF is used to measure the original time-domain latent space representation z T and the original frequency-domain latent space representation z F , the enhanced time-domain latent space representation and the enhanced frequency-domain latent space representation The similarity between them encourages the features in different domains to be close in the latent space. The final weighted loss function value loss is calculated by the formula loss = lam * (loss_t + loss_f) + I_TF, where lam is the weighting factor, and by default, it can be taken as 0.2. In each iteration, the model parameters of the initial anomaly detection model are gradually optimized according to the weighted loss function loss through backpropagation, and it is evaluated on the validation set whether the updated anomaly detection model is overfitting. If the performance on the validation set no longer improves within a preset number of update cycles, the training is stopped in a timely manner.
[0099] Through the above solution, a weighted loss function is obtained based on the time-domain loss function, the frequency-domain loss function, and the cross-domain contrast loss function. In this way, the similarities between time-domain feature representations, the similarities between frequency-domain feature representations, and the similarities between cross-domain feature representations can be comprehensively considered, making the weighted loss function more accurate.
[0100] In some embodiments, after step 104, it further includes:
[0101] Step 105, obtaining the time series to be detected, and inputting the time series to be detected into the network time series anomaly detection model.
[0102] Step 106, using the network time series anomaly detection model to map the multi-dimensional data in the time series to be detected into the latent space, and determining the loss value of the time series to be detected in the latent space.
[0103] Step 107, determining whether the loss value is greater than a preset loss threshold, where the preset loss threshold is determined according to the hyperparameters of the network time series anomaly detection model.
[0104] Step 108, in response to determining that the loss value is greater than the preset loss threshold, determining that the time series to be detected is an abnormal time series.
[0105] Specifically, when implementing, after completing the model training and optimization, the abnormal sequences of the network time series on the test set are detected by setting the loss threshold. During model training, the weighted loss function loss of each batch is saved in the threshold list in ascending order, and the hyperparameter anormly_ratio is set as the abnormal ratio value of the data set. The loss value at the corresponding percentage position in the threshold list is found according to the hyperparameter as the preset loss threshold.
[0106] Obtain the corresponding data in the test set (the time series to be detected), and input the corresponding data in the test set into the trained network time series anomaly detection model, so that the multi-dimensional data such as CPU, bandwidth, memory, etc. are finally mapped to the same latent space, and the corresponding loss value loss is obtained. Finally, the multi-dimensional data mapping of the network time series is calculated as a unified loss value loss, so that the abnormal data of the network time series can be judged as a whole. If the loss value loss is greater than the preset loss threshold, it means that the current network time series data is an abnormal time series sequence.
[0107] Use time-frequency consistency as a dual perspective for comparative learning. The network time series is converted to the frequency domain to obtain different representations of the same data, namely the time domain and the frequency domain. Then, through the Transformer's multi-attention head mechanism, the multi-dimensional information of the network time series is simultaneously focused on, so as to better identify anomalies. Using these two perspectives for comparative learning, richer feature representations can be learned without relying on negative samples. In this way, there is no need to label the network time series, which can reduce costs, capture the multi-scale features of the data, improve the model's sensitivity to abnormal patterns, and enhance the accuracy and robustness of anomaly detection.
[0108] Fourier transform is used to construct the dual perspectives required for contrastive learning. By focusing on multi-dimensional information such as CPU, bandwidth, and memory of complex network time series in the neural network, and bringing the representation of the same sample in the time domain and frequency domain closer, the overall model training is completed. Finally, the anomaly detection of the network time series is completed through the threshold method.
[0109] Through the above scheme, the pre-set loss threshold can be accurately determined according to the hyperparameters of the network time series anomaly detection model. The network time series anomaly detection model is used to map the multidimensional data in the time series to be detected to the latent space, and determine the loss value of the time series to be detected in the latent space. When the loss value is greater than the pre-set loss threshold, the time series to be detected is determined to be an abnormal time series. In this way, the network time series anomaly detection model can accurately detect anomalies in the time series to be detected.
[0110] Through the above embodiments, a network time series is obtained, multi-dimensional data is extracted from the network time series, and the multi-dimensional data is subjected to filling processing and conversion processing to obtain standard data. The standard data is transformed from the time domain to the frequency domain through Fourier transform to obtain time-domain features and frequency-domain features. The time-domain features are enhanced to obtain time-domain samples, and the frequency-domain features are enhanced to obtain frequency-domain samples. Through a contrastive learning algorithm, an initial anomaly detection model is trained according to the time-domain samples and frequency-domain samples to obtain a network time series anomaly detection model. In this way, by transforming the standard data of the network time series from the time domain to the frequency domain, time series features can be extracted from multiple perspectives, the ability to capture abnormal patterns can be enhanced, and it is not necessary to label abnormal data when constructing an abnormal sequence detection model, thereby reducing costs and ensuring accurate identification of abnormal patterns in the case of unlabeled data or complex features.
[0111] It should be noted that the embodiments of the present disclosure can be further described in the following manner:
[0112] The main application scenario of the embodiments of the present disclosure is to support the anomaly detection of network time series. A neural network model trained through contrastive learning is used to accurately identify abnormal network time series according to the calculated threshold. The overall algorithm architecture is as Figure 2 .
[0113] Before training the contrastive learning model, it is necessary to perform data preprocessing and enhancement on the two required perspectives. The algorithm flowchart is as Figure 3 , and the processing steps are as follows:
[0114] Step A: Obtain multi-dimensional information such as CPU, bandwidth, and memory of the network time series and corresponding labels from the original dataset, perform corresponding missing value and outlier processing, convert the data into a unified tensor format, and adjust the dimensions to ensure that the channel position meets the input requirements of the model.
[0115] Step B: Use most of the data in the training data as the training set and the remaining part as the validation set.
[0116] Step C: After the model reads the dataset and the validation set, obtain network time series samples and corresponding labels from them, and randomly shuffle these data to ensure the randomness of data input during the model training process and reduce the bias caused by the data order.
[0117] Step D: Align the lengths of the network time series samples to unify the lengths of all input network time series. By setting a parameter with a fixed length, only keep the data of the first several network time points of each sample to ensure the consistency of the input data in the time dimension.
[0118] Step E: Use the Fourier transform Convert the multi-dimensional data such as CPU, bandwidth, and memory of the original network time series from the time domain to the frequency domain, extract the frequency domain features, and construct the dual perspectives required for contrastive learning. Among them, x(n) represents the time domain signal, X(f) represents the corresponding frequency domain signal, N is the total number of signal samples, and j is the imaginary unit. This process decomposes the time domain signal into the superposition of different frequency components, revealing the frequency characteristics of the signal.
[0119] Step F, in the training stage, generate variant samples with multiple perspectives by performing data augmentation on the time domain and frequency domain data respectively. In the time domain part, adding Gaussian white noise is a commonly used data augmentation method. Gaussian white noise is a random noise with a Gaussian distribution, whose mean is 0 and the variance can be adjusted according to the actual situation. The specific formula is: μ is the mean, which determines the central position of the data distribution; σ is the standard deviation, which measures the degree of dispersion of the data. By adding Gaussian white noise, the noise interference in the actual environment can be simulated, enabling the model to learn more robust feature representations and improving the model's adaptability to data noise. In the frequency domain, data augmentation is performed by removing frequency components and adding frequency noise. These augmented data are used for contrastive learning. By comparing the data features from different perspectives, the model can better learn the internal laws of the data, thereby enhancing the detection ability and robustness for abnormal situations.
[0120] The algorithm flow chart for model training and abnormal sequence identification through contrastive learning is as Figure 4 , and the detailed steps are as follows:
[0121] Step a, read the time domain samples and frequency domain samples after data processing, and set the number of iterative training times. In each iteration, the model receives the original data data and its augmented version aug1, and simultaneously processes the frequency domain data data_f and its augmented version aug1_f. These data are input into two independent encoders respectively, which are used to extract the feature representations of the time domain and frequency domain to obtain four embedding vectors. The encoder can be composed of multiple stacked Transformer layers, and the architecture of each layer is as Figure 5 .
[0122] Step b, the input time series data first enters the self-attention mechanism module. For each element in the sequence, a query vector (Query, Q), a key vector (Key, K), and a value vector (Value, V) are respectively generated through linear transformation. Then, through the self-attention formula: Obtain the output of the self-attention mechanism, d k is the dimension of the Key vector, For scaling the attention scores. In the multi - head self - attention mechanism, multiple different attention heads execute the above - mentioned calculation process in parallel, and its formula is MultiHead(Q, K, V)=Concat(head i ,…,head h ), and after combination, the final output of the multi - head self - attention mechanism is obtained.
[0123] Step c, the output of the multi - head attention mechanism is connected with the original input through a residual connection, which can retain the original input information and avoid problems such as gradient vanishing. Then, a Layer Norm layer normalization operation is performed to normalize each sample's feature dimension, adjusting the data distribution to a range with a mean of 0 and a variance of 1, which helps to accelerate model training and improve stability.
[0124] Step d, the output after layer normalization enters the feed - forward neural network module. The feed - forward neural network consists of two linear transformations, and usually uses the ReLU activation function in the middle. It further transforms the input data features, increases the model's expressive ability through the non - linear activation function, maps the input to a new feature space to generate a high - dimensional feature representation h T and h F , enhances the time - domain and frequency - domain feature representations of the data and The four embeddings of the output enter the mapper as shown in Figure 6 and are finally mapped to the same latent space.
[0125] Step e, the embeddings input to the mapper first enter the first linear layer (Linear), which linearly transforms the input vector through matrix multiplication, maps it to a new feature space, and changes the dimension or feature representation form of the vector.
[0126] Step f, then it enters the batch normalization layer, which normalizes each feature dimension on the batch data to make the data distribution stable, with a mean of 0 and a variance of 1, helping to accelerate training and prevent problems such as gradient vanishing or explosion.
[0127] Step g, the data after batch normalization enters the ReLU activation function layer. The ReLU function sets all negative input values to 0 and keeps positive input values unchanged, introducing non - linear factors and enhancing the model's expressive ability.
[0128] Step h, finally, a linear transformation is performed again through the second linear layer to further adjust the feature representation, and finally the four embedding vectors are mapped to the same latent space to generate the original time - frequency domain latent space representation z T 、z F and the enhanced time - frequency domain latent space representation and
[0129] Step i, the model calculates the contrastive learning loss through NTXentLoss (Normalized Temperature Scaled Cross Entropy Loss). The specific formula is where the similarity function is the cosine similarity, which is used to measure the angular similarity between two vectors u and v. Its value ranges from -1 to 1, where 1 indicates complete similarity and -1 indicates complete dissimilarity. denotes a binary function, whose value is 1 when i≠j, indicating that the sample itself is excluded during calculation and only the comparison with other samples is considered. The temperature parameter τ is a scaling factor used to control the scaling of the similarity. A smaller τ will amplify the differences between similarities and affect the learning dynamics of the model. x j ∈D ta represents other samples used in contrastive learning or their augmented samples.
[0130] Step j, different losses are calculated through the loss function. The selection of positive and negative samples is as shown in Figure 2 . The red dashed line represents the positive sample pair, and the blue dashed line represents the negative sample pair. Among them, the loss loss_t of the time-domain feature calculates the similarity between the original time-domain feature h T and the augmented feature . The loss loss_f of the frequency-domain feature calculates the contrastive loss between h F and . The cross-domain contrastive loss I_TF is used to measure the similarity between the time-domain latent representation z T and the frequency-domain latent representation z F , the augmented time-domain latent representation and the augmented frequency-domain latent representation , encouraging the features in different domains to be close in the latent space. The final loss value is calculated by the formula loss = lam * (loss_t + loss_f) + I_TF, where lam is the weighted weight, which can be taken as 0.2 by default. In each iteration, the model parameters are gradually optimized through backpropagation according to the loss, and it is evaluated on the validation set whether the model is overfitting. If the performance on the validation set no longer improves within several epochs, the training is stopped in time.
[0131] Step k, after completing the model training and optimization, the abnormal sequences of the network time series on the test set are detected by setting thresholds. When the model is trained, the loss function of each batch is saved in the threshold list in ascending order. The hyperparameter anormly_ratio is set as the abnormal ratio value of the dataset, and the loss value corresponding to the percentage position of the threshold list is found as the threshold according to this hyperparameter.
[0132] Test the model with the corresponding data in the test set, so as to finally map multi-dimensional data such as CPU, bandwidth, and memory to the same latent space, and obtain the corresponding loss. Finally, the multi-dimensional data mapping calculation of the network time series is calculated as a unified loss value, so that the abnormal data of the network time series can be judged overall. If the loss is greater than the set threshold, it means that the data here is an abnormal network time series.
[0133] Through the above embodiments, in view of the deficiencies of the traditional network time series anomaly detection method in multi-perspective feature extraction and data latent law learning, the embodiments of the present disclosure propose a network time series anomaly detection method based on time-frequency consistent dual perspectives and contrast learning. By fusing the time domain and frequency domain information of time series data and using the multi-attention head mechanism of Transformer to simultaneously focus on the complex multi-dimensional network information of the network time series, time series features can be extracted from multiple perspectives, enhancing the ability to capture abnormal patterns. In the anomaly detection process, by combining dual-perspective data augmentation and contrast learning, the adaptability of the model to different feature perspectives is effectively improved, reducing the risk of information loss caused by single-perspective feature extraction. The solution of the embodiments of the present disclosure improves the robustness and generalization ability of the model, ensuring that anomalies can still be accurately identified in the case of unlabeled data or complex features, and improving the detection accuracy and stability of the system.
[0134] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0135] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0136] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a device for generating a network time series anomaly detection model.
[0137] Refer to Figure 7 , the device for generating a network time series anomaly detection model includes:
[0138] The extraction module 301 is configured to obtain a network time series, extract multi-dimensional data from the network time series, and perform padding processing and conversion processing on the multi-dimensional data to obtain standard data;
[0139] The conversion module 302 is configured to convert the standard data from the time domain to the frequency domain through Fourier transform to obtain time domain features and frequency domain features;
[0140] The enhancement module 303 is configured to perform enhancement processing on the time domain features to obtain time domain samples, and perform enhancement processing on the frequency domain features to obtain frequency domain samples;
[0141] The training module 304 is configured to train an initial anomaly detection model based on the time domain samples and the frequency domain samples through a contrast learning algorithm to obtain a network time series anomaly detection model.
[0142] In some embodiments, the enhancement module 303 includes:
[0143] The time domain enhancement unit is configured to perform enhancement processing on the time domain features by adding Gaussian white noise to the time domain features to obtain time domain samples;
[0144] The frequency domain enhancement unit is configured to perform enhancement processing on the frequency domain features by moving out the frequency part in the frequency domain features and adding frequency noise to the frequency domain features to obtain frequency domain samples.
[0145] In some embodiments, the training module 304 includes:
[0146] The encoding unit is configured to input the time domain samples into a time domain encoder to obtain a time domain feature representation, and input the frequency domain samples into a frequency domain encoder to obtain a frequency domain feature representation;
[0147] The mapping unit is configured to input the time domain feature representation into a mapper to obtain a time domain latent space representation, and input the frequency domain feature representation into a mapper to obtain a frequency domain latent space representation;
[0148] The loss function determination unit is configured to obtain a weighted loss function for contrast learning based on the time domain feature representation and the frequency domain feature representation;
[0149] The update unit is configured to update the model parameters of the initial anomaly detection model through backpropagation according to the weighted loss function to obtain an updated anomaly detection model, and use a validation set to determine whether the updated anomaly detection model is overfitting;
[0150] The first model generation unit is configured to, in response to determining that the updated anomaly detection model is overfitted, use the updated anomaly detection model as the network time series anomaly detection model;
[0151] The second model generation unit is configured to, in response to determining that the updated anomaly detection model is not overfitted but the number of iterations reaches a preset number threshold, use the updated anomaly detection model as the network time series anomaly detection model.
[0152] In some embodiments, the encoding unit is specifically configured to:
[0153] Input the time domain feature and the time domain sample into the time domain encoder, and process the time domain sample through the self-attention mechanism to obtain a time domain vector;
[0154] Perform normalization processing on the time domain vector to obtain a normalized time domain vector;
[0155] Perform mapping processing on the normalized time domain vector through a feed-forward neural network module to obtain a high-dimensional time domain feature representation and an enhanced time domain feature representation;
[0156] Input the frequency domain feature and the frequency domain sample into the frequency domain encoder, and process the frequency domain sample through the self-attention mechanism to obtain a frequency domain vector;
[0157] Perform normalization processing on the frequency domain vector to obtain a normalized frequency domain vector;
[0158] Perform mapping processing on the normalized frequency domain vector through a feed-forward neural network module to obtain a high-dimensional frequency domain feature representation and an enhanced frequency domain feature representation.
[0159] In some embodiments, the mapping unit is specifically configured to:
[0160] Input the time domain feature representation into the mapper, and map the time domain feature representation to the latent space through the first linear layer to obtain a mapped time domain feature representation;
[0161] Perform normalization processing on the mapped time domain feature representation through a normalization layer to obtain a normalized time domain feature representation;
[0162] Set the time domain feature representation with a negative value to zero through the activation function layer, and keep the time domain feature representation with a positive value to obtain a set time domain feature representation;
[0163] Perform linear transformation processing on the set time domain feature representation through the second linear layer to obtain an original time domain latent space representation and an enhanced time domain latent space representation;
[0164] Input the frequency-domain feature representation into a mapper, and map the frequency-domain feature representation to a latent space through a first linear layer to obtain a mapped frequency-domain feature representation;
[0165] Normalize the mapped frequency-domain feature representation through a normalization layer to obtain a normalized frequency-domain feature representation;
[0166] Set the frequency-domain feature representations with negative values to zero through an activation function layer, and keep the frequency-domain feature representations with positive values to obtain a set frequency-domain feature representation;
[0167] Perform a linear transformation on the set frequency-domain feature representation through a second linear layer to obtain an original frequency-domain latent space representation and an enhanced frequency-domain latent space representation.
[0168] In some embodiments, the loss function determination unit is specifically configured to:
[0169] Based on the similarity between the high-dimensional time-domain feature representation and the enhanced time-domain feature representation in the time-domain feature representation, obtain a time-domain loss function for the time-domain feature;
[0170] Based on the similarity between the high-dimensional frequency-domain feature representation and the enhanced frequency-domain feature representation in the frequency-domain feature representation, obtain a frequency-domain loss function for the frequency-domain feature;
[0171] Based on the similarity between the time-domain latent space representation and the frequency-domain latent space representation, obtain a cross-domain contrast loss function;
[0172] Obtain a weighted loss function according to the time-domain loss function, the frequency-domain loss function, and the cross-domain contrast loss function.
[0173] In some embodiments, after training the neural network model according to the time-domain samples and the frequency-domain samples through a contrast learning algorithm to obtain a network time series anomaly detection model, the apparatus further includes:
[0174] An acquisition module, configured to acquire a time series to be detected, and input the time series to be detected into the network time series anomaly detection model;
[0175] A loss value determination module, configured to map the multi-dimensional data in the time series to be detected to a latent space by using the network time series anomaly detection model, and determine the loss value of the time series to be detected in the latent space;
[0176] A judgment module, configured to judge whether the loss value is greater than a preset loss threshold, where the preset loss threshold is determined according to the hyperparameters of the network time series anomaly detection model;
[0177] The detection module is configured to determine that the to-be-detected time series is an abnormal time series in response to determining that the loss value is greater than a preset loss threshold.
[0178] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0179] The device in the above embodiment is used to implement the method for generating the corresponding network time series anomaly detection model in any one of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0180] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for generating the network time series anomaly detection model in any one of the above embodiments.
[0181] Figure 8 FIG. shows a more specific schematic hardware structure diagram of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0182] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0183] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0184] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure), or can be externally connected to the device to provide corresponding functions. Among them, the input devices can include keyboards, mice, touchscreens, microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.
[0185] The communication interface 1040 is used to connect to the communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB (Universal Serial Bus), network cable, etc.), or can also achieve communication through wireless means (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).
[0186] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0187] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and do not necessarily include all the components shown in the figure.
[0188] The electronic device of the above embodiment is used to implement the method for generating the corresponding network timing sequence anomaly detection model in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0189] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method for generating the network timing sequence anomaly detection model as described in any of the above embodiments.
[0190] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0191] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the method for generating the network timing sequence anomaly detection model described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0192] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides a computer program product, including computer program instructions, which, when running on a computer, cause the computer to execute the method for generating the network timing sequence anomaly detection model described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0193] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0194] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be executed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0195] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0196] It is understood that the above-mentioned notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0197] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above. For the sake of brevity, they are not provided in detail.
[0198] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in the form of a block diagram in order not to make the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation manner of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (that is, these details should be completely within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0199] Although the present disclosure has been described in connection with specific embodiments of the present disclosure, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) can be used with the embodiments discussed.
[0200] The embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the present disclosure. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for generating an abnormal detection model for network time series, characterized in that, The method includes: Obtain a network time series, extract multi-dimensional data from the network time series, and perform padding processing and transformation processing on the multi-dimensional data to obtain standard data; Convert the standard data from the time domain to the frequency domain through Fourier transform to obtain time domain features and frequency domain features; Perform enhancement processing on the time domain features to obtain time domain samples, and perform enhancement processing on the frequency domain features to obtain frequency domain samples; Through a contrastive learning algorithm, train an initial anomaly detection model based on the time domain samples and the frequency domain samples to obtain a network time series anomaly detection model.
2. The method according to claim 1, characterized in that, The performing enhancement processing on the time domain features to obtain time domain samples and performing enhancement processing on the frequency domain features to obtain frequency domain samples includes: Perform enhancement processing on the time domain features by adding Gaussian white noise to the time domain features to obtain time domain samples; Perform enhancement processing on the frequency domain features by moving the frequency part in the frequency domain features and adding frequency noise to the frequency domain features to obtain frequency domain samples.
3. The method according to claim 1, characterized in that, The through a contrastive learning algorithm, training an initial anomaly detection model based on the time domain samples and the frequency domain samples to obtain a network time series anomaly detection model includes: Input the time domain samples into a time domain encoder to obtain a time domain feature representation, and input the frequency domain samples into a frequency domain encoder to obtain a frequency domain feature representation; Input the time domain feature representation into a mapper to obtain a time domain latent space representation, and input the frequency domain feature representation into a mapper to obtain a frequency domain latent space representation; Obtain a weighted loss function for contrastive learning based on the time domain feature representation and the frequency domain feature representation; According to the weighted loss function, update the model parameters of the initial anomaly detection model through backpropagation to obtain an updated anomaly detection model, and use a validation set to determine whether the updated anomaly detection model is overfitting; In response to determining that the updated anomaly detection model is overfitting, then use the updated anomaly detection model as the network time series anomaly detection model; In response to determining that the updated anomaly detection model is not overfitting but the number of iterations reaches a preset number threshold, then use the updated anomaly detection model as the network time series anomaly detection model.
4. The method according to claim 3, characterized in that, The inputting the time domain samples into a time domain encoder to obtain a time domain feature representation and inputting the frequency domain samples into a frequency domain encoder to obtain a frequency domain feature representation includes: Input the time domain features and the time domain samples into a time domain encoder, and process the time domain samples through a self-attention mechanism to obtain a time domain vector; Perform normalization processing on the time domain vector to obtain a normalized time domain vector; Perform mapping processing on the normalized time domain vector through a feed-forward neural network module to obtain a high-dimensional time domain feature representation and an enhanced time domain feature representation; Input the frequency domain features and the frequency domain samples into a frequency domain encoder, and process the frequency domain samples through a self-attention mechanism to obtain a frequency domain vector; Perform normalization processing on the frequency domain vector to obtain a normalized frequency domain vector; The normalized frequency-domain vector is processed by a feedforward neural network module to obtain a high-dimensional frequency-domain feature representation and an enhanced frequency-domain feature representation.
5. The method according to claim 3, characterized in that The inputting the time-domain feature representation into a mapper to obtain a time-domain latent space representation, and inputting the frequency-domain feature representation into a mapper to obtain a frequency-domain latent space representation includes: Inputting the time-domain feature representation into a mapper, and mapping the time-domain feature representation to a latent space through a first linear layer to obtain a mapped time-domain feature representation; Performing normalization processing on the mapped time-domain feature representation through a normalization layer to obtain a normalized time-domain feature representation; Setting the time-domain feature representation with a negative value to zero through an activation function layer, and keeping the time-domain feature representation with a positive value to obtain a set time-domain feature representation; Performing a linear transformation on the set time-domain feature representation through a second linear layer to obtain an original time-domain latent space representation and an enhanced time-domain latent space representation; Inputting the frequency-domain feature representation into a mapper, and mapping the frequency-domain feature representation to a latent space through a first linear layer to obtain a mapped frequency-domain feature representation; Performing normalization processing on the mapped frequency-domain feature representation through a normalization layer to obtain a normalized frequency-domain feature representation; Setting the frequency-domain feature representation with a negative value to zero through an activation function layer, and keeping the frequency-domain feature representation with a positive value to obtain a set frequency-domain feature representation; Performing a linear transformation on the set frequency-domain feature representation through a second linear layer to obtain an original frequency-domain latent space representation and an enhanced frequency-domain latent space representation.
6. The method according to claim 3, characterized in that, The obtaining a weighted loss function for contrast learning based on the time-domain feature representation and the frequency-domain feature representation includes: Obtaining a time-domain loss function of the time-domain feature based on the similarity between the high-dimensional time-domain feature representation and the enhanced time-domain feature representation in the time-domain feature representation; Obtaining a frequency-domain loss function of the frequency-domain feature based on the similarity between the high-dimensional frequency-domain feature representation and the enhanced frequency-domain feature representation in the frequency-domain feature representation; Obtaining a cross-domain contrast loss function based on the similarity between the time-domain latent space representation and the frequency-domain latent space representation; Obtaining a weighted loss function according to the time-domain loss function, the frequency-domain loss function, and the cross-domain contrast loss function.
7. The method according to claim 1, characterized in that After training a neural network model according to the time-domain samples and the frequency-domain samples by a contrast learning algorithm to obtain a network time series anomaly detection model, it further includes: Obtaining a time series to be detected, and inputting the time series to be detected into the network time series anomaly detection model; Using the network time series anomaly detection model to map the multi-dimensional data in the time series to be detected to a latent space, and determining a loss value of the time series to be detected in the latent space; Judging whether the loss value is greater than a preset loss threshold, where the preset loss threshold is determined according to hyperparameters of the network time series anomaly detection model; In response to determining that the loss value is greater than the preset loss threshold, determining that the time series to be detected is an abnormal time series.
8. An apparatus for generating an abnormal detection model for network time series, characterized in that, Includes: An extraction module, configured to obtain a network timing sequence, extract multi-dimensional data from the network timing sequence, and perform padding processing and conversion processing on the multi-dimensional data to obtain standard data; A conversion module, configured to convert the standard data from the time domain to the frequency domain through Fourier transform to obtain time domain features and frequency domain features; An enhancement module, configured to perform enhancement processing on the time domain features to obtain time domain samples, and perform enhancement processing on the frequency domain features to obtain frequency domain samples; A training module, configured to train an initial anomaly detection model according to the time domain samples and the frequency domain samples through a contrastive learning algorithm to obtain a network timing sequence anomaly detection model.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method described in any one of claims 1 to 7.
Citation Information
Cited By
Microservice system anomaly detection method, device and equipment based on time-frequency feature contrast enhancement, and storage medium
CN121144144A