A Network Traffic Anomaly Detection Method Based on Transformer Generative Adversarial Network

By constructing a network traffic anomaly detection method based on Transformer generated adversarial network, the problem of low detection accuracy of new threats and unknown attacks in the prior art is solved, efficient and accurate network traffic anomaly detection is achieved, and false alarm rate and data collection costs are reduced.

CN119496638BActive Publication Date: 2025-07-29UNIV OF ELECTRONICS SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411540737.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-07-29
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

When facing new threats and unknown attacks, the existing network intrusion detection system has low detection accuracy, is sensitive to dynamic changes in the network environment, has a high false alarm rate, making it difficult to effectively identify complex and changeable attack methods.

Method used

The method of generating an adversarial network based on Transformer is adopted to convert network traffic data into traffic feature matrix, and an encoder, generator and discriminator are built. Normal traffic features are learned through adversarial training, and timing features are processed using Transformer's self-attention mechanism to reduce dependence on labeled attack data, and an unsupervised detection model is built.

Benefits of technology

It improves the detection ability of unknown attacks and variant attacks, reduces the false positive rate, enhances the response ability to new threats, improves the accuracy and efficiency of detection, and reduces the complexity and cost of data collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119496638B_ABST
    Figure CN119496638B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting network traffic anomalies based on a Transformer generative adversarial network. First, a number of traffic data are collected in the normal operating state of the communication network, and traffic feature matrices corresponding to each traffic data are generated to form a training sample set. A Transformer generative adversarial network including an encoder, a generator, and a discriminator is constructed and trained using the training sample set. For the network traffic to be detected, a traffic feature matrix is generated in the same way, and this traffic feature matrix is input into the trained Transformer generative adversarial network. The difference value between the traffic feature matrix of the network traffic to be detected and the generated traffic feature matrix is calculated to determine whether the network traffic to be detected is abnormal. The present invention converts network traffic data into traffic feature matrices, and uses the Transformer generative adversarial network to learn the characteristics of normal traffic data, thereby achieving efficient anomaly detection and improving the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security artificial intelligence detection, and more specifically, relates to a network traffic anomaly detection method based on a Transformer generative adversarial network. Background Art

[0002] Currently, network intrusion detection systems (IDSs) have been widely used in the field of network security. Traditional IDS technologies are mainly divided into the following categories:

[0003] Signature-based Detection: Such IDSs rely on a pre-defined attack signature database and detect potential threats by matching the data in network traffic with known signatures. This method is very effective in detecting known attacks, but is powerless against new or variant attacks (such as zero-day attacks). In addition, the maintenance of the signature database requires continuous updating, and as the number of signatures increases, the consumption of system resources also increases. Moreover, signature-based detection methods cannot identify unknown attacks, especially zero-day attacks, because such attacks do not have pre-defined signatures. Therefore, in the face of new threats, traditional signature matching methods appear powerless.

[0004] Anomaly-based Detection: This method detects abnormal behaviors that deviate from the baseline of normal network behavior by establishing a baseline of normal network behavior. Although this method can discover unknown attack types, its false positive rate is usually high because the dynamic changes in the network environment may cause normal behaviors to be misjudged as abnormal, reducing the practical application value of the system.

[0005] Statistical and Rule-based Detection: These methods rely on statistical models or pre-defined rules to identify potential threats. Although they perform well in certain specific environments, it is difficult to maintain a high detection rate in the face of complex and variable attack means. Especially in the context of the continuous evolution of network attack means, the adaptability and scalability of such methods are limited.

[0006] Machine Learning-based Detection: In recent years, machine learning algorithms have been introduced into IDSs to improve the automation and intelligence level of detection. By classifying network traffic data, these systems can identify known and unknown attacks. However, the performance of machine learning algorithms depends severely on the quality and quantity of the training data set. Especially when dealing with imbalanced data sets (i.e., normal traffic is much more than attack traffic), it may lead to a decline in the detection effect.

[0007] With its powerful generation ability, Generative Adversarial Networks (GANs) have also been applied to network intrusion detection systems to some extent. Such technical solutions train a generator and a discriminator. The generator attempts to generate forged network traffic data, while the discriminator attempts to distinguish between real network traffic and the forged traffic generated by the generator. Through such adversarial training, the system can learn the characteristics of network traffic and thus be used to detect potential network attacks. Summary of the Invention

[0008] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for detecting network traffic anomalies based on a Transformer Generative Adversarial Network. The network traffic data is converted into a traffic feature matrix, and the Transformer Generative Adversarial Network is used to learn the characteristics of normal traffic data, thereby achieving efficient anomaly detection and improving the detection accuracy.

[0009] To achieve the above-mentioned invention purpose, the method for detecting network traffic anomalies based on a Transformer Generative Adversarial Network of the present invention includes the following steps:

[0010] S1: Collect a number of traffic data to obtain a traffic data set under the normal operating state of the communication network; extract the feature values of each traffic data respectively to obtain a traffic feature vector , where represents the th feature value, ; copy the traffic feature vector times to obtain an extended vector with a length of , and then convert it into a traffic feature matrix , where and represent the height and width of the matrix respectively, represents the number of channels; use the traffic feature matrix corresponding to each traffic data as a training sample, thereby constituting a training sample set;

[0011] S2: Construct a Transformer Generative Adversarial Network, including an encoder, a generator, and a discriminator, where:

[0012] The encoder is used to encode the traffic feature matrix to obtain a noise matrix and send it to the generator;

[0013] The generator is used to generate a reconstructed traffic feature matrix according to the noise matrix ​and send it to the discriminator; the generator includes a noise matrix block module, a position encoding module, a vector superposition module, a Transformer encoder module, and a generator output module, where:

[0014] The noise matrix block module is used to divide the noise matrix into multiple non-overlapping noise data blocks of a fixed size, flatten each noise data block into a vector, and then embed the noise vector into a vector space of a fixed dimension through linear projection to obtain a noise embedding vector , , denotes the number of data blocks, and send noise embedding vectors to the vector superposition module;

[0015] The position encoding module is used to generate a position encoding vector of the same dimension as the noise embedding vector for each noise data block and send it to the vector superposition module;

[0016] The vector superposition module is used to superpose the position encoding vector onto the noise embedding vector respectively to obtain a feature vector and send it to the Transformer encoder module;

[0017] The Transformer encoder module includes Transformer encoders, which are used to encode feature vectors to obtain encoded features and send them to the generator output module;

[0018] The generator output module is used to process the encoded features to generate a reconstructed traffic feature matrix ;

[0019] The discriminator is used to discriminate between the real traffic feature matrix or the reconstructed traffic feature matrix to obtain a discrimination score belonging to the real data;

[0020] S3: Use the training sample set obtained in step S1 to train the Transformer generative adversarial network constructed in step S2 to obtain a trained Transformer generative adversarial network. The training method is as follows:

[0021] First, alternately train the generator and the discriminator using the generative adversarial network training method, and then fix the parameters of the generator and the discriminator and train the encoder;

[0022] S4: For the network traffic to be detected, generate a traffic feature matrix using the same method as in step S1 , and then input it into the Transformer generative adversarial network trained in step S3 to obtain a generated two-dimensional traffic feature matrix , and then calculate the two-dimensional traffic feature matrix and the two-dimensional traffic feature matrix to obtain the difference value between them . If the difference value is greater than the preset difference value threshold , then the network traffic is abnormal network traffic, otherwise it is normal network traffic.

[0023] The network traffic anomaly detection method based on the Transformer generative adversarial network of the present invention first collects a number of traffic data in the normal operating state of the communication network, generates a traffic feature matrix corresponding to each traffic data to form a training sample set, constructs a Transformer generative adversarial network including an encoder, a generator and a discriminator, and trains it using the training sample set; for the network traffic to be detected, generate a traffic feature matrix using the same method, and input the traffic feature matrix into the trained Transformer generative adversarial network, and calculate the difference value between the traffic feature matrix of the network traffic to be detected and the generated traffic feature matrix to determine whether the network traffic to be detected is abnormal.

[0024] The present invention has the following beneficial effects:

[0025] 1) The present invention constructs a Transformer generative adversarial network, and uses the self-attention mechanism in the Transformer to effectively process the temporal features in the network traffic. Especially in the analysis of long-sequence data, compared with the traditional RNN or LSTM structure, it has higher accuracy and efficiency, thus effectively improving the performance when dealing with complex network behaviors and improving the accuracy of anomaly detection;

[0026] 2) The present invention combines the generation ability of GAN and the feature extraction ability of Transformer, and proposes an improved solution for the data imbalance problem by only using positive sample data for modeling, that is, only relying on normal traffic data for training, reducing the dependence on a large amount of labeled attack data, thus reducing the complexity and cost of data collection;

[0027] 3) The present invention uses Transformer to construct a generative adversarial network to learn the network traffic features, which can significantly reduce false alarms, improve the accuracy of detection, ensure that normal network activities will not be mislabeled as abnormal or malicious behaviors, thus reducing the interference to network administrators;

[0028] 4) The present invention uses an unsupervised Transformer generative adversarial network, which can learn complex traffic patterns through adversarial training, so as to more effectively detect unknown attacks or variant attacks, and improve the response ability of the method of the present invention to new threats. Description of the Drawings

[0029] Figure 1 is a flowchart of the specific implementation manner of the network traffic anomaly detection method based on the Transformer generative adversarial network of the present invention;

[0030] Figure 2 is a structural diagram of the traffic anomaly detection model of the Transformer generative adversarial network in the present invention;

[0031] Figure 3 is a structural diagram of the encoder in this embodiment;

[0032] Figure 4 is a structural diagram of the generator in this embodiment;

[0033] Figure 5 is a structural diagram of the discriminator in this embodiment;

[0034] Figure 6 is a schematic diagram of the calculation of the difference value in this embodiment. Detailed Description of the Invention

[0035] The following describes the specific implementation manner of the present invention with reference to the drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.

[0036] Embodiment

[0037] Figure 1 is a flowchart of the specific implementation manner of the network traffic anomaly detection method based on the Transformer generative adversarial network of the present invention. As Figure 1 shown, the specific steps of the network traffic anomaly detection method based on the Transformer generative adversarial network of the present invention include:

[0038] S101: Obtain training samples:

[0039] Under the normal operating state of the communication network, a number of traffic data are collected to obtain a traffic data set. The feature values of each traffic data are extracted respectively to obtain a traffic feature vector , where represents the th feature value, . The traffic feature vector Copy times to obtain an extended vector of length , and then convert it into a traffic feature matrix , where and represent the height and width of the matrix respectively, represents the number of channels. Take the traffic feature matrix corresponding to each traffic data as a training sample, thus forming a training sample set.

[0040] In this embodiment, the method for extracting the traffic feature vector is: extract the source IP address, destination IP address, port number, protocol type, TCP flag bits, statistical features of the traffic, and time features of the traffic from the traffic data, where the statistical features of the traffic include the average size of the packets and the duration of the flow, and the time features include the time interval between packet arrivals and the active time of the flow.

[0041] In this embodiment, in order to improve the quality of the training sample set, after collecting the traffic data, it is also necessary to perform cleaning processing on the traffic data. Data cleaning includes operations such as removing noise data, handling missing values and outliers, etc., to ensure the accuracy and consistency of the data. Remove irrelevant or harmful data by detecting and deleting duplicate samples or abnormally long traffic sessions. Then perform feature selection, aiming to select the features that are most discriminative for the detection task. Through the above data cleaning steps, it can be ensured that the constructed intrusion detection system has higher detection accuracy and efficiency in actual applications.

[0042] In addition, in order to ensure that the features are compared on the same scale, it is also necessary to perform normalization processing on each feature value in the dataset. In this embodiment, the "maximum-minimum" normalization method is adopted, and each is scaled to a specified range, which is [0, 1]. That is, for the traffic feature vector each feature value is the normalized feature value, and the normalization calculation formula is:

[0043]

[0044] where, the original numerical value of the th feature, , represent the maximum and minimum values of this feature item in all traffic data.

[0045] S102: Construct a Transformer generative adversarial network:

[0046] In order to improve the accuracy of traffic anomaly detection, the present invention constructs a traffic anomaly detection model of a Transformer generative adversarial network. Figure 2It is the structural diagram of the traffic anomaly detection model of the Transformer generative adversarial network in the present invention. As Figure 2 shown, the traffic anomaly detection model of the Transformer generative adversarial network in the present invention includes an encoder, a generator, and a discriminator. Next, each module will be described in detail.

[0047] The encoder is used to encode the traffic feature matrix to obtain a noise matrix and send it to the generator. Figure 3 It is the structural diagram of the encoder in this embodiment. As Figure 3 shown, the encoder in the present invention includes a first linear module, a second linear module, and a third linear module, where:

[0048] The first linear module includes a cascaded linear transformation layer (Linear), a LeakyReLU activation layer, and a Dropout module, and is used to perform linear processing on the traffic feature matrix to obtain a matrix and send it to the second linear module.

[0049] The second linear module includes a cascaded linear transformation layer, a LeakyReLU activation layer, and a Dropout module, and is used to perform linear processing on the matrix to obtain a matrix and send it to the third linear module.

[0050] The third linear module includes a cascaded linear transformation layer and a Tanh activation layer, and is used to perform linear processing on the matrix to obtain a noise matrix and output it.

[0051] According to the above description, it can be known that the operation of the encoder in this embodiment can be expressed by the following formula:

[0052]

[0053] where , , respectively represent the linear transformation coefficients of the first linear module, the second linear module, and the third linear module, , , respectively represent the biases of the first linear module, the second linear module, and the third linear module, , respectively represent the LeakyRelu activation functions in the first linear module and the second linear module, , respectively represent the Dropout module operations in the first linear module and the second linear module, represent the Tanh activation function in the first linear module.

[0054] The generator is used to generate a reconstructed traffic feature matrix based on the noise matrix generated by the encoder and send it to the discriminator. Figure 4 is the structural diagram of the generator in this embodiment. As Figure 2 shown, the generator in this embodiment uses a Transformer as the backbone structure, including a noise matrix block division module, a position encoding module, a vector superposition module, a Transformer encoder module, and a generator output module, where:

[0055] The noise matrix block division module is used to divide the noise matrix into multiple non-overlapping noise data blocks of a fixed size, flatten each noise data block into a vector, and then embed the noise vector into a vector space of a fixed dimension through linear projection to obtain a noise embedding vector , , represents the number of data blocks, and send noise embedding vectors to the vector superposition module. By dividing the noise matrix into blocks, high-dimensional data can be effectively processed, local spatial information can be retained, and it can be converted into a format suitable for Transformer processing. Assume that a two-dimensional noise matrix is divided into noise data blocks of size , and the number of data blocks . Then each noise data block is unfolded to obtain a vector of length , and then linearly projected into a vector space of a fixed dimension. This process converts the matrix blocks into a set of vector representations, called Patch Embeddings.

[0056] The position encoding module is used to generate a position encoding vector of the same dimension as the noise embedding vector for each noise data block and send it to the vector superposition module. Since the Transformer structure itself does not have position information, in order to retain the position information of the data blocks in the noise feature matrix, the position encoding is added and superimposed on the noise embedding vector, so as to retain the position information of the noise data blocks and enhance the model's understanding ability of the spatial structure of the noise data.

[0057] The vector superposition module is used to superimpose the position encoding vector on the noise embedding vector respectively to obtain a feature vector and send it to the Transformer encoder module.

[0058] The Transformer encoder module includes a number of Transformer encoders, which are used to a number of feature vectors for encoding to obtain encoded features and send them to the generator output module. As Figure 4 shown, in this embodiment, the Transformer encoder includes a multi-head self-attention (Multi Head Attention) module, a first feature fusion module, a first layer normalization (Layer Norm) module, a feed-forward network (Feed-Forward Net, FFN), a second feature fusion module, and a second layer normalization module, where:

[0059] The multi-head self-attention module is used to process the input feature matrix using the multi-head self-attention mechanism, and send the obtained attention matrix to the first feature fusion module. The multi-head self-attention mechanism embeds each matrix block to interact with other blocks, and captures the global relationships in the matrix by calculating self-attention. The self-attention mechanism maps the input features to three matrices of query (Q), key (K), and value (V) respectively, calculates the dot product between them to generate attention scores. This process is executed in parallel on multiple independent heads, and the outputs of each head will be concatenated and projected back to the input dimension through a linear layer. The input undergoes three linear transformations to generate the query matrix Q, the key matrix K, and the value matrix V respectively. The core of the attention mechanism is to calculate the similarity measure between the query and the key as the attention score to construct the attention matrix. The specific formula is:

[0060]

[0061] where, is the dimension of the key vector, is the scaling factor used to stabilize the gradient.

[0062] The first feature fusion module is used to superimpose the input feature matrix and the attention matrix to obtain the fused feature and send it to the first layer normalization module.

[0063] The first layer normalization module is used to perform normalization processing on the fused feature to obtain the normalized feature and send it to the feed-forward network. The specific operation of the layer normalization module can be expressed by the following formula:

[0064]

[0065] Among them, represents the input, represents the operation of calculating the mean value, represents the operation of calculating the standard deviation, represents the scaling parameter, represents the offset parameter.

[0066] The feed-forward network is used to map the received normalized features to obtain features and send them to the second feature fusion module. The feed-forward network is independently applied to the embedding of each matrix block, that is, the calculations of each matrix block are parallel. In this embodiment, the feed-forward network includes two fully connected layers and one layer of Dropout neuron random inactivation layer, and there is a non-linear activation function between them. The specific operation of the feed-forward network can be expressed by the following formula:

[0067]

[0068] Among them, and respectively represent the operations of two fully connected layers, represents the Relu activation function layer, represents the Dropout operation.

[0069] The second feature fusion module is used to superimpose the normalized features and the features to obtain the fused features and send them to the second layer normalization module.

[0070] The second layer normalization module is used to normalize the fused features to obtain the normalized features and output them as the result of the current Transformer encoder.

[0071] The generator output module is used to process the encoded features to generate the reconstructed traffic feature matrix . In this embodiment, the generator output module adopts a cascaded layer normalization layer and a fully connected layer, that is, first perform layer normalization processing on the encoded features , and then use a fully connected layer (Linear) to inversely synthesize the output result into a reconstructed two-dimensional traffic feature matrix with the same dimension as the original two-dimensional traffic feature matrix.

[0072] The discriminator is used to discriminate the real traffic feature matrix or the reconstructed two-dimensional traffic feature matrix to obtain the discrimination score belonging to the real data.Figure 5 is the structural diagram of the discriminator in this embodiment. As Figure 5 shown, in this embodiment, the discriminator also uses the Transformer as the backbone structure, including a traffic feature matrix block division module, a first position encoding module, a first vector superposition module, a Transformer encoder module, an encoded feature embedding, a second position encoding module, a second vector superposition module, a Transformer decoder module, a splicing module, and a discriminator output module, where:

[0073] The traffic feature matrix block division module is used to divide the real traffic feature matrix or the reconstructed two-dimensional traffic feature matrix into multiple non-overlapping traffic feature data blocks of a fixed size, flatten each traffic feature data block into a vector, and then embed the traffic feature vector into a vector space of a fixed dimension through linear projection to obtain a traffic embedding vector , and send traffic embedding vectors to the first vector superposition module. In this embodiment, the block division process of the traffic feature matrix block division module is the same as that of the noise matrix block division module.

[0074] The first position encoding module is used to generate a position encoding vector of the same dimension as the traffic embedding vector for each traffic feature data block and send it to the first vector superposition module;

[0075] The first vector superposition module is used to respectively superimpose the position encoding vector onto the traffic embedding vector to obtain a feature vector and send it to the Transformer encoder module.

[0076] The Transformer encoder module includes Transformer encoders, which are used to encode feature vectors to obtain encoded features and send them to the Transformer decoder module. In this embodiment, the Transformer encoder module in the discriminator uses the same number and structure of Transformer encoders as the Transformer encoder module in the generator.

[0077] The encoded feature embedding module performs a linear projection on encoded features to obtain encoded embedding vectors and send them to the second vector superposition module.

[0078] The second position encoding module is used to generate a position encoding vector of the same dimension for each encoded embedding vector and send it to the second vector superposition module;

[0079] The second vector superposition module is used to separately superpose the position encoding vectors onto the traffic embedding vectors to obtain feature vectors and send them to the Transformer decoder module.

[0080] The Transformer decoder module includes a number of Transformer decoders, which are used to decode a number of encoded features to obtain reconstructed features and send them to the splicing module. As Figure 5 shown, in this embodiment, the structure of the Transformer decoder module is somewhat similar to that of the Transformer encoder module, including a masked multi-head self-attention module, a first feature fusion module, a first layer normalization module, a multi-head self-attention module, a second feature fusion module, a second layer normalization module, a feed-forward network, a third feature fusion module, and a third layer normalization module, where:

[0081] The masked multi-head self-attention module is used to process the input feature matrix using the masked multi-head self-attention mechanism, and send the obtained attention matrix to the first feature fusion module. The masked multi-head self-attention mechanism is almost the same as the standard multi-head self-attention mechanism, and the only difference lies in its attention mechanism masking process. That is, in the sequence generation task, to prevent the model from "seeing" future information during generation, an upper triangular mask matrix is used to mask the information of future time steps (tokens). Ensure that when the model generates the output of the current time step, it only focuses on past time steps.

[0082] The first feature fusion module is used to superpose the input feature matrix and the attention matrix to obtain a fused feature and send it to the first layer normalization module.

[0083] The first layer normalization module is used to perform normalization processing on the fused feature to obtain a normalized feature and send it to the multi-head self-attention module.

[0084] The multi-head self-attention module is used to use the multi-head self-attention mechanism to process the normalized feature Process it to obtain the attention matrix and send it to the second feature fusion module.

[0085] The second feature fusion module is used to stack the normalized features and the attention matrix to obtain the fused feature and send it to the second layer normalization module.

[0086] The second layer normalization module is used to perform normalization processing on the fused feature to obtain the normalized feature and send it to the feed-forward network.

[0087] The feed-forward network is used to perform a mapping operation on the received normalized feature to obtain the feature and send it to the third feature fusion module.

[0088] The third feature fusion module is used to stack the normalized feature and the feature to obtain the fused feature and send it to the third layer normalization module.

[0089] The third layer normalization module is used to perform normalization processing on the fused feature to obtain the normalized feature and output it as the result of the current Transformer decoder.

[0090] The splicing module is used to splice reconstruction features to obtain the feature and send it to the discriminator output module.

[0091] The discriminator output module is used to process the feature to obtain the score that the input traffic feature vector is the real traffic . In this embodiment, the discriminator output processing module includes a cascaded fully connected layer and a Softmax layer, so the score can be expressed by the following formula:

[0092]

[0093] S103: Train the Transformer generative adversarial network:

[0094] Use the training sample set obtained in step S1 to train the Transformer generative adversarial network constructed in step S2 to obtain the trained Transformer generative adversarial network. The training method is:

[0095] First, the generator and discriminator are alternately trained using the generative adversarial network training method. Then, with the parameters of the generator and discriminator fixed, the encoder is trained.

[0096] To improve the training effect of the Transformer generative adversarial network, in this embodiment, the calculation method of the loss function is improved, and the training stability of the network is enhanced through the gradient penalty mechanism. In the training process of the traditional generative adversarial network, problems such as vanishing gradients or exploding gradients may occur, resulting in difficulties in model convergence or unsatisfactory generation results. To solve this problem, in this embodiment, new samples are generated by interpolating between real samples and generated samples, and the gradients of the discriminator are calculated for these interpolated samples, thereby introducing a gradient penalty term to optimize the calculation of the loss function. The discriminator loss function in this embodiment The specific calculation method is as follows:

[0097] For each real traffic feature matrix in the current training batch (batch) and the generated traffic feature matrix generated by the Transformer generative adversarial network , , denotes the number of samples in the training batch, and a weight matrix is randomly generated, where the value range of each element is [0, 1]. Through this weight, linear interpolation is performed between the real sample and the generated sample to obtain the interpolated traffic feature matrix :

[0098]

[0099] Each interpolated traffic feature matrix is input into the discriminator to obtain the corresponding score and its corresponding gradient is calculated, thereby obtaining the gradient vector , and then the gradient penalty term is calculated using the following formula:

[0100]

[0101] where denotes taking the L2 norm.

[0102] The design goal of the gradient penalty term gradient penalty (gp) is to make the gradient norm of the discriminator close to 1, thereby ensuring the smoothness of the gradient.

[0103] The discriminator loss is calculated using the following formula :

[0104]

[0105] Among them, and respectively represent the real traffic feature matrix and the generated traffic feature matrix of the discriminant score. is a hyperparameter used to adjust the weight of the gradient penalty term. In this embodiment, is 10.

[0106] The generator loss The calculation formula is:

[0107]

[0108] In this embodiment, when the encoder is trained, the loss function The calculation method is:

[0109]

[0110] Among them, represents the difference loss between the real traffic feature matrix and the generated traffic feature matrix . In this embodiment, , represents calculating the mean square error. represents the real traffic feature matrix discriminant score and the difference loss between the discriminant scores of the generated traffic feature matrix . In this embodiment, . represents a hyperparameter used to adjust the difference loss of the discriminant score. In this embodiment, is set to 1.

[0111] S104: Network traffic detection:

[0112] For the network traffic to be detected, use the same method in step S101 to generate the traffic feature matrix , and then input it into the Transformer generative adversarial network trained in step S103 to obtain the generated two-dimensional traffic feature matrix , and then calculate the difference value between the two-dimensional traffic feature matrix and the two-dimensional traffic feature matrix . If the difference value is greater than the preset difference value threshold , then the network traffic is abnormal network traffic, otherwise it is normal network traffic.

[0113] Figure 6 is the schematic diagram of the calculation of the difference value in this embodiment. As Figure 6As shown, in this embodiment, the encoder loss function is used as the difference value , that is, the difference value is calculated by the following method:

[0114]

[0115] Among them, represents the difference between the true traffic feature matrix and the generated traffic feature matrix , represents calculating the mean square error, represents the difference between the discrimination score of the true traffic feature matrix and the discrimination score of the generated traffic feature matrix . represents the hyperparameter used to adjust the discrimination score difference.

[0116] In this embodiment, the difference value threshold is determined by the following method:

[0117] In addition, a number of traffic data when the network is in an attack state are collected. The attack state can be set according to actual needs. In this embodiment, the attack states include DoS / DDoS attacks, port scanning attacks, and Web attacks. The traffic feature matrix of each traffic data in the attack state is generated by the same method as in step S101. The traffic feature matrix in the normal running state and the traffic feature matrix in the attack state are respectively input into the Transformer generative adversarial network, and the corresponding difference value is calculated, and the threshold value with the highest judgment accuracy for the two types of traffic feature matrices is searched as the finally used difference value threshold .

[0118] To illustrate the technical effects of the present invention, a specific example is used to experimentally verify the present invention. In this embodiment, the CIC-IDS-2017 dataset is used. This dataset is released by the Canadian Institute of Cybersecurity and simulates a real enterprise network environment, including multiple clients and servers. Traffic collection covers multiple protocols and services, making it close to the real scenario. The dataset contains a variety of known network attack types, including but not limited to: brute force cracking of FTP and SSH, DoS and DDoS, Web attacks, and Heartbleed attacks, etc.

[0119] First, preprocess the dataset. In this embodiment, after data cleaning and feature selection, 78 most important feature items are finally retained and normalized. The normalized data is copied 5 times to form a one-dimensional feature vector with a length of 390. Subsequently, the one-dimensional feature vector is padded with 0s to a length of 400 and converted into a two-dimensional feature matrix of 20*20. Finally, the preprocessed data is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1, and the training set is used to train the transformer-based generative adversarial network.

[0120] During training, only normal data is used for training. During validation and testing, both normal and abnormal data are used for testing. The validation set is used to validate the model, and the score of each piece of data is recorded. Finally, the F1-score is used as a reference to search for the optimal scoring threshold, and this threshold is used as the boundary between normal data and abnormal data. Data higher than this threshold is normal data, and data lower than this threshold is abnormal.

[0121] In this embodiment, three existing network intrusion detection methods are used as comparison methods to compare and validate with the present invention, which are respectively:

[0122] MSCNN-LSTM-AE: For details, see the literature "Abou El Houda Z, Senhaji Hafid A, Khoukhi L. A novel unsupervised learning method for intrusion detection in software-defined networks[M] / / Computational Intelligence in Recent Communication Networks. Cham: Springer International Publishing, 2021: 103-117."

[0123] FID-GAN: For details, see the literature "Singh A, Jang-Jaccard J. Autoencoder-based unsupervised intrusion detection using multi-scale convolutional recurrent networks[J]. arXiv preprint arXiv:2204.03779, 2022."

[0124] IDS-IF: For details, see the literature "de Araujo-Filho P F, Naili M, Kaddoum G, et al. Unsupervised gan-based intrusion detection system using temporal convolutional networks and self-attention[J]. IEEE Transactions on Network and Service Management, 2023, 20(4): 4951-4963."

[0125] In the performance evaluation of network intrusion detection methods, common evaluation metrics include accuracy, recall, precision, F1-score, and area under the curve (AUC). These metrics can comprehensively reflect the effectiveness of the model in detecting network attacks. In practical applications, recall is mainly used as the standard, which can truly reflect whether there is any missed detection of abnormal data. Table 1 is a comparison table of the evaluation metrics of the method of the present invention and three other methods in this embodiment.

[0126] Method Accuracy Precision Recall F1-Score MSCNN-LSTM-AE 0.92 0.88 0.83 0.85 FID-GAN 0.916 0.899 0.899 0.899 IDS-IF 0.92 0.92 0.92 0.92 The present invention 0.916 0.912 0.995 0.952

[0127] Table 1

[0128] As shown in Table 1, the method of the present invention performs excellently in terms of recall, reaching as high as 0.995, which is significantly better than other methods. This means that its ability to detect attacks is very strong and it can effectively identify almost all attacks. However, the precision is 0.912, slightly lower than that of the IDS-IF method, indicating that there is a certain risk of false alarms. Generally, the F1-score is 0.952, showing that this method can maintain a good balance while accurately detecting attacks, and it is a very suitable detection method for scenarios with extremely high security requirements.

[0129] Although the above-described illustrative specific embodiments of the present invention are described to facilitate the understanding of the present invention by those skilled in the art, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

Claims

1. A network traffic anomaly detection method based on a Transformer generative adversarial network, characterized in that It includes the following steps: S1: Collect a number of traffic data to obtain a traffic data set under the normal operation state of the communication network; extract the eigenvalues of each traffic data respectively to obtain a traffic feature vector , where represents the th eigenvalue, ; copy the traffic feature vector times to obtain an extended vector with a length of , and then convert it into a traffic feature matrix , where and represent the height and width of the matrix respectively, represents the number of channels; regard the traffic feature matrix corresponding to each traffic data as a training sample, thus constituting a training sample set;​ S2: Construct a Transformer generative adversarial network, including an encoder, a generator, and a discriminator, where: The encoder is used to encode the traffic feature matrix to obtain the noise matrix and send it to the generator; The generator is used to generate a reconstructed traffic feature matrix according to the noise matrix generated by the encoder and send it to the discriminator; the generator includes a noise matrix block module, a positional encoding module, a vector superposition module, a Transformer encoder module, and a generator output module, where: ​ The noise matrix block module is used to divide the noise matrix into multiple non-overlapping noise data blocks of a fixed size, flatten each noise data block into a vector, and then embed the noise vector into a vector space of a fixed dimension through linear projection to obtain a noise embedding vector , , denotes the number of data blocks, and send noise embedding vectors to the vector superposition module; The position encoding module is used to generate a position encoding vector with the same dimension as the noise embedding vector for each noise data block and send it to the vector superposition module; The vector superposition module is used to separately superpose the position encoding vector onto the noise embedding vector to obtain the feature vector and send it to the Transformer encoder module; The Transformer encoder module includes Transformer encoders for feature vectors to be encoded to obtain encoded features and sent to the generator output module; The generator output module is used to process the encoded features to generate a reconstructed traffic feature matrix ; The discriminator is used to discriminate the real traffic feature matrix or the reconstructed traffic feature matrix and obtain the discrimination score belonging to the real data; S3: Use the training sample set obtained in step S1 to train the Transformer generative adversarial network constructed in step S2 to obtain a trained Transformer generative adversarial network. The training method is as follows: First, alternately train the generator and the discriminator using the generative adversarial network training method, and then fix the parameters of the generator and the discriminator and train the encoder; S4: For the network traffic to be detected, generate a traffic feature matrix using the same method as in step S1 , and then input it into the Transformer generative adversarial network trained in step S3 to obtain a generated two-dimensional traffic feature matrix , and then calculate the two-dimensional traffic feature matrix and the two-dimensional traffic feature matrix to obtain the difference value . If the difference value is greater than the preset difference value threshold , then the network traffic is abnormal network traffic, otherwise it is normal network traffic.

2. The network traffic anomaly detection method according to claim 1, wherein The features of the traffic data in step S1 include the source IP address, destination IP address, port number, protocol type, TCP flag bits, statistical features of the traffic, and time features. The statistical features of the traffic include the average size of the packets and the duration of the flow, and the time features include the time interval between packet arrivals and the active time of the flow.

3. The network traffic anomaly detection method according to claim 1, wherein The encoder in step S2 includes a first linear module, a second linear module, and a third linear module, where: The first linear module includes cascaded linear transformation layers, LeakyReLU activation layers, and Dropout modules, and is used to perform linear processing on the traffic feature matrix to obtain a matrix and send it to the second linear module; The second linear module includes cascaded linear transformation layers, LeakyReLU activation layers, and Dropout modules for linearly processing the matrix to obtain the matrix and sending it to the third linear module; The third linear module includes a cascaded linear transformation layer and a Tanh activation layer, which are used to perform linear processing on the matrix to obtain a noise matrix and output it.

4. The network traffic anomaly detection method according to claim 1, characterized in that, The Transformer encoder in step S2 includes a multi-head self-attention module, a first feature fusion module, a first layer normalization module, a feed-forward network, a second feature fusion module, and a second layer normalization module, where: The multi-head self-attention module is used to process the input feature matrix by adopting the multi-head self-attention mechanism and send the obtained attention matrix to the first feature fusion module; The first feature fusion module is used to superimpose the input feature matrix and the attention matrix to obtain the fused feature and send it to the first layer normalization module; The first - layer normalization module is used to normalize the fused features to obtain normalized features and send them to the feed - forward network; The feedforward network is used to process the received normalized features to perform a mapping operation to obtain features and send them to the second feature fusion module; The second feature fusion module is used to superimpose the normalized feature and the feature to obtain a fused feature and send it to the second layer normalization module; The second-layer normalization module is used to normalize the fused features to obtain normalized features and output them as the result of the current Transformer encoder.

5. The network traffic anomaly detection method according to claim 1, wherein The discriminator in step S2 includes a traffic feature matrix block module, a first position encoding module, a first vector superposition module, a Transformer encoder module, an encoded feature embedding, a second position encoding module, a second vector superposition module, a Transformer decoder module, a splicing module, and a discriminator output module, where: The traffic feature matrix block module is used to divide the real traffic feature matrix or the reconstructed two-dimensional traffic feature matrix into multiple non-overlapping traffic feature data blocks of a fixed size, flatten each traffic feature data block into a vector, and then embed the traffic feature vector into a vector space of a fixed dimension through linear projection to obtain a traffic embedding vector , and send traffic embedding vectors to the first vector superposition module; The first position encoding module is used to generate a position encoding vector with the same dimension as the traffic embedding vector for each traffic feature data block and send it to the first vector superposition module; The first vector superposition module is used to superpose the positional encoding vectors onto the traffic embedding vectors respectively to obtain feature vectors and send them to the Transformer encoder module; The Transformer encoder module includes Transformer encoders, which are used to feature vectors for encoding to obtain encoded features and send them to the Transformer decoder module; The encoding feature embedding module performs linear projections on encoding features to obtain encoding embedding vectors and sends them to the second vector superposition module; The second position encoding module is used to generate a position encoding vector of the same dimension for each encoded embedding vector and send it to the second vector superposition module; The second vector superposition module is used to separately superpose the position encoding vector onto the traffic embedding vector to obtain a feature vector and send it to the Transformer decoder module; The Transformer decoder module includes Transformer decoders, which are used to decode the encoded features to obtain the reconstructed features and send them to the splicing module; The splicing module is used to reconstruction features splice to obtain a feature and send it to the discriminator output module; The discrimination output processing module is used to process the feature to obtain the score of the input traffic feature vector as the real traffic .

6. The network traffic anomaly detection method according to claim 5, wherein, The Transformer decoder includes a masked multi-head self-attention module, a first feature fusion module, a first layer normalization module, a multi-head self-attention module, a second feature fusion module, a second layer normalization module, a feed-forward network, a third feature fusion module, and a third layer normalization module, where: The masked multi-head self-attention module is used to process the input feature matrix by means of the masked multi-head self-attention mechanism and send the obtained attention matrix to the first feature fusion module; The first feature fusion module is used to superimpose the input feature matrix and the attention matrix to obtain the fused feature and send it to the first layer normalization module; The first-layer normalization module is used to normalize the fused features to obtain normalized features and send them to the multi-head self-attention module; The multi-head self-attention module is used to process the normalized features by adopting the multi-head self-attention mechanism and send the obtained attention matrix to the second feature fusion module; The first feature fusion module is used to superimpose the normalized features and the attention matrix to obtain the fused feature and send it to the second layer normalization module; The second - layer normalization module is used to normalize the fused features to obtain normalized features and send them to the feed - forward network; The feedforward network is used to process the received normalized features to perform a mapping operation to obtain features and send them to the third feature fusion module; The third feature fusion module is used to superimpose the normalized features and the features to obtain the fused features and send them to the third layer normalization module; The third-layer normalization module is used to normalize the fused features to obtain normalized features, which are output as the result of the current Transformer decoder.

7. The network traffic anomaly detection method according to claim 1, wherein The calculation method of the loss function used in the training of the Transformer generative adversarial network in step S3 is as follows: Discriminator loss function The specific calculation method is as follows: For each real traffic feature matrix in the current training batch and the generated traffic feature matrix generated by the Transformer generative adversarial network , , denotes the number of samples in the training batch, and randomly generate a weight matrix whose elements range from [0, 1]; Perform linear interpolation between the real samples and the generated samples using this weight to obtain an interpolated traffic feature matrix : , Input each interpolated flow feature matrix into the discriminator to obtain the corresponding scores and calculate its corresponding gradients to obtain the gradient vector , and then calculate the gradient penalty term using the following formula : , Among them, represents the calculation of the L2 norm; The discriminator loss is calculated using the following formula : , Among them, , respectively represent the real traffic feature matrix and the generated traffic feature matrix of the discriminant score, is a hyperparameter used to adjust the weight of the gradient penalty term, represents taking the mean; Generator loss The calculation formula is as follows: , Loss function during encoder training The calculation method is as follows: , Among them, represents the real traffic feature matrix and the generated traffic feature matrix of the difference loss, represents the real traffic feature matrix discriminant score and the generated traffic feature matrix of the difference loss between discriminant scores, represents the hyperparameter used to adjust the difference loss of the discriminant score.

8. The network traffic anomaly detection method according to claim 1, characterized in that The difference value in step S4 is calculated by the following method: , Among them, represents the difference between the real traffic feature matrix and the generated traffic feature matrix ; represents the difference between the discriminant score of the real traffic feature matrix and the discriminant score of the generated traffic feature matrix ; represents the hyperparameter used to adjust the difference in discriminant scores.

Citation Information

Patent Citations

  • Two-dimensional image sewage flow detection method based on generative adversarial network

    CN110675374A

  • Network intrusion detection and classification method based on twin Transformer

    CN118740502A