Data enhancement method based on classifier TransGAN and space channel collaborative self-attention

By adopting a data enhancement method based on TransGAN and spatial channel collaborative self-attention in IoT network traffic detection, the data imbalance problem is solved, the detection accuracy and generalization ability of the model are improved, and the detection ability of DDoS attacks is significantly enhanced.

CN120067694APending Publication Date: 2025-05-30JIANGSU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510230397.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The imbalance of network traffic data generated by IoT devices leads to the excellent performance of traditional machine learning models when detecting DDoS attacks, but they are prone to missed or false positives when facing small-scale or rare attack types, and the generalization ability of the model is reduced.

Method used

The data augmentation method based on the classifier TransGAN and spatial channel coordinated self-attention is adopted. Through data cleaning, feature extraction and normalization processing, network traffic data is converted into RGB images, local features are captured using the spatial channel coordinated self-attention mechanism, and auxiliary classifiers and adaptive loss functions are introduced to improve the quality and diversity of generated samples.

Benefits of technology

It significantly improves the accuracy and generalization ability of the model when detecting abnormal traffic of different categories, avoids pattern collapse, and enhances the detection ability of malicious traffic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067694A_ABST
    Figure CN120067694A_ABST
Patent Text Reader

Abstract

The invention provides a data enhancement method based on a classifier TransGAN and space channel collaborative self-attention. Comprising the following steps: step 1, carrying out data cleaning on a network traffic data set, extracting key features of malicious traffic by utilizing mutual information, and converting the network traffic into an RGB image in a CSV form in a value mapping superposition mode; step 2, capturing multi-semantic information of network traffic data in the RGB image by using a space channel collaborative self-attention mechanism, and inputting the multi-semantic information into a TransGAN generator in combination with category information to generate a traffic sample with a label; and step 3, additionally introducing a category output layer to the generator, distinguishing whether an input sample is from a real data set or a labeled sample generated by the generator, and guiding training of the generator and a discriminator by dynamically balancing an adaptive loss function and combining category loss and judgment loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network traffic detection, and relates to a data augmentation method based on the classifier TransGAN and spatial-channel collaborative self-attention. Background Art

[0002] As an emerging technology, the Internet of Things (IoT) is rapidly changing our way of life. The large-scale application of IoT smart devices has also led to an exponential growth trend in the scale of network traffic generated by them. However, IoT devices usually have problems such as limited computing power and weak security protection mechanisms, making them extremely vulnerable to network attacks. According to statistics, in recent years, network attack events against IoT devices such as data leakage, malware propagation, and device hijacking have shown a significant upward trend. These attacks not only threaten users' privacy and data security, but may also lead to the destruction of the stability of the entire IoT ecosystem.

[0003] Among the numerous network attacks against IoT devices, the Denial of Service (DoS) attack and its derivative form - the Distributed Denial of Service (DDoS) attack are particularly common. Such attacks occupy the target device or network resources through a large number of malicious requests to make it unable to provide normal services, thereby achieving the attack purpose. Due to the large number and wide distribution of IoT devices, attackers can easily use multiple infected devices to launch DDoS attacks to form a "botnet", thus causing a huge impact on the target system in a very short time. In order to reduce the risk of IoT being attacked, firewalls, intrusion detection systems, and intrusion prevention systems are often used for IoT devices. They use a static rule library to detect malicious traffic in the network traffic generated by IoT devices to discover attack behaviors, thereby ensuring data transmission security. However, in the face of a complex and changing network environment, the static rule-based method is difficult to defend against complex DDoS attacks. Therefore, it is urgent to introduce artificial intelligence technology to improve the detection accuracy of malicious traffic to enhance the ability to resist network attacks.

[0004] Frequent DDoS attacks on IoT devices have led to a significant increase in the scale of this type of network traffic compared to that generated by other types of attacks, resulting in data imbalance (i.e., a certain type of attack sample occupies a large proportion of all attack samples, while other types of attack samples are relatively few). The emergence of data imbalance not only affects the data distribution during model training but also leads to poor performance of the model in detecting other low-frequency attacks. Specifically, traditional machine learning models tend to overfit attack types with a majority of samples, such as DDoS attacks, when faced with imbalanced data. This tendency makes the model perform excellently in identifying DDoS attacks but prone to false negatives or false positives when dealing with small-scale or rare attack types. In addition, due to the model's inability to fully learn attack patterns with low frequencies in the training data, data imbalance leads to a decline in the generalization ability of the model during actual deployment.

[0005] To address the data imbalance problem, researchers have proposed various solutions. In the field of traditional machine learning, common solutions include undersampling and oversampling. Among them, undersampling balances the dataset by reducing the number of majority-class samples, while oversampling increases the number of minority-class samples by replicating or synthesizing them. These methods are relatively simple in concept but may lead to underfitting or overfitting problems during model training, thus affecting the detection performance. To solve these problems, researchers have proposed improved sampling methods such as Synthetic Minority Over-sampling Technique (SMOTE) and Adaptive Synthetic Sampling Technique, which enhance the representativeness of the minority class by generating synthetic samples. In addition, adjusting the learning algorithm itself is also a common strategy, which can make the model pay more attention to minority-class samples during training by modifying the decision boundary of the classifier or adjusting the loss function, thereby improving the model's recognition ability. Compared with machine learning, deep learning has advantages such as strong feature learning ability and strong adaptability to large amounts of data. Therefore, generative adversarial networks in the field of deep learning have become an effective method for solving data imbalance that has received much attention in recent years. Specifically, generative adversarial networks increase the number of minority-class samples by generating synthetic data similar to minority-class samples, thereby improving the performance of intrusion detection systems when dealing with imbalanced data. In the field of network intrusion, data augmentation methods based on generative adversarial networks have been proven to be superior to traditional undersampling and oversampling techniques.

[0006] When applying generative adversarial networks (GANs) to network traffic data augmentation, the main processing involves three traffic formats: PCAP, CSV, and Gray-scale images. Among them, PCAP files are large, and parsing them is complex and requires a large amount of storage space and computing resources for processing; the representability of the CSV data structure is poor, resulting in the inability to fully learn the characteristics of network traffic when dealing with high-dimensional data with strong complexity and the loss of detailed information in low-dimensional data; gray-scale images are more suitable for the model to detect malicious traffic than the first two data formats, but they will cause the model to lose some details during the feature learning process and lose some hidden features during the conversion process. In addition, due to the wide variety of network attacks against IoT devices, when converting the characteristics shown by various attack methods into images, there are deficiencies such as huge differences between images and many peaks in the data distribution. These deficiencies make it easy for GANs to have mode collapse problems during the learning process, resulting in the generator generating images of specific patterns while ignoring other patterns, thus restricting the diversity of the generated samples.

[0007] Based on this, the present invention proposes a data augmentation method based on a classifier TransGAN and spatial-channel collaborative self-attention. First, data cleaning techniques are used to extract useful features from IoT network traffic, and the correlation between the extracted features and attack category features is evaluated to obtain the non-linear relationship between the features; then, the features are normalized and stacked to form RGB images to elevate the two-dimensional features of the grayscale images to three dimensions, thereby improving the representational ability of network traffic; next, a spatial-channel collaborative self-attention mechanism is introduced into TransGAN to capture sufficient local features of network traffic, thereby improving the quality of sample generation; finally, an auxiliary classifier is introduced into TransGAN to enable both the generator and the discriminator to learn specified features according to input conditions and output specified images according to the conditions, thereby significantly improving the quality of the generated samples and avoiding mode collapse. In addition, an adaptive loss function that can self-feedback according to the quality of sample generation is designed to better guide the direction of sample generation. Summary of the Invention

[0008] The object of the present invention is to more accurately detect different categories of abnormal traffic by solving the problem of network traffic data imbalance. For this purpose, the present invention proposes a data augmentation method based on a classifier TransGAN and spatial-channel collaborative self-attention.

[0009] The present invention provides a data augmentation method based on a classifier TransGAN and spatial-channel collaborative self-attention, including:

[0010] Step 1: Clean the network traffic dataset, extract the key features of malicious traffic using mutual information, and convert the network traffic in CSV format into an RGB image by means of value mapping superposition.

[0011] Step 2: Use the spatial-channel collaborative self-attention mechanism to capture the multi-semantic information of the network traffic data in the RGB image, and combine the category information to input it into the generator of TransGAN to generate labeled traffic samples.

[0012] Step 3: Introduce an additional category output layer to the generator to distinguish whether the input sample comes from the real dataset or the labeled sample generated by the generator. Guide the training of the generator and discriminator by dynamically balancing the adaptive loss function and combining the category loss and the judgment loss.

[0013] In the first aspect, the specific steps of the above Step 1 are as follows:

[0014] Step 1.1: Analyze the network traffic dataset to identify missing values and unreasonable samples. For the missing values in the large-scale normal network traffic, they are directly deleted, while for the missing values in the small-scale malicious traffic, they are processed by mean filling. In addition, for unreasonable values such as nan, -inf, +inf, etc., they are directly deleted to avoid introducing noise.

[0015] Step 1.2: Delete the features that are meaningless for determining malicious traffic generated by attacks and normal traffic generated by normal access behaviors from the network traffic dataset after traffic cleaning, including flow ID, source IP, source port, destination IP, destination port, protocol, timestamp, sub-flow related features, and Flag marks. Calculate the mutual information I between the remaining feature columns and malicious traffic and normal traffic, sort the features in the network traffic according to the value of I, and select the 32 most relevant features.

[0016] Step 1.3: Separate all normal traffic and malicious traffic into two groups of data. Use Min-Max normalization to map each feature column to the range of 0-255 to solve the problem of uneven distribution of traffic feature values. Then, round down and iteratively select 32 rows of traffic to form a sub-table. Each sub-table forms a 32×32×1-dimensional vector. Then, every three selected sub-tables form the R, G, and B channels of the image. Finally, use the OpenCV library to combine the three sub-tables together to form a 32×32×3 RGB image.

[0017] In the second aspect, the specific steps of the above Step 2 are as follows:

[0018] Step 2.1: Multiply the noise vector by the label vector to form a labeled noise vector and input it into the Generator network. Then use a multi-layer Transformer Encoder to reduce the embedding dimension and learn the hierarchical features of the RGB image.

[0019] Step 2.2: In each layer of the Transformer Encoder, use the spatial-channel collaborative self-attention mechanism to replace the multi-head attention mechanism. Guide the channel attention learning through its spatial attention, strengthen the capture of local features while capturing global context dependencies, and significantly reduce the computational complexity. Finally, calculate the FID evaluation metric of the Generator.

[0020] In the third aspect, the specific steps of the above step 3 are as follows:

[0021] Step 3.1: Segment the input image into 8×8 sub-images and enhance the sub-images to increase the model's sensitivity to local features.

[0022] Step 3.2: Modify the network output layer in the way of class probability output, so that the Discriminator outputs both class probability and true / false probability at the same time, and calculate the performance loss, gradient loss, classification loss of the Discriminator on real images and generated images, as well as the performance loss of the Generator.

[0023] Step 3.3: Dynamically balance the performance loss and the classification loss according to the FID metric, so that the model can balance the realism and class consistency of the images.

[0024] Compared with the prior art, the beneficial effects of the present invention are:

[0025] 1. A novel RGB image feature representation method for IoT network traffic is proposed. First, use data cleaning technology to extract useful features in IoT network traffic, and evaluate the correlation between the extracted features and attack category features to obtain the non-linear relationship between features. Then, normalize the features and form an RGB image in a stacked manner to upgrade the two-dimensional features of the grayscale image to three dimensions, thereby improving the representation ability of network traffic.

[0026] 2. Aiming at the characteristics of IoT network traffic RGB images and the deficiencies brought about by the lack of convolutional layers in TransGAN, a spatial-channel collaborative self-attention mechanism is introduced into TransGAN to capture sufficient local features of network traffic, thereby improving the quality of sample generation.

[0027] 3. Regarding the mode collapse problem of generative adversarial networks, an auxiliary classifier is introduced in TransGAN to enable both the generator and the discriminator to learn specified features according to input conditions and output specified images according to conditions, thereby significantly improving the quality of generated samples and avoiding mode collapse. In addition, an adaptive loss function that can self-feedback according to the generation quality of samples is designed to better guide the direction of sample generation. Description of the Drawings

[0028] Figure 1 is the overall flowchart of a data augmentation method CT-SSSA (Classifier TransGAN and Spatial-channel Synergistic Self-Attention) based on classifier TransGAN and spatial-channel synergistic self-attention.

[0029] Figure 2 is the detailed flowchart of a data augmentation method CT-SSSA based on classifier TransGAN and spatial-channel synergistic self-attention.

[0030] Figure 3 is the information of the network traffic dataset CSE-CIC-IDS2018 used in the experimental section of the present invention.

[0031] Figure 4 is an example of converting the CSE-CIC-IDS2018 dataset into RGB using mutual information in the present invention.

[0032] Figure 5 is the improvement effect of the CT-SSSA model proposed in the present invention using mutual information to preprocess into RGB images, including the comparison of the true positive rate (TPR), false positive rate (FPR), F1-score, and accuracy of four detection models, namely CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory Network), and TCN (Temporal Convolutional Network), for three network traffic types.

[0033] Figure 6 is the detection confusion matrix result obtained by the method proposed in the present invention using CNN and RNN models on the original dataset, where the left figure is the detection effect of CNN and the right figure is the detection effect of RNN.

[0034] Figure 7 is the confusion matrix result obtained by the method proposed in the present invention using LSTM and TCN models on the original dataset, where the left figure is the detection effect of LSTM and the right figure is the detection effect of TCN.

[0035] Figure 8 It is the detection confusion matrix result obtained by the comparative TransGAN method using CNN and RNN models on the original dataset. The left figure shows the detection effect of CNN, and the right figure shows the detection effect of RNN.

[0036] Figure 9 It is the confusion matrix result obtained by the comparative TransGAN method using LSTM and TCN models on the original dataset. The left figure shows the detection effect of LSTM, and the right figure shows the detection effect of TCN.

[0037] Figure 10 It is the ablation experiment result of CT-SSSA proposed by the present invention. The ordinate is the FID value, and the abscissa is the number of rounds. The TransGAN method is used as the base model. TransGAN_classes means modifying the output layer of the Discriminator in the TransGAN method to multi-classification. AcTransGAN_SCSSA means the method after introducing the Ac (Auxiliary classifier) auxiliary classifier and SCSSA (Spatial Channel Collaborative Self-Attention) on the basis of TransGAN. AcTransGAN_SCSSA_adaption means the method after introducing the adaption adaptive loss function on the basis of the AcTransGAN_SCSSA method.

[0038] Figure 11 It is the best FID result of each ablation experiment of CT-SSSA proposed by the present invention. Detailed implementation manners

[0039] The present invention will be further described below in conjunction with the accompanying drawings and implementation cases. It should be noted that the described implementation cases are only for facilitating the understanding of the present invention and do not limit it in any way.

[0040] The present invention aims at malicious network traffic and proposes a data augmentation method based on the classifier TransGAN and spatial channel collaborative self-attention to effectively perform data augmentation on malicious network attack behaviors. The invention provides a perfect malicious network traffic data augmentation model and conducts sufficient experiments to prove the feasibility and effectiveness of the method.

[0041] As Figure 1 shown, a data augmentation method based on the classifier TransGAN and spatial channel collaborative self-attention of the present invention includes:

[0042] Step 201 performs data cleaning on the network traffic dataset, extracts the key features of malicious traffic using mutual information, and converts the network traffic in CSV format into RGB images by means of value mapping and superposition.

[0043] In the embodiments of the present invention, the purpose of constructing RGB image data for network traffic is that the local, global, and correlation relationships between network traffics can all be represented by pixel values, which can greatly facilitate data augmentation of network traffic in the form of images.

[0044] Step 2011 analyzes the network traffic dataset to identify missing values and unreasonable samples. Among them, missing values in large-scale normal network traffic are directly deleted, while missing values in small-scale malicious traffic are processed by mean filling. In addition, unreasonable values such as nan, -inf, and +inf are directly deleted to avoid introducing noise.

[0045] Step 2012 deletes features that are meaningless for determining malicious traffic generated by attacks and normal traffic generated by normal access behaviors from the network traffic dataset after traffic cleaning, including flow ID, source IP, source port, destination IP, destination port, protocol, timestamp, sub-flow related features, and Flag marks, and calculates the mutual information I between the remaining feature columns and malicious traffic and normal traffic. Sort the features in the network traffic according to the value of I and select the 32 most relevant features.

[0046] Step 2013 separates all normal network traffic and malicious traffic into two groups of data, uses Min-Max normalization to map each feature column to the range of 0 to 255 to solve the problem of uneven distribution of traffic feature values, and then iteratively selects 32 rows of traffic by rounding down to form a sub-table. Each sub-table forms a vector of 32×32×1 dimension. Then, every three selected sub-tables form the three channels R, G, and B of the image. Finally, use the OpenCV library to combine the three sub-tables together to form a 32×32×3 RGB image.

[0047] Step 202 uses a spatial-channel collaborative self-attention mechanism to capture multi-semantic information of network traffic data in the RGB image, and combines the category information to input it into the generator of TransGAN to generate labeled traffic samples.

[0048] Step 2021 multiplies the noise vector by the label vector to form a labeled noise vector and inputs it into the Generator network, and then uses a multi-layer Transformer Encoder to reduce the embedding dimension and learn the hierarchical features of the RGB image.

[0049] Step 2022 uses a spatial-channel collaborative self-attention mechanism to replace the multi-head attention mechanism in each layer of the Transformer Encoder. It guides the channel attention learning through its spatial attention, strengthens the capture of local features while capturing global context dependencies, and significantly reduces the computational complexity. Finally, calculate the evaluation metric FID of the Generator;

[0050] Among them, using spatial attention to guide the channel attention learning process includes:

[0051] (1) First, split the input sequence N = H×W into S H and S W along the height and width, and evenly split the channel C into four subsets of the same size to form the following independent sub-feature sets

[0052]

[0053] where i represents the serial number of the i-th subset after splitting, and C represents the length of the image channel.

[0054] (2) To enrich semantic information and enhance semantic coherence, apply depth 1D convolutions with sizes 3, 5, 7, and 9 respectively to the four sub-features to learn and construct semantic associations in the height and width dimensions and extract multi-semantic spatial information. Then, aggregate the two dimensions Att H and Att W through group normalization (GN) and activation functions as shown below to obtain the semantic information Att S ;

[0055]

[0056] where DWConv1d represents the depthwise separable convolution kernel, and represent and the sub-feature sets after being processed by DWConv1d (convolution kernel sizes are 3, 5, 7, and 9) respectively. Sigmoid represents the activation function in the neural network. GN H and GN W represent group normalization along the height and width respectively. Concat represents serial concatenation of the content it contains, including the four sub-feature sets split along the height and the four sub-feature sets split along the width S represents the original image feature set, ⊕ represents the matrix multiplication operation, and Att H and Att WPerform a matrix multiplication operation with S to output the final spatial semantic information Att S .

[0057] (3) Mine the dependencies F between channels through GN and 1×1 convolution operations. Self-attention calculates Q (query vector Query), K (key vector Key), V (value vector Value), and their attention weights Att along the channel dimension C ; Finally, the results are output after layer normalization (LN) and the MLP layer. The process is as follows.

[0058] Q = F Q (Att S ), K = F K (Att S ), V = F V (Att S )

[0059]

[0060] where represents depthwise separable convolution using a 1x1 convolution kernel along the channel, and Softmax represents the normalized exponential function. The final channel semantic information Att is output after Softmax C .

[0061] In step 203, an additional class output layer is introduced to the generator to distinguish whether the input sample comes from the real dataset or the labeled sample generated by the generator. By dynamically balancing the adaptive loss function and combining the class loss and the discrimination loss, the training of the generator and the discriminator is guided.

[0062] In step 2031, the input image is segmented into 8×8 sub-images, and the sub-images are enhanced to increase the model's sensitivity to local features;

[0063] Specifically, DiffAugment (differentiable augmentation) is used to enhance the sub-images, and then they are converted into 1D sequences through a Linear layer to obtain the tokens carrying the key information of the image, enabling the Discriminator to perform more effective feature learning and classification in subsequent processing.

[0064] In step 2032, the network output layer is modified by using the class probability output method, so that the Discriminator outputs both class probabilities and true / false probabilities at the same time, and the performance loss, gradient loss, classification loss of the Discriminator on real images and generated images, as well as the performance loss of the Generator are calculated;

[0065] For multi-class and class-imbalanced generation tasks, simple binary classification output will perform poorly because the model tends to output the class with a larger number in the case of class imbalance and ignores the importance of the less represented classes. Therefore, the output layer of the network is modified to use class probability output, which provides the probability distribution of each class in addition to the traditional True or False output, enabling the Discriminator to consider the distribution characteristics of each class while discriminating the generated images, thereby improving the overall judgment ability of the model and the accuracy of the generation effect.

[0066] Among them, the construction processes of the loss functions of the Generator and the Discriminator are as follows:

[0067] (1) Calculate the performance loss and classification loss of the Generator;

[0068]

[0069] where D is the Discriminator network, G is the Generator network, Z is the input random noise vector, C F is the current class label of Z, and X fake is the image generated by the Generator according to Z and C F . denotes the mathematical expectation of the random variable Z under the distribution, and denotes the mathematical expectation of the random variables (Z, C and ) sampled from the distributions F . P(C F |X fake ) represents the conditional probability given the fake sample X fake , indicating the probability that the generated sample belongs to a certain condition C F , and L f represents the Generator classification loss.

[0070] (2) Calculate the performance loss L D , gradient loss L GP , and classification loss L r of the Discriminator;

[0071]

[0072] where D is the Discriminator network, X real is the real image, and X fake is the image generated by the Generator according to Z and CF The generated image, C T is X real the true class label of, denotes in the distribution of the random variable X real the mathematical expectation of, x is a sample sampled from the distribution, which is obtained by interpolating between real data and generated data, denoting the expectation of the real data distribution of, the coefficient λ is a hyperparameter that controls the strength of the gradient penalty, denotes denotes in the distribution of the mathematical expectation of the random variable x, denotes the gradient of the discriminator D's output for the real sample x, denotes the L 2 norm of the gradient, measuring the magnitude of the gradient, denotes for the real sample X sampled from the real data distribution and the label C real the expectation of, P(C T ∣X T ) denotes the probability distribution of the label C given the real sample X real , logP(C real ∣X T ) denotes the logarithm of the conditional probability P(C T ∣X real ) T ∣X real )

[0073] Step 2033 dynamically balances the performance loss and the classification loss according to the FID metric, so that the model can balance the fidelity and class consistency of the images.

[0074] To better guide the training of the Generator and the Discriminator, during the training process, when the quality of the generated images is poor (high FID value), the model gives priority to the fidelity of the Generator, and when the quality of the generated images is high (low FID value), the model pays more attention to the class consistency of the generated images. For this reason, coefficients λ r and λ f are added to dynamically adjust the weights of the loss function, so that the model can balance the fidelity and class consistency of the images and ensure that λ is always non-negative. The calculation formula is as follows:

[0075]

[0076] λ r= 1 - λ f

[0077] where represents the maximum value of the parameter λ f and k is a positive hyperparameter that controls the slope of the exponential function in the formula and affects the parameter λ f 's sensitivity to changes in the FID value. FID represents the Fréchet Inception Distance, a metric used to evaluate the quality of generated images. A lower value indicates that the generated images are more similar to the real images. FID mid is the intermediate FID value that determines the center point of the weight conversion.

[0078] The introduction of the loss coefficient makes the relationship between λ f and the FID value smoother and more flexible and always keeps λ non - negative. When FID > FID mid , λ f → 0, giving priority to focusing on the realism of the generated images; when FID < FID mid , λ f → 1, giving priority to focusing on the class consistency of the generated images; finally, the following final loss function is obtained and

[0079]

[0080] The present invention mainly performs data augmentation on malicious network traffic and uses the network traffic data CSE - CIC - IDS2018 for effect testing. Figure 5 It shows that the mutual information extracts the key features of network traffic, and through value mapping, the traffic is superimposed and processed into RGB image representation. From the comparison of the true positive rate (TPR), false positive rate (FPR), F1 - score, and accuracy of three types of network traffic, it can be found that this method can better retain the features of network traffic, thereby improving the effectiveness of malicious traffic detection.

[0081] To verify the data augmentation effect of the data augmentation method CT - SSSA based on the classifier TransGAN and spatial - channel collaborative self - attention proposed by the present invention, some generated samples are selected from the above - mentioned network traffic dataset for effect display. The data augmentation effect is as shown in Figure 6 and Figure 7 shown, and the generation effect of the compared model TransGAN is as shown in Figure 8 and Figure 9As shown; it can be observed that the enhancement effect of CT-SSSA has a higher accuracy rate than that of TransGAN in the detection of each model. In addition, in order to verify the improvement effect of each part of CT-SSSA, the statistical FID line chart and the FID value results are as Figure 10 and Figure 11 shown.

Claims

1. A data enhancement method based on classifier TransGAN and spatial channel collaborative self-attention, characterized in that: The steps include: Step 1: Clean the network traffic data set, extract the key features of malicious traffic using mutual information, and convert the network traffic CSV format into RGB images by value mapping and superposition; Step 2: Use the spatial channel collaborative self-attention mechanism to capture the multi-semantic information of the network traffic data in the RGB image, and input it into the generator of TransGAN to generate labeled traffic samples in combination with the category information; In step 3, an additional category output layer is introduced to the generator to distinguish whether the input sample comes from the real dataset or the labeled sample generated by the generator. The training of the generator and discriminator is guided by dynamically balancing the adaptive loss function and combining the category loss and judgment loss.

2. The method according to claim 1, characterized in that The specific implementation of step 1 includes the following steps: Step 1.1: Analyze the network traffic data set to identify missing values ​​and unreasonable samples. Missing values ​​in large-scale normal network traffic are directly deleted, while missing values ​​in small-scale malicious traffic are processed by mean filling. In addition, unreasonable values ​​such as nan, -inf, +inf are directly deleted to avoid introducing noise. Step 1.2, delete the features that are meaningless for judging malicious traffic generated by attacks and normal traffic generated by normal access behaviors from the network traffic data set after traffic cleaning, including flow ID, source IP, source port, destination IP, destination port, protocol, timestamp, sub-traffic related features and Flag tags, and calculate the mutual information I between the remaining feature columns and malicious traffic and normal traffic, sort the features in the network traffic according to the value of I and select the 32 features with the highest correlation; In step 1.3, all normal traffic and malicious traffic are separated into two sets of data. Min-Max normalization is used to map each feature column to the range of 0 to 255 to solve the problem of uneven distribution of traffic feature values. Then, 32 rows of traffic are iteratively selected to form a sub-table by rounding down. Each sub-table constitutes a 32×32×1-dimensional vector. Then, three sub-tables are selected to form the R, G, and B channels of the image. Finally, the three sub-tables are combined together using the OpenCV library to form a 32×32×3 RGB image.

3. The method according to claim 1, characterized in that The specific implementation of step 2 includes the following steps: Step 2.1, multiply the noise vector by the label vector to form a labeled noise vector and input it into the Generator network, then use a multi-layer Transformer Encoder to reduce the embedding dimension and learn the hierarchical features of the RGB image; In step 2.2, the spatial channel collaborative self-attention mechanism is used to replace the multi-head attention mechanism in each layer of Transformer Encoder. The spatial attention is used to guide the channel attention learning, and the capture of local features is strengthened while capturing global context dependencies, while significantly reducing the computational complexity. Finally, the evaluation indicator FID (Fréchet Inception Distance) of the Generator is calculated.

4. The method according to claim 1, characterized in that The specific implementation of step 3 includes the following steps: Step 3.1, split the input image into 8×8 sub-images and enhance the sub-images to enhance the model’s sensitivity to local features; Step 3.2, modify the network output layer by using the category probability output method, so that the Discriminator outputs the category probability and the true and false probability at the same time, and calculate the performance loss, gradient loss and classification loss of the Discriminator on the real image and the generated image, as well as the performance loss of the Generator; In step 3.3, the performance loss and classification loss are dynamically balanced according to the FID metric, so that the model can balance the realism and category consistency of the image.

Citation Information

Cited By

  • Intrusion detection method based on visual self-attention mechanism

    CN121984749A