Traffic data category balancing method based on VAE-CGAN fusion module
By using the traffic data category balance method of the VAE-CGAN fusion module in network traffic detection, the original data is preprocessed and data sample expansion is solved, and the problem of data sample category imbalance is improved. The abnormal traffic detection model detects intrusions of a few categories is improved.
Patent Information
- Application Number
- CN202510156232.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The prior art has poor performance in the face of a few categories of intrusion samples in network traffic detection due to data sample categories imbalance.
The traffic data category balance method based on the VAE-CGAN fusion module is adopted to preprocess the original data and expand the data sample to generate a few new intrusion data to ensure the consistency of the data distribution, thereby obtaining the balanced data set.
While solving the problem of data imbalance, the sample diversity is increased, and the generated data is more accurate and diversified, making the abnormal traffic detection model have higher detection capabilities when dealing with a few types of intrusions.
Smart Images

Figure CN120074894A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent network security defense, and particularly relates to a method for balancing the categories of traffic data based on a VAE-CGAN fusion module. Background Art
[0002] Current AI-based network traffic detection methods show potential effects in new intrusion detection by learning the characteristics of normal traffic and intrusion traffic. Through technologies such as deep learning, these methods can extract key features from massive data to assist network security personnel in timely discovering and dealing with various threats. However, although these methods show great potential in theory, they still face some severe challenges in practical applications. Among them, the imbalance of data sample categories is a technical bottleneck of current network abnormal traffic detection methods. Since normal traffic is much larger than abnormal traffic, there is an obvious imbalance in the training data, resulting in poor performance of the model when facing a small number of intrusion samples of certain categories.
[0003] Current processing of imbalanced data mainly focuses on methods such as oversampling and undersampling of data. However, the oversampling method is to repeatedly learn the same data, which may lead to overfitting of the classifier. At the same time, the undersampling method will cause the loss of some information. In addition, for some samples of minority-class intrusions, traditional processing methods cannot solve the problem of category imbalance. Therefore, conducting research on AI-based network traffic detection methods to solve the problem of category imbalance in traffic data has significant research significance for improving network security defense capabilities and strengthening the prediction and prevention of new network threats. This will not only help improve the actual effect of network security technologies but also provide technical support for building a more powerful and intelligent network security system, and can better maintain the network security and stability in the digital age. Summary of the Invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides a method for balancing the categories of traffic data based on a VAE-CGAN fusion module. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0005] In a first aspect, the present invention provides a method for balancing the categories of traffic data based on a VAE-CGAN fusion module, including:
[0006] Preprocess the original data in the obtained abnormal dataset for network intrusion detection, and label the intrusion behavior types of the abnormal data in the original data to obtain an original dataset;
[0007] Using a pre-trained traffic data class balance model based on the VAE-CGAN fusion module, the abnormal data in the original data set is augmented with data samples on the premise of maintaining the consistency of the data distribution, and a balanced data set is obtained;
[0008] Among them, the traffic data class balance model based on the VAE-CGAN fusion module includes:
[0009] An encoder module, an unbalanced data filtering module, a generator module, and a discriminator module.
[0010] In an embodiment of the present invention, the preprocessing of the original data in the obtained network intrusion detection abnormal data set includes:
[0011] Using one-hot encoding to convert non-numerical features in the network intrusion detection abnormal data set into numerical features;
[0012] Using the isolation forest method to exclude outliers in the converted network intrusion detection abnormal data set;
[0013] The converted network intrusion detection abnormal data set is processed by Min-Max maximum-minimum normalization, and each data in the network intrusion detection abnormal data set is divided according to the label corresponding to the data.
[0014] In an embodiment of the present invention, using the isolation forest method to exclude outliers in the converted network intrusion detection abnormal data set includes:
[0015] Using the isolation forest method to process the sample data in the converted network intrusion detection abnormal data set, traversing each tree iTree to obtain the path length h(x) of the sample data;
[0016] Using the first formula, according to the path length h(x), the outlier score s(x, n) corresponding to each sample data in the network intrusion detection abnormal data set is obtained; where the first formula is as follows:
[0017]
[0018] E(h(x)) represents the expectation of the path length h(x), h(x) represents the number of splits required to separate a sample data in the process of obtaining the path length of the sample data, and c(n) represents the average height of each tree iTree;
[0019] According to the gap between the outlier score s(x, n) corresponding to each sample data in the network intrusion detection abnormal data set and the preset threshold, the outliers in the converted network intrusion detection abnormal data set are excluded.
[0020] In one embodiment of the present invention, the process of obtaining the pre-trained traffic data class balance model based on the VAE-CGAN fusion module includes:
[0021] Training the traffic data class balance model based on the VAE-CGAN fusion module;
[0022] Evaluating the trained traffic data class balance model based on the VAE-CGAN fusion module, and using the traffic data class balance model based on the VAE-CGAN fusion module that meets the requirements as the pre-trained traffic data class balance model based on the VAE-CGAN fusion module.
[0023] In one embodiment of the present invention, the process of training the traffic data class balance model based on the VAE-CGAN fusion module includes:
[0024] S01, encoding the sample data in the original dataset using the encoder to obtain latent variables corresponding to all sample data;
[0025] S02, screening the minority class data in the sample data in the original dataset using the imbalanced data filtering module to obtain target sample data;
[0026] S03, using the generator module to obtain pseudo-sample data according to the latent variables corresponding to the target sample data;
[0027] S04, using the discriminator module to judge the target sample data and the pseudo-sample data, and outputting the class label of the predicted pseudo-sample data;
[0028] S05, training the generator module according to the error between the predicted class label and the true class label to generate new pseudo-sample data;
[0029] S06, repeatedly executing steps S03 - S05 until the number of loops reaches the preset loop value or the loss value of the traffic data class balance model based on the VAE-CGAN fusion module reaches the preset loss value, outputting the final pseudo-sample data, and optimizing the gradient update process of the traffic data class balance model based on the VAE-CGAN fusion module using the SGD algorithm during the repeated execution of steps S03 - S05 to obtain the trained traffic data class balance model based on the VAE-CGAN fusion module.
[0030] In one embodiment of the present invention, the process of evaluating the trained traffic data class balance model based on the VAE-CGAN fusion module includes:
[0031] Input the original dataset into the trained traffic data class balance model based on the VAE-CGAN fusion module to obtain the balanced dataset;
[0032] Use the pre-trained CNN abnormal traffic detection model to classify the original dataset and the balanced dataset to obtain the classification results;
[0033] Obtain the evaluation metrics of the traffic data class balance model based on the VAE-CGAN fusion module based on the classification results;
[0034] Evaluate the trained traffic data class balance model based on the VAE-CGAN fusion module according to the evaluation metrics of the model. If the preset requirements are met, use the trained traffic data class balance model based on the VAE-CGAN fusion module as the pre-trained traffic data class balance model based on the VAE-CGAN fusion module. If the preset requirements are not met, continue to train the trained traffic data class balance model based on the VAE-CGAN fusion module until the obtained traffic data class balance model based on the VAE-CGAN fusion module meets the preset requirements.
[0035] In an embodiment of the present invention, obtaining the evaluation metrics of the model based on the classification results includes:
[0036] According to the classification results, obtain the number of samples TP with correctly predicted normal traffic type, the number of samples TN with correctly predicted intrusion traffic type, the number of samples FP with incorrectly predicted normal traffic type, and the number of samples FN with incorrectly predicted intrusion traffic type;
[0037] According to the number of samples TP with correctly predicted normal traffic type, the number of samples TN with correctly predicted intrusion traffic type, the number of samples FP with incorrectly predicted normal traffic type, and the number of samples FN with incorrectly predicted intrusion traffic type, obtain the evaluation metrics of the model; where the evaluation metrics of the model include: accuracy Accuray, precision Precison, recall Recall, and F1-score F1-score; where,
[0038] The expression of accuracy Accuray is as follows:
[0039]
[0040] The expression of precision Precison is as follows:
[0041]
[0042] The expression of recall Recall is as follows:
[0043]
[0044] The expression of the F1-score is as follows:
[0045]
[0046] In one embodiment of the present invention, by using a pre-trained traffic data class balance model based on the VAE-CGAN fusion module, on the premise of maintaining the consistency of the data distribution, the abnormal data in the original data set is used for data sample augmentation to obtain a balanced data set, including:
[0047] Using the pre-trained traffic data class balance model based on the VAE-CGAN fusion module to generate corresponding pseudo-sample data as intrusion samples according to the minority-class data in the original data set;
[0048] Merging the intrusion samples with the original data set to obtain a balanced data set.
[0049] In one embodiment of the present invention, the loss function of the generator module in the traffic data class balance model based on the VAE-CGAN fusion module is as follows:
[0050]
[0051] where z represents the encoded latent variable, z ∼ P z represents the distribution followed by the data generated by the generator module, y′ represents the label value of the sample data filtered by the unbalanced data filtering module, and (G(z, y′), y′) represents the sample generated by the generator module;
[0052] The loss function of the discriminator module is as follows:
[0053]
[0054] where x ∼ Pr represents the distribution followed by the sample data, and D((G(z, y′), y′)|y′) represents the discrimination probability of the discriminator module for the sample (G(z, y′), y′) generated by the generator module under the given conditions.
[0055] In a second aspect, the present invention provides a classification method for traffic data class balance based on the VAE-CGAN fusion module, including:
[0056] Obtaining a network intrusion detection abnormal data set;
[0057] Using the traffic data class balance method based on the VAE-CGAN fusion module as described in the first aspect to perform class balance processing on the original data in the network intrusion detection abnormal data set to obtain a balanced data set;
[0058] Use the pre-trained abnormal traffic detection model to classify the balanced dataset and the network intrusion detection abnormal dataset to obtain the classification results.
[0059] Advantages of the present invention:
[0060] In the solution provided by the present invention, by preprocessing the original data, the outlier monitoring of the normal class sample data in the original data is realized, and the boundary overlap with the minority intrusion class samples is reduced. Then, the pre-trained traffic data class balance model based on the VAE-CGAN fusion module is used to expand the data samples of the original dataset obtained by preprocessing, and new minority intrusion class data is synthesized on the premise of ensuring the consistency of data distribution, so as to obtain the balanced dataset; while solving the problem of data imbalance, the sample diversity is increased, and the generated data is more accurate and diverse, so that the abnormal traffic detection model has higher detection ability when dealing with minority class intrusions. Description of the Drawings
[0061] Figure 1 It is a schematic diagram of the steps of a traffic data class balance method based on the VAE-CGAN fusion module provided by an embodiment of the present invention;
[0062] Figure 2 It is a schematic diagram of the structure of a traffic data class balance model based on the VAE-CGAN fusion module provided by an embodiment of the present invention;
[0063] Figure 3 It is a schematic diagram of the steps of the training process of a traffic data class balance model based on the VAE-CGAN fusion module provided by an embodiment of the present invention;
[0064] Figure 4 It is a data distribution diagram of the CICIDS2017 dataset before and after balancing provided by an embodiment of the present invention;
[0065] Figure 5 It is a data distribution diagram of the NSL-KDD dataset before and after balancing provided by an embodiment of the present invention;
[0066] Figure 6 It is a performance improvement effect diagram of the CICIDS2017 dataset provided by an embodiment of the present invention in four evaluation indicators of accuracy, recall rate, precision rate and F1 value;
[0067] Figure 7 It is a performance improvement effect diagram of the NSL-KDD dataset provided by an embodiment of the present invention in four evaluation indicators of accuracy, recall rate, precision rate and F1 value. Detailed Embodiments
[0068] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0069] In a first aspect, an embodiment of the present invention provides a traffic data class balance method based on a VAE-CGAN fusion module, as Figure 1 shown, which may include:
[0070] S1. Preprocess the original data in the obtained network intrusion detection abnormal dataset, and label the intrusion behavior types of the abnormal data in the original data to obtain an original dataset.
[0071] Regarding S1, preprocessing the original data in the obtained network intrusion detection abnormal dataset may include:
[0072] S11. Use one-hot encoding to convert non-numerical features in the network intrusion detection abnormal dataset into numerical features;
[0073] S12. Use the isolation forest method to exclude outliers in the converted network intrusion detection abnormal dataset;
[0074] Specifically, regarding S12, the following process may be included:
[0075] S121. Use the isolation forest method to process the sample data in the converted network intrusion detection abnormal dataset, traverse each tree iTree, and obtain the path length h(x) of the sample data, which may include:
[0076] S1211. Randomly select n sample data from the converted network intrusion detection abnormal dataset and put them into the root node of a tree iTree; set the maximum value of the selected sample data as X max , and the minimum value as X min . Randomly select a dimension in the sample data and a certain split point p(X min <p<X max ), and divide the sample data into two subspaces: put the sample data with values less than p in the left node, and put the sample data with values greater than or equal to p in the right node;
[0077] S1212. Recursively repeat step S1211 in the node to continuously generate new nodes until there is only one sample data in each node;
[0078] S1213. Loop through steps S121 - S122 to continue generating and training the next tree iTree until a preset number of tree iTrees are generated;
[0079] S1214. Calculate the average height c(i) of each tree iTree. For each sample data, calculate its average path length among all the trees iTree. Among them, the expression of the average height c(i) is as follows:
[0080]
[0081] Among them, H(i) represents the harmonic number, and the expression of H(i) is as follows:
[0082] H(i) = ln(i) + δ, where δ ≈ 0.577.
[0083] S122. Use the first formula to obtain the anomaly score s(x, n) corresponding to each sample data in the network intrusion detection anomaly dataset according to the path length h(x).
[0084] Among them, the first formula is as follows:
[0085]
[0086] Among them, E(h(x)) represents the expectation of the path length h(x). h(x) is equal to the number of splits required to separate a sample data during the process of obtaining the path length of the sample data. c(n) represents the average height of each tree iTree.
[0087] S123. Exclude the outliers in the converted network intrusion detection anomaly dataset according to the gap between the anomaly score s(x, n) corresponding to each sample data in the network intrusion detection anomaly dataset and the preset threshold.
[0088] It can be understood that the closer the anomaly score s(x, n) is to 1, the greater the possibility that the sample data is abnormal. Therefore, when performing step S123, the preset threshold can be set to 1. According to the gap between the anomaly score s(x, n) and the preset threshold, the sample data with a smaller gap is excluded as an outlier. By constructing a series of random trees iTree to isolate data points. Different from traditional methods that rely on the distance or density of data points, the isolation forest gradually splits the dataset by randomly selecting features and split values to form a tree structure. Due to their rarity and distinctive characteristics, outliers are more likely to be isolated and thus will show a shorter path length in the isolation forest.
[0089] S13. Perform Min - Max min - max normalization processing on the network intrusion detection anomaly dataset after excluding outliers, and divide each data in the network intrusion detection anomaly dataset according to the label corresponding to the data.
[0090] It can be understood that the data ranges of different feature values in the dataset may vary greatly. In some cases, the prediction performance of the abnormal traffic detection model depends on the features with larger data ranges, while ignoring the features with smaller data ranges. Therefore, in order to eliminate the influence between different feature values and ensure that the contribution degrees of each feature value to the model are equivalent, the embodiments of the present invention adopt Min-Max maximum-minimum normalization processing to map the numerical values of different features into the range of [0, 1], so that different features contribute equally to the training and prediction of the model, and avoid some data features occupying too large weights. The formula is as follows:
[0091]
[0092] Among them, x represents the original feature value, and x v represents the value after normalization processing, X′ max represents the minimum value of the dimension, X′ min the minimum value of the dimension, and then for each data in the network intrusion detection abnormal dataset, data division is performed according to the label corresponding to the data, and the data corresponding to each class label in the network intrusion detection abnormal dataset is correspondingly divided into a group of data for subsequent calculation and processing.
[0093] It can be understood that by preprocessing the original data, the detection of outliers in the normal class sample data in the original data is realized, and the boundary overlap with the minority intrusion class samples is reduced.
[0094] S2. Using the pre-trained traffic data class balance model based on the VAE-CGAN fusion module, on the premise of maintaining the consistency of data distribution, data sample augmentation is performed on the abnormal data in the original dataset to obtain a balanced dataset.
[0095] Among them, the traffic data class balance model based on the VAE-CGAN fusion module, as Figure 2 shown, may include:
[0096] An encoder module, an unbalanced data filtering module, a generator module, and a discriminator module.
[0097] Specifically, the process of obtaining the pre-trained traffic data class balance model based on the VAE-CGAN fusion module may include:
[0098] Training the traffic data class balance model based on the VAE-CGAN fusion module;
[0099] Evaluate the trained traffic data class balance model based on the VAE-CGAN fusion module, and use the traffic data class balance model based on the VAE-CGAN fusion module that meets the requirements as the pre-trained traffic data class balance model based on the VAE-CGAN fusion module.
[0100] The process of training the traffic data class balance model based on the VAE-CGAN fusion module, as Figure 3 shown, may include:
[0101] S01, Use the encoder to encode the sample data in the original dataset to obtain the latent variables corresponding to all sample data;
[0102] S02, Use the imbalanced data filtering module to screen the minority class data in the sample data of the original dataset to obtain the target sample data;
[0103] S03, Use the generator module to obtain pseudo-sample data according to the latent variables corresponding to the target sample data;
[0104] S04, Use the discriminator module to judge the target sample data and the pseudo-sample data, and output the class label of the predicted pseudo-sample data;
[0105] Specifically, the discriminator module judges the target sample data and the pseudo-sample data respectively, outputs the sample classification probability values of real samples and pseudo-samples, and then converts the sample classification probability values into the class labels of the predicted pseudo-sample data through a preset activation function.
[0106] S05, Train the generator module according to the error between the predicted class label and the real class label to generate new pseudo-sample data;
[0107] It can be understood that during the training process, the generator module is trained according to the error given by the discriminator module. While improving the discrimination ability of the discriminator module, the generator module is used to train the latent variable noise with minority class labels to generate pseudo-sample data with higher simulation degree.
[0108] S06, Loop through steps S03 - S05 until the number of loops reaches the preset loop value or the loss value of the traffic data class balance model based on the VAE-CGAN fusion module reaches the preset loss value, output the final pseudo-sample data, and optimize the gradient update process of the traffic data class balance model based on the VAE-CGAN fusion module using the SGD algorithm during the loop execution of steps S03 - S05 to obtain the trained traffic data class balance model based on the VAE-CGAN fusion module.
[0109] Specifically, in the traffic data category balance model, the target sample data is the minority class sample data. The minority class sample data and the corresponding class labels are augmented according to the training formula, and the training formula is as follows:
[0110]
[0111] Among them, D represents the discriminator module in CGAN, G represents the generator module in CGAN, x' represents the sample data filtered by the unbalanced data filtering module, y' represents the label value of the sample data filtered by the unbalanced data filtering module, z represents the encoded latent variable, and x ∼ P r represents the distribution that the sample data follows, and z ∼ P z represents the distribution that the data generated by the generator module follows. That is, the goal in the VAE-CGAN training stage is for the data generated by the generator module to fit the distribution of the minority class data, and the minority class data is augmented to obtain the augmented data.
[0112] The process of evaluating the trained traffic data category balance model based on the VAE-CGAN fusion module may include:
[0113] Input the original dataset into the trained traffic data category balance model based on the VAE-CGAN fusion module to obtain the balanced dataset;
[0114] Use the pre-trained CNN abnormal traffic detection model to classify the original dataset and the balanced dataset to obtain the classification results;
[0115] Obtain the evaluation metrics of the traffic data category balance model based on the VAE-CGAN fusion module based on the classification results;
[0116] Evaluate the trained traffic data category balance model based on the VAE-CGAN fusion module according to the evaluation metrics of the model. If the preset requirements are met, the trained traffic data category balance model based on the VAE-CGAN fusion module is used as the pre-trained traffic data category balance model based on the VAE-CGAN fusion module. If the preset requirements are not met, continue to train the trained traffic data category balance model based on the VAE-CGAN fusion module until the obtained traffic data category balance model based on the VAE-CGAN fusion module meets the preset requirements.
[0117] It can be understood that the pre-trained CNN abnormal traffic detection model is an existing model that can be used for classification, which will not be elaborated here. For details, please refer to the prior art. By inputting the original dataset obtained through S1 processing and the balanced dataset obtained through S2 processing into the pre-trained CNN abnormal traffic detection model, the corresponding classification results can be obtained.
[0118] Based on the classification results, the evaluation metrics of the model can include:
[0119] According to the classification results, obtain the number of samples TP correctly predicted as the normal type of traffic, the number of samples TN correctly predicted as the intrusion type of traffic, the number of samples FP wrongly predicted as the normal type of traffic, and the number of samples FN wrongly predicted as the intrusion type of traffic;
[0120] According to the number of samples TP correctly predicted as the normal type of traffic, the number of samples TN correctly predicted as the intrusion type of traffic, the number of samples FP wrongly predicted as the normal type of traffic, and the number of samples FN wrongly predicted as the intrusion type of traffic, obtain the evaluation metrics of the model; among them, the evaluation metrics of the model include: accuracy Accuray, precision Precison, recall Recall, and F1-score F1-score. Among them,
[0121] The expression of accuracy Accuray is as follows:
[0122]
[0123] The expression of precision Precison is as follows:
[0124]
[0125] The expression of recall Recall is as follows:
[0126]
[0127] The expression of F1-score F1-score is as follows:
[0128]
[0129] The embodiments of the present invention use accuracy Accuray, precision Precison, recall Recall, and F1-score F1-score as the evaluation metrics of the model to evaluate the performance of the model, including the performance changes before and after balancing and the performance changes of different data balancing methods.
[0130] For S2, it can include:
[0131] S21, using a pre-trained traffic data class balance model based on the VAE-CGAN fusion module to generate corresponding pseudo-sample data as intrusion samples according to the minority class data in the original dataset;
[0132] S22, merging the intrusion samples with the original dataset to obtain a balanced dataset.
[0133] The loss function of the generator module in the traffic data class balance model based on the VAE-CGAN fusion module is as follows:
[0134]
[0135] Among them, z represents the latent variable after encoding, z ∼ P z represents the distribution followed by the data generated by the generator module, y′ represents the label value of the sample data filtered by the unbalanced data filtering module, and (G(z, y′), y′) represents the sample generated by the generator module;
[0136] The loss function of the discriminator module is as follows:
[0137]
[0138] Among them, x ∼ Pr represents the distribution followed by the sample data, and D((G(z, y′), y′)|y′) represents the discrimination probability of the discriminator module for the sample (G(z, y′), y′) generated by the generator module under the given conditions.
[0139] It can be understood that by using the pre-trained traffic data class balance model based on the VAE-CGAN fusion module to perform data sample augmentation on the original data set obtained by preprocessing, new minority intrusion class data is synthesized on the premise of ensuring data distribution consistency, so as to obtain a balanced data set; while solving the data imbalance problem, the sample diversity is increased, and the generated data is more accurate and diverse, making the abnormal traffic detection model have higher detection ability when dealing with minority class intrusions.
[0140] In order to verify the effectiveness of the traffic data class balance method based on the VAE-CGAN fusion module proposed in the embodiments of the present invention, a large number of experiments were carried out on two public data sets, CICIDS2017 and NSL-KDD.
[0141] Figure 4 This is the data distribution diagram of the CICIDS2017 data set before and after balancing provided by the embodiments of the present invention, Figure 5The data distribution diagrams of the NSL-KDD dataset before and after balancing provided by the embodiments of the present invention are as follows. It can be seen that the embodiments of the present invention evaluate five categories of the CICIDS2017 dataset, namely BENIGN, Dos, Port Scan, Brute Force, and Web Attack, and evaluate five categories of the NSL-KDD dataset, namely Normal, Dos, Probe, U2R, and R2L, to ensure the universality of the traffic data category balance model based on the VAE-CGAN fusion module. The performance improvement effect diagrams of the CICIDS2017 dataset in four evaluation metrics, namely accuracy, recall, precision, and F1 value, are as shown in Figure 6 As shown, the performance improvement effect diagrams of the NSL-KDD dataset in four evaluation metrics, namely accuracy, recall, precision, and F1 value, are as shown in Figure 7 As shown. It can be clearly seen from Figure 6 and Figure 7 that for both the CICIDS2017 dataset and the NSL-KDD dataset, the four evaluation metrics of accuracy, recall, precision, and F1 value have been greatly improved compared to the data without processing. Experimental comparison found that the data after category balance processed by the model proposed in the present invention is better than the dataset with category imbalance in the four evaluation metrics of accuracy, recall, precision, and F1 value, and has a good recognition and classification effect. This is because when training a model with an imbalanced dataset, the model will be biased towards majority-class samples and ignore minority-class samples, resulting in underfitting and overfitting problems. After balancing the categories of the dataset, the above problems can be effectively alleviated. Moreover, the traffic data category balance method based on the VAE-CGAN fusion module proposed in the embodiments of the present invention can synthesize new minority intrusion class samples on the premise of ensuring data distribution consistency, solve the data imbalance problem while increasing sample diversity.
[0142] When different models are used to process the CICIDS2017 dataset, the results of the evaluation metrics are shown in Table 1, the corresponding evaluation metrics table of the CICIDS2017 dataset;
[0143] Table 1 Corresponding evaluation metrics table of the CICIDS2017 dataset
[0144]
[0145] When different models are used to process the NSL-KDD dataset, the results of the evaluation metrics are shown in Table 2, the corresponding evaluation metrics table of the NSL-KDD dataset.
[0146] Table 2 Corresponding evaluation metrics table of the NSL-KDD dataset
[0147]
[0148] As can be seen from Table 1 and Table 2, in addition to verifying the performance improvement of the algorithm before and after balancing, the embodiments of the present invention also conduct horizontal comparison experiments on other models to verify the superiority of the traffic data class balancing method based on the VAE-CGAN fusion module proposed by the embodiments of the present invention. Under the condition of the same detection model, the model uses classic data enhancement algorithms such as Random Over Sampler (ROS), Synthetic Minority Oversampling Technique (SMOTE), Adaptive Synthetic (ADASYN), etc. and CGAN to process the imbalanced data set. From the data in the figure, it can be seen that the embodiments of the present invention are superior to other data balancing methods in the four evaluation indexes of accuracy, precision, recall rate and F1 value. This is because ROS only performs simple resampling on the original data, and ADASYN and SMOTE randomly synthesize the original data according to the k-nearest neighbor principle. Neither of them considers the nature of the original data. The present invention introduces VAE into CGAN to synthesize new minority intrusion class samples on the premise of ensuring data distribution consistency, increases sample diversity while solving the data imbalance problem, and enables the model to have higher detection ability when dealing with minority class intrusions.
[0149] In a second aspect, an embodiment of the present invention provides a classification method for traffic data class balancing based on a VAE-CGAN fusion module, which may include:
[0150] Obtain a network intrusion detection abnormal data set;
[0151] Use the traffic data class balancing method based on the VAE-CGAN fusion module as described in the first aspect to perform class balancing processing on the original data in the network intrusion detection abnormal data set to obtain a balanced data set;
[0152] Use a pre-trained abnormal traffic detection model to classify the balanced data set and the network intrusion detection abnormal data set to obtain a classification result.
[0153] Among them, the abnormal traffic detection model may select a CNN abnormal traffic detection model.
[0154] In the embodiments of the present invention, by preprocessing the original data, outlier monitoring of the normal class sample data in the original data is realized, and the boundary overlap with the minority intrusion class samples is reduced. Then, the pre-trained traffic data class balance model based on the VAE-CGAN fusion module is used to expand the data samples of the original data set obtained by preprocessing, and new minority intrusion class data is synthesized on the premise of ensuring the consistency of the data distribution, so as to obtain a balanced data set; while solving the data imbalance problem, the sample diversity is increased, and the generated data is more accurate and diverse, so that the abnormal traffic detection model has higher detection ability when dealing with minority class intrusions.
[0155] It should be noted that in the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.
[0156] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A traffic data category balancing method based on VAE-CGAN fusion module, characterized in that: include: Preprocess the original data in the acquired network intrusion detection anomaly data set, and mark the intrusion behavior type of the anomaly data in the original data to obtain the original data set; Using the pre-trained traffic data category balancing model based on the VAE-CGAN fusion module, the data samples of the abnormal data in the original data set are expanded under the premise of maintaining the consistency of data distribution to obtain a balanced data set; The traffic data category balance model based on the VAE-CGAN fusion module includes: Encoder module, imbalanced data filtering module, generator module, and discriminator module.
2. According to claim 1, a flow data category balancing method based on a VAE-CGAN fusion module is characterized in that: The preprocessing of the original data in the acquired network intrusion detection anomaly data set includes: Using one-hot encoding to convert non-numerical features in the network intrusion detection anomaly data set into numerical features; The isolation forest method is used to exclude outliers from the transformed network intrusion detection anomaly dataset; After excluding outliers, the network intrusion detection anomaly data set is processed by Min-Max maximum and minimum normalization, and each data in the network intrusion detection anomaly data set is divided according to the label corresponding to the data.
3. According to claim 2, a flow data category balancing method based on a VAE-CGAN fusion module is characterized in that: The method of using the isolation forest method to exclude outliers in the converted network intrusion detection anomaly data set includes: The isolation forest method is used to process the sample data in the converted network intrusion detection anomaly data set, traverse each tree iTree, and obtain the path length h(x) of the sample data; Using the first formula, according to the path length h(x), the anomaly score s(x, n) corresponding to each sample data in the network intrusion detection anomaly data set is obtained; wherein the first formula is as follows: E(h(x)) represents the expectation of the path length h(x), h(x) represents the number of splits required to separate a sample data in the process of obtaining the path length of the sample data, and c(n) represents the average height of each tree iTree; According to the difference between the anomaly score s(x, n) corresponding to each sample data in the network intrusion detection anomaly data set and the preset threshold, the outliers in the converted network intrusion detection anomaly data set are excluded.
4. According to the method of claim 1, the flow data category balancing method based on the VAE-CGAN fusion module is characterized in that: The process of obtaining the pre-trained traffic data category balance model based on the VAE-CGAN fusion module includes: Training the traffic data category balancing model based on the VAE-CGAN fusion module; The trained traffic data category balancing model based on the VAE-CGAN fusion module is evaluated, and the traffic data category balancing model based on the VAE-CGAN fusion module that meets the requirements is used as the pre-trained traffic data category balancing model based on the VAE-CGAN fusion module.
5. According to claim 4, a flow data category balancing method based on a VAE-CGAN fusion module is characterized in that: The process of training the traffic data category balancing model based on the VAE-CGAN fusion module includes: S01, using the encoder to encode the sample data in the original data set to obtain hidden variables corresponding to all the sample data; S02, using the unbalanced data filtering module to filter the minority class data in the sample data in the original data set to obtain target sample data; S03, using the generator module to obtain pseudo sample data according to the hidden variables corresponding to the target sample data; S04, using the discriminator module to judge the target sample data and the pseudo sample data, and output the predicted category label of the pseudo sample data; S05, training the generator module according to the error between the predicted category label and the actual category label to generate new pseudo sample data; S06, looping through steps S03-S05 until the number of loops reaches a preset loop value or the loss value of the traffic data category balance model based on the VAE-CGAN fusion module reaches a preset loss value, outputting the final pseudo sample data, and in the process of looping through steps S03-S05, optimizing the gradient update process of the traffic data category balance model based on the VAE-CGAN fusion module using the SGD algorithm to obtain a trained traffic data category balance model based on the VAE-CGAN fusion module.
6. A flow data category balancing method based on VAE-CGAN fusion module according to claim 4, characterized in that: The process of evaluating the trained traffic data category balance model based on the VAE-CGAN fusion module includes: Inputting the original data set into the trained traffic data category balancing model based on the VAE-CGAN fusion module to obtain a balanced data set; Using a pre-trained CNN abnormal traffic detection model to classify the original data set and the balanced data set to obtain a classification result; Based on the classification results, an evaluation index of a traffic data category balance model based on a VAE-CGAN fusion module is obtained; The trained traffic data category balancing model based on the VAE-CGAN fusion module is evaluated according to the evaluation index of the model. If the preset requirements are met, the trained traffic data category balancing model based on the VAE-CGAN fusion module is used as the pre-trained traffic data category balancing model based on the VAE-CGAN fusion module. If the preset requirements are not met, the trained traffic data category balancing model based on the VAE-CGAN fusion module continues to be trained until the obtained traffic data category balancing model based on the VAE-CGAN fusion module meets the preset requirements.
7. A flow data category balancing method based on VAE-CGAN fusion module according to claim 6, characterized in that: The evaluation index of the model is obtained based on the classification result, including: According to the classification results, the number of samples TP correctly predicted as normal traffic, the number of samples TN correctly predicted as intrusion traffic, the number of samples FP wrongly predicted as normal traffic, and the number of samples FN wrongly predicted as intrusion traffic are obtained; According to the number of samples TP of correctly predicted traffic as normal type, the number of samples TN of correctly predicted traffic as intrusion type, the number of samples FP of incorrectly predicted traffic as normal type and the number of samples FN of incorrectly predicted traffic as intrusion type, the evaluation index of the model is obtained; wherein the evaluation index of the model includes: accuracy Accuray, precision Precison, recall Recall and F1 score F1-score; wherein, The expression of accuracy Accuray is as follows: The expression of precision is as follows: The expression of recall rate Recall is as follows: The expression of F1 score F1-score is as follows:
8. The method for balancing traffic data categories based on the VAE-CGAN fusion module according to claim 1, characterized in that: The pre-trained traffic data category balancing model based on the VAE-CGAN fusion module is used to expand the data samples of the abnormal data in the original data set under the premise of maintaining the consistency of data distribution, so as to obtain a balanced data set, including: Generate corresponding pseudo sample data as intrusion samples according to the minority class data in the original data set using the pre-trained traffic data category balance model based on the VAE-CGAN fusion module; The intrusion sample is merged with the original data set to obtain a balanced data set.
9. The method for balancing traffic data categories based on the VAE-CGAN fusion module according to claim 1, characterized in that: The loss function of the generator module in the traffic data category balancing model based on the VAE-CGAN fusion module is as follows: Among them, z represents the encoded latent variable, represents the distribution of the data generated by the generator module, y′ represents the label value of the sample data after being filtered by the imbalanced data filtering module, and (G(z,y′),y′) represents the sample generated by the generator module; The loss function of the discriminator module is as follows: Among them, x~P r represents the distribution that the sample data obeys, and D((G(z,y′),y′)y′) represents the probability of the discriminator module distinguishing the sample (G(z,y′),y′) generated by the generator module under given conditions.
10. A classification method for traffic data category balance based on VAE-CGAN fusion module, characterized in that: include: Obtain network intrusion detection anomaly dataset; Using the traffic data category balancing method based on the VAE-CGAN fusion module as described in any one of claims 1 to 9, the original data in the network intrusion detection anomaly data set is subjected to category balancing processing to obtain a balanced data set; The balanced data set and the network intrusion detection anomaly data set are classified using a pre-trained anomaly traffic detection model to obtain a classification result.
Citation Information
Patent Citations
Efficient approximate query processing algorithm based on conditional generative model
CN113177078A
Network traffic anomaly detection method and system based on VAE-CWGAN model, and storage medium
CN117336036A
Intrusion detection method based on incremental training
CN118282707A
Abnormal traffic classification detection method based on CGAN and TabTransform
CN118861910A
Network intrusion detection method fusing Balanced WCGAN-GP and IBA0A feature selection
CN118972177A