Network intrusion detection model based on generative adversarial and multi-stack ensemble learning

By introducing generative adversarial networks and multi-stack integrated learning strategies into the network intrusion detection model, the data imbalance problem is solved, the accuracy and stability of detection are improved, and more effective malicious traffic detection is achieved.

CN120223375AInactive Publication Date: 2025-06-27TAISHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312487.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing network intrusion detection model has data imbalance problem when processing network traffic data, resulting in the model preferring normal traffic categories with a large amount of data and ignoring the small number of malicious traffic categories, affecting the accuracy of detection.

Method used

A network intrusion detection model based on generative adversarial network (GAN) and multi-stack integrated learning is adopted to improve the accuracy and generalization ability of detection through iterative search feature selection, data augmentation and multi-stack integrated learning strategies.

Benefits of technology

It improves the accuracy, accuracy and stability of network intrusion detection, is better than the existing network intrusion detection methods based on integrated learning, and can detect malicious traffic more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223375A_ABST
    Figure CN120223375A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of network intrusion detection, and discloses a network intrusion detection model based on generative adversarial and multi-stack ensemble learning, comprising the following steps: step 1, preprocessing network traffic data to select more important features for a network intrusion detection task; step 2, learning distribution of network traffic characteristics for the preprocessed network traffic data by using an improved generative adversarial network, thereby expanding a network traffic sample in a small-scale category; and step 3, inputting the enhanced network traffic sample in the step 2 into a network intrusion detection model based on multi-stack ensemble learning for training so as to obtain a more accurate detection result. According to the method, a public available network flow data set is preprocessed according to the step 1, data enhancement is realized according to the step 2, and then enhanced network flow data is put into an intrusion detection model of multi-stack ensemble learning to obtain results with better accuracy, precision, F1-score, FPR and detection stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network intrusion detection, and relates to a network intrusion detection model based on generative adversarial and multi-stack ensemble learning. Background Art

[0002] With the popularization of network technology and the continuous development of network applications, network intrusion events such as zero-day vulnerabilities, advanced persistent threats (APTs), large-scale DDoS attacks, and malicious botnets occur more and more frequently, posing a serious threat to national security, enterprise operations, and personal privacy. Therefore, it is urgent to study an efficient and accurate network intrusion detection model to maintain network security.

[0003] Most existing network intrusion detection methods are based on machine learning and deep learning models. Such methods use deep learning models to analyze network traffic and determine whether the input traffic is malicious or benign based on the learned network traffic characteristics. However, the scale of malicious traffic generated by intrusion behaviors is much smaller than that of normal network traffic. When training an intrusion detection model, this uneven data distribution will cause the model to tend to prefer the normal traffic category with a larger data volume and ignore those malicious traffic categories with a smaller number, which has an adverse effect on the detection accuracy.

[0004] In recent years, many scholars have used sampling techniques to balance the network traffic data distribution. However, although this method can reduce the imbalance rate of network traffic, it reduces the number of network traffic samples, which will lead to a decrease in the detection accuracy of the model. Generative adversarial networks have been widely applied in the field of image generation due to their unique ability to generate new data. In generative adversarial networks, when the generator outputs the probability distribution of all possible class values, the original class values are directly encoded in a single hot vector, making it easier for the discriminator to distinguish real data and synthetic data by comparing the sparsity of their distributions. However, since the network traffic data objects are mostly tabular, the traditional generative adversarial network model is not applicable to network traffic detection.

[0005] Network traffic intrusion detection methods based on ensemble learning have been widely adopted by many scholars. These methods integrate the detection results of multiple base models through strategies such as boosting, bagging, and voting, thereby improving the overall detection performance of the model. Most existing network intrusion detection models based on ensemble learning combine the prediction results of multiple classifiers for intrusion detection, and they can achieve higher detection accuracy than a single classifier; in addition, ensemble learning-based models usually have better robustness and still show good detection performance when facing network traffic with noise or abnormal data. However, the network intrusion detection method based on ensemble learning inevitably requires more time and computing resources to train the model, and there is also a risk of overfitting in the case of insufficient data.

[0006] To solve the problems existing in the process of network intrusion detection, the present invention proposes a network intrusion detection model GMSEL based on generative adversarial and multi-stack ensemble learning. This method consists of a network traffic feature selection module, a data augmentation module, and an intrusion detection module based on ensemble learning, which jointly complete the network intrusion detection task. First, the method iteratively searches for all possible feature combinations and uses the GBDT model to detect network traffic containing the corresponding feature groups, so as to select feature combinations with greater relevance to network intrusion detection from a large number of network traffic features according to the final detection accuracy. Then, by using a data augmentation module based on generative adversarial networks, new network traffic samples are generated to balance the distribution of different categories of network traffic. Finally, through a multi-stack ensemble learning strategy, the advantages of each base model are fully utilized to improve the accuracy and generalization ability of network intrusion detection. The GMSEL model proposed by the present invention is superior to existing network intrusion detection methods based on ensemble learning in terms of accuracy, precision, FPR, and F1-score, and has better stability. Summary of the Invention

[0007] Aiming at the problems existing in the existing network intrusion detection models, such as redundant features of network traffic, inability to effectively solve data imbalance in network traffic, and poor generalization and stability caused by using a single detection model, a method based on generative adversarial networks and multi-stack ensemble learning is proposed to realize network intrusion detection.

[0008] The technical solution of the present invention:

[0009] A network intrusion detection model based on generative adversarial and multi-stack ensemble learning, comprising the following steps:

[0010] Step 1, preprocess the original network traffic data set using an iterative search-based feature selection algorithm to select network traffic features that are more important for the network intrusion detection task, so as to reduce the impact of invalid features on intrusion detection and improve the time efficiency of subsequent data augmentation and intrusion detection;

[0011] Step 1.1, check and eliminate the features containing missing values and the features irrelevant to the network traffic itself in the original network traffic data set, and then use the hot encoding method to encode the character-based discrete features in the network traffic data set and convert them into numerical representations to obtain a new network traffic data set F;

[0012] Step 1.2, perform secondary processing on the network traffic data set F using iterative search; for the features [f1, f2,..., f in the network traffic data set F K, first determine the dimension d of the feature (d ∈ (1, K]), where K is the dimension of the feature; then search for d-dimensional features and group them into feature groups, and the search process continues until all features are searched; finally, use the GBDT model to detect the network traffic data of all feature groups to obtain the average detection accuracy of each feature group; among them, the feature group with the highest average detection accuracy is the optimal feature subset F for constructing the GMSEL model. ′ ;

[0013] Step 2: Use the improved generative adversarial network to learn the distribution of the network traffic features for the network traffic features preprocessed in Step 1, so as to expand the samples of the network traffic features in the small-scale categories, balance the distribution of different categories of network traffic, and further alleviate the problem of insufficient training caused by the small-scale training data set.

[0014] Step 2.1: Data generation: Introduce a strategy that combines a conditional generator with sample-based training in the generator network to improve the quality of the generated samples; the generator network regularly samples all categories in the discrete features during training and restores the distribution of the original network traffic data during the test process, so as to generate the conditional distribution of the original network traffic data according to the hash value; ensure uniform sampling of all categories of discrete features during the training process, so as to generate network traffic samples that satisfy the distribution shown in formula (1); the conditional generator needs to learn the conditional distribution of the network traffic samples before preprocessing, as shown in formula (2).

[0015] r ∼ P g (raw|D i = k) (1)

[0016] P g (raw|D i = k)= P(raw|D i = k) (2)

[0017] Among them, r represents the network traffic sample generated by the generator network, P g represents the conditional distribution of the generated network service samples, P represents the conditional distribution of the original network traffic data samples, Di is the i-th discrete column of the original network traffic samples, k is the hash value of the i-th Di, and raw represents the original network traffic data.

[0018] Step 2.2: Data discrimination: The discriminator network receives the network traffic samples generated by the generator network and the real network traffic samples, and identifies the difference between the generated network service and the real network service by comparison; the discriminator network continuously optimizes to improve its ability to identify and generate network traffic, thereby driving the generator network to generate more real network traffic.

[0019] Step 2.3, Cycle Stability: The generator network and the discriminator network are alternately trained and compete with each other, continuously adjusting their parameters until a dynamic equilibrium state is reached; when the network traffic samples generated by the generator network are indistinguishable from the real network traffic samples in the eyes of the discriminator network, it is considered that the generator network and the discriminator network have reached a stable state, and the training ends here; the specific process is shown in Equation (3).

[0020]

[0021] Among them, L(G,D) represents the cross-entropy loss function, E represents the mathematical expectation, p data represents the distribution of the original network traffic data, p z represents the distribution of the generated network traffic, z represents the input noise of the generator network, G(z) represents the output of the generator network, and D(x) represents that the discriminator network judges the input data x as the original network traffic data.

[0022] Modify the ReLU activation function of the discriminator network and the generator network to the ELU activation function, as shown in Equation (4).

[0023]

[0024] That is, when the input is negative, the ELU produces a non-zero output, avoiding the situation where neurons "die" and are completely inactive.

[0025] Among them, α is a hyperparameter with a value of 1.

[0026] Step 3, Input the enhanced network traffic samples in Step 2 into the network intrusion detection model based on multi-stack ensemble learning for training to obtain more accurate detection results; specifically include: on the publicly available network traffic datasets CIRA, CICID, and NSL-KDD, perform preprocessing according to Step 1, implement data enhancement according to Step 2, and then put the enhanced network traffic data into the network intrusion detection model based on multi-stack ensemble learning to obtain better results in terms of accuracy, precision, F1-score, FPR, and detection stability.

[0027] Step 3.1, First Layer Stacking, Preliminary Detection by the Basic Model: The network traffic dataset X=(x1, x2,..., x m)It is input into five pre-trained base models for preliminary detection, and the output results of each base model are integrated into a feature vector [r1, r2, r3, r4, r5] of dimension 5. This feature vector is fed into the second layer for final classification prediction. Among them, the five pre-trained base models include DT, RF, CatBoost, GBDT, and MLP. The integration process of the outputs of the base models in the first-layer stack is shown in formula (5);

[0028] [r1, r2, r3, r4, r5] = [DT(x j ), RF(x j ), CatBoost(x j ), GBDT(x j ), MLP(x j )](5)

[0029] Among them, x j represents the j-th data in the network traffic data, and r n represents the output of the n-th base model, j ∈ {1, 2, 3, …, m}, n ∈ {1, 2, 3, …, 5};

[0030] Step 3.2, second-layer stacking, meta-model LR classification prediction: The feature vector R = [r1, r2, r3, r4, r5] obtained from the first-layer stack is input into the pre-trained meta-model logistic regression for classification, and the final classification result y is obtained through the output of the meta-model logistic regression, as shown in formula (6);

[0031]

[0032] Among them, w is the training parameter, and T is the transpose of w.

[0033] Advantages of the present invention:

[0034] 1. By introducing an improved GAN network model to achieve data augmentation of network traffic, higher-quality network traffic samples can be generated, which can improve the accuracy of network intrusion detection methods based on machine learning and deep learning models.

[0035] 2. Aiming at the advantages of five machine learning models (including DT, RF, CatBoost, GBDT, and MLP), a multi-stack assembly strategy is used to integrate them. This model can effectively utilize the advantages of the base models and detect malicious traffic more accurately, thus obtaining more effective detection results than using a single base model.

[0036] 3. In the second stacking process of the intrusion detection module, a linear regression model is used, which has a faster calculation speed and higher calculation efficiency. It can quickly find more effective features for classification, further improving the detection accuracy of network intrusion, thus solving the bias and variance phenomena existing in the detection process of the single-machine learning model. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is the overall flowchart of a network intrusion detection model based on generative adversarial and multi-stack ensemble learning.

[0038] Figure 2 are the network intrusion detection results without data augmentation on the CIRA dataset. Among them, (a) is the accuracy comparison result, (b) is the precision comparison result, (c) is the false positive rate comparison result, and (d) is the F1-value comparison result.

[0039] Figure 3 are the network intrusion detection results without data augmentation on the CICID dataset. Among them, (a) is the accuracy comparison result, (b) is the precision comparison result, (c) is the false positive rate comparison result, and (d) is the F1-value comparison result.

[0040] Figure 4 are the network intrusion detection results on the NSL-KDD dataset without data augmentation. Among them, (a) is the accuracy comparison result, (b) is the precision comparison result, (c) is the false positive rate comparison result, and (d) is the F1-value comparison result.

[0041] Figure 5 is the stability of the intrusion detection method based on ensemble learning. Among them, (a) is the accuracy rate ratio result, (b) is the precision comparison result, (c) is the false positive rate comparison result, and (d) is the F1-value comparison result. DETAILED DESCRIPTION OF THE INVENTION

[0042] The following further describes the specific implementation manners of the present invention in conjunction with the drawings and technical solutions.

[0043] The present invention aims to address the problems of poor interpretability, high resource consumption, weak model generalization ability, and data imbalance existing in the existing network intrusion detection technology. A network intrusion detection model GMSEL based on generative adversarial network and multi-stack ensemble learning is proposed. By introducing the generative adversarial network GAN technology, the data imbalance problem is solved, and a multi-stack ensemble learning strategy is innovatively adopted to improve the model generalization ability.

[0044] As Figure 1 shown, the overall process of a network intrusion detection model based on generative adversarial and multi-stack ensemble learning proposed by the present invention includes:

[0045] Step 201, use the feature selection model FS based on iterative search to extract features from the original network traffic dataset to obtain the optimal feature subset.

[0046] Step 2011, check and eliminate the features containing missing values and the features irrelevant to the network traffic itself in the network traffic dataset, and then use the one-hot encoding method to encode the character-based discrete features in the data and convert them into numerical representations, so as to obtain a new network traffic dataset F;

[0047] Step 2012, considering that not all features contribute to the detection performance, use iterative search to perform secondary processing on F. Specifically, for the features [f1, f2,..., f K (where K is the dimension of the features), first determine the network traffic sample feature dimension d (d ∈ (1, K]), and then search for these d-dimensional features and group them into feature groups. This process will continue until all features are searched. Finally, use the GBDT model to detect the network traffic containing this feature group to obtain the average detection accuracy. Among them, the subset with the highest average precision in the feature group is the optimal feature subset F' for constructing the GMSEL model.

[0048] Step 202, use the generative adversarial network model NT-GAN to enhance the network traffic dataset containing features to obtain a new network traffic database, including:

[0049] Step 2021, data generation: Introduce a strategy that combines a conditional generator with sample-based training in the traditional generator network to improve the quality of the generated samples. The generator network regularly samples all network traffic categories in the discrete attributes during training and restores the data distribution of the original network traffic during testing, so as to generate the conditional distribution of the original data according to the specific values of the hash. Uniform sampling is performed on the discrete attributes of all categories during the training process, so as to generate network traffic samples that satisfy the distribution shown in formula (1). The conditional generator needs to learn the conditional distribution of the original network traffic, as shown in formula (2);

[0050] r ∼ P g (raw|D i = k) (1)

[0051] P g (raw|D i = k)=P(raw|D i = k) (2)

[0052] Among them, r represents the network traffic sample generated by the generator, Pg represents the conditional distribution of the generated network data sample, P represents the conditional distribution of the original network traffic, Di is the i-th discrete column of the network traffic, k is the hash value of the i-th Di, and raw represents the original network traffic.

[0053] Step 2022, data discrimination: The discriminator network receives the network traffic samples generated by the generator network and the real network traffic samples, and then identifies the differences between the generated network services and the real network services through comparative analysis. The discriminator network improves its ability to identify and generate network traffic through continuous optimization, thereby driving the generator network to generate more real network traffic;

[0054] Step 2023, cycle stability: The generator network and the discriminator network are alternately trained and compete with each other, continuously adjusting parameters until a dynamic equilibrium state is reached. When the network traffic generated by the generator network is indistinguishable from the real network traffic in the eyes of the discriminator network, it is considered that the generator network and the discriminator network have reached a stable state, and the training ends here. The specific process is shown in formula (3):

[0055]

[0056] Among them, L represents the cross-entropy loss function, E represents the mathematical expectation, P data represents the distribution of the original network traffic, P Z represents the distribution of the generated network traffic, Z represents the input noise of the generator, G(z) represents the output of the generator, and D(x) represents that the discriminator judges the input data x as the original network traffic.

[0057] In order to solve the problem of causing neurons to "die", the present invention modifies the ReLU activation function of the discriminator and the generator to the ELU activation function, as shown in formula (4).

[0058]

[0059] That is, when the input is negative, ELU can produce a non-zero output, avoiding the situation where neurons "die" and are completely inactive.

[0060] Step 203, putting the network traffic data into a network intrusion detection model based on multi-stack ensemble learning to detect malicious traffic, including:

[0061] Step 2031, First - layer stacking, preliminary detection by the base models: Input the network traffic dataset X=(x1, x2,..., xm) containing m samples into 5 pre - trained base models (including DT, RF, CatBoost, GBDT, and MLP) for preliminary detection, and integrate the output results of each base model into a feature vector [r1, r2, r3, r4, r5] with a dimension of 5. This feature vector is fed to the second layer for final classification prediction. The specific process is shown in formula (5), which describes the integration process of the outputs of the base models in the first - layer stacking.

[0062] [r1, r2, r3, r4, r5]=[DT(x j ), RF(x j ), CatBoost(x j ), GBDT(x j ), MLP(x j )](5)

[0063] where, x j represents the i - th data in the network traffic data, and r j represents the output of the j - th base model, j∈{1, 2, 3,…, 5}.

[0064] Specifically, the decision - tree model (DT) divides the network traffic into two subsets at the root node according to the source IP of the network traffic, one from known attackers and the other from unknown attackers; at each subsequent node, the data is further segmented according to the destination IP of the network traffic to identify suspicious network traffic with the destination IP. Preferably, the CART tree and the Gini coefficient (Gini) minimization criterion shown in formula (6) are used to split the features and obtain the classification result of the network traffic.

[0065]

[0066] where, A represents the number of classes of the network traffic samples, and P a represents the probability of each class.

[0067] Specifically, the random forest (RF) constructs each decision tree by randomly selecting a subset of samples and a subset of features. This diversity ensures that the model can capture various patterns and relationships in the network traffic, thereby improving the model's ability to identify complex intrusion patterns.

[0068] The absolute majority voting method shown in formula (7) is used for integration;

[0069]

[0070] where, T(X) represents the set of decision trees, and t b(X) represents a single CART tree, Y represents the output variable, and I(·) represents the indicator function.

[0071] Specifically, the Gradient Boosting Decision Tree (GBDT) constructs a set of decision trees by iteratively improving the prediction results of the model and uses the gradient descent algorithm to minimize the loss function of the model. Formula (8) shows the predicted results of the N fitted trees;

[0072]

[0073] where I(·) is the indicator function, ien is the corresponding leaf node region, J is the number of leaf nodes, and C en is the negative gradient fitting value.

[0074] Specifically, the scale of normal network traffic is often significantly larger than that of malicious traffic, resulting in a serious data imbalance problem. The CatBoost model can better handle this imbalance by adjusting the class weights, thereby improving the detection ability of minority attack samples. To better cope with the network intrusion detection task, the loss function shown in formula (9) is used to measure the difference between the model prediction value and the true value, and then the parameters are adjusted according to the difference;

[0075]

[0076] where z represents the number of network traffic samples, L(y i , F(x i )) represents the training error of the i-th sample, L represents the number of decision trees, and Ω(f l ) is the regularization term of the L-th tree.

[0077] Specifically, the basic structure of the Multi-Layer Perceptron (MLP) includes an input layer, an output layer, and at least one hidden layer. Each layer consists of multiple neurons, and each neuron generates an output by summing the input values with weights and applying an activation function, as shown in formula (10).

[0078] O = (XW h + b h )W o + b o (10)

[0079] where O represents the output of the output layer, h represents the number of hidden units, W h represents the weights of the hidden layer, W o represents the weights of the output layer, and b o represents the bias of the output layer.

[0080] Step 2032, second-layer stacking, meta-model LR classification prediction: Input the feature vector [r1, r2, r3, r4, r5] obtained by the first-layer stacking into the pre-trained meta-model logistic regression (LR) for classification, and obtain the final classification result y through the output of the meta-model LR, as shown in formula (11):

[0081]

[0082] where w is a training parameter.

[0083] The present invention mainly focuses on detecting network intrusion malicious traffic, and proposes a network intrusion detection model based on generative adversarial and multi-stack ensemble learning. Three datasets, namely CIRA, CICID, and NSL-KDD, are selected for testing. Among them, the network traffic in the CIRA dataset has 34-dimensional features, including four categories: non-DoH, DoH, malicious, and benign; the CICID dataset contains benign traffic and the latest common attack traffic, and it contains 78 features; the NSL-KDD dataset is an improved version of the KDD Cup 1999 dataset, containing 41 features and 1 label.

[0084] As can be seen from the experimental results in Table 1, for the three publicly available network traffic datasets CIRA, CICID, and NSL-KDD, the detection results of the GBDT model for the feature combination E selected by the proposed iterative search-based network traffic feature selection method are better than those of feature combinations A / B / C / D. On the CIRA dataset, the detection accuracy of feature combination E is 78.47%, while the detection accuracies of feature combinations A, B, C, and D are only 78.02%, 63.93%, 75.27%, and 76.20% respectively; the precision of the model for feature combination E is 78.59%, while the precisions for feature combinations A, B, C, and D are only 74.29%, 64.45%, 73.03%, and 77.65% respectively. On the CICID dataset, the accuracy of the model for feature combination E is 90.83%, the precision is 91.33%, the FPR is 9.95%, and the F1 score is 89.53%; the model performs second best on feature combination A, with an accuracy of 81.82%, a precision of 66.95%, an FPR of 20.00%, and an F1 score of 73.64%; the model performs worst on feature combination D, with an accuracy of only 73.49%, a precision of 74.17%, an FPR of 16.93%, and an F1 score of 73.27%. On the NSL-KDD dataset, the accuracy of the model for feature combination E is 71.88%, which is approximately 3.60%, 0.69%, 1.40%, and 1.84% higher than the accuracies of feature combinations A, B, C, and D respectively; the FPR of feature combination E is 9.25%, which is approximately 0.04%, 0.16%, 0.60%, and 0.72% lower than the FPRs of feature combinations A, B, C, and D respectively. The experimental results show that the proposed iterative search-based network traffic feature selection method can iteratively search all possible feature combinations, thus avoiding missing some feature combinations and selecting better feature combinations. This significantly improves the effectiveness of network intrusion detection.

[0085] Table 1 Detection Efficiency of Gradient Boosting Decision Tree Model for Different Feature Combinations

[0086]

[0087] From Figure 2 、 3 、4 Experimental results show that after data augmentation using the improved GAN, the detection results of the publicly available network traffic datasets CIRA, CICID, and NSL-KDD are generally higher than those in the network traffic datasets without data augmentation and in the network traffic datasets with data augmentation using the traditional GAN. From Figure 3(a) It can be seen that on the CIRA dataset, after data augmentation using the improved GAN, the accuracy of feature combination a is 78.36%, which is 0.34% higher than the case without data augmentation and 0.21% higher than that after data augmentation using the traditional GAN; the accuracy of the improved GAN for data augmentation of feature combination E is 78.71%, which is 0.24% higher than the case without data augmentation and 0.19% higher than the result of data augmentation using the traditional GAN; as Figure 3 (c) shows, the FPR after data augmentation using the improved GAN on feature combination A is 9.46%, which is 3.00% lower than the FPR without data augmentation on feature combination B; it can be seen from Table 2 that on the CICID dataset, the accuracies of feature combinations A / B / C / D / E without data augmentation are 66.95% / 66.44% / 78.98% / 74.17% / 91.33%, and the accuracies of feature combinations A / B / C / D / E after data augmentation using the improved GAN are 71.91% / 66.48% / 80.11% / 83.82% / 94.75%. After data augmentation using the traditional GAN on feature combination A, the accuracies are 71.91% / 66.48% - 80.11% / 83.75% / 94.75%. The ratios of B / C / D / E are 82.97% / 80.94% / 82.77% / 89.43% / 96.82%; on feature combinations A / B / C / D / E, the F1 scores after data augmentation using the improved GAN are 83.89% / 78.84% / 81.96% / 76.96% / 96.78%, which are approximately 10.25% / 5.63% / 4.43% / 3.69% / 7.25% higher than the F1 scores without data augmentation on feature combinations A / B / C / D / E. Figure 4 It shows that on the NSL-KDD dataset, the detection performance of the network traffic dataset after data augmentation using the improved GAN is also better than that without data augmentation and after data augmentation using the traditional GAN.

[0088] Table 2 Detection efficiency of different integration strategies in network intrusion detection

[0089]

[0090] As can be seen from the experimental results in Table 2, on the three public network traffic datasets, the detection performance of the multi-stack ensemble strategy is better than that of the other five ensemble learning strategies. On the CIRA dataset, the accuracy of the multi-stack ensemble strategy is approximately 82.39%, the precision is approximately 81.15%, the FPR is approximately 6.85%, and the F1-score is approximately 81.75%; on the CICID dataset, the accuracy of the multi-stack ensemble strategy is 99.24%, which is 4.80%, 0.30%, 0.61%, 0.30%, and 0.29% higher than the accuracies of the Bagging, Boosting, Blending, Voting (soft), and Voting (hard) strategies respectively; in terms of precision, it is 99.26%, which is 3.42%, 0.27%, 0.58%, 0.28%, and 0.30% higher than the precisions of the Bagging, Boosting, Blending, Voting (soft), and Voting (hard) strategies respectively; in terms of FPR, it is 0.18%, which is 1.12%, 0.07%, 0.14%, 0.07%, and 0.06% lower than that of the Bagging, Boosting, Blending, Voting (soft), and Voting (hard) strategies; the F1-score is 99.24%, which is 4.95%, 0.29%, 0.61%, 0.30%, and 0.28% higher than the scores of the Bagging, Boosting, Blending, Voting (soft), and Voting (hard) strategies respectively. On the NSL-KDD dataset, the accuracy of the multi-stack ensemble strategy is 5.62% higher than that of the Bagging strategy, 5.29% higher than that of the Boosting strategy, 2.54% lower than that of the blending strategy, and 3.06% higher than that of the Voting (soft) strategy.

[0091] Table 3 shows the experimental comparison results of the multi-stack ensemble learning strategy and the five most commonly used ensemble learning strategies in the field of network intrusion (including Bagging, Boosting, Blending, Voting (soft), and Voting (hard)). It can be seen that on the three public network traffic datasets, the detection performance of the multi-stack ensemble strategy is better than that of the other five ensemble learning strategies. The experimental results show that the multi-stack ensemble learning strategy used can achieve higher accuracy, precision, and F1-score, as well as lower FRP, in the network intrusion detection task, and its stability is also better.

[0092] Table 3 Detection Efficiency of the State-of-the-Art Intrusion Detection Models

[0093]

[0094] Figure 5 Shown on these three network traffic datasets, compared with five ensemble learning-based network intrusion detection models, the GMSEL model proposed by the present invention has the smallest interquartile range and no outliers in four metrics (Accuracy, Precision, FPR, and F1-score), indicating that the GMSEL model has better detection stability. As Figure 5 (a) shows, GMSEL is superior to ADFSE, AIDSEL, EID, HELID, and VGID in terms of precision; among them, the AIDSEL, HELID, and VGID models have poor stability and there are outliers. As Figure 5 (b) and 5(d) show, GMSEL still maintains the highest precision and F1-score, indicating that GMSEL has better performance in network intrusion detection tasks. Among them, the VGID model performs relatively weakly on the NSL-KDD dataset and shows outliers. As Figure 5 (c) shows, the GMSEL model is in the lowest position in terms of FPR, indicating that GMSEL can reduce the possibility of detecting normal network activities as malicious attack behaviors.

Claims

1. A network intrusion detection model based on generative adversarial and multi-stack ensemble learning, characterized in that: The steps include: Step 1, use the iterative search-based feature selection algorithm to preprocess the original network traffic data set to select the network traffic features that are more important for the network intrusion detection task; Step 1.1, check and eliminate the features containing missing values ​​and features irrelevant to the network traffic itself in the original network traffic dataset, and then use hot encoding to encode the character-based discrete features in the network traffic dataset, convert them into numerical representation, and obtain a new network traffic dataset F; Step 1.2, use iterative search to perform secondary processing on the network traffic dataset F; for the features [f1,f2,...,f K ], first determine the dimension d of the feature (d∈(1,K]), K is the dimension of the feature; then search for d-dimensional features and group them into feature groups, and the search process continues until all features are searched; finally, use the GBDT model to detect the network traffic data of all feature groups to obtain the average detection accuracy of each feature group; among them, the feature group with the highest average detection accuracy is the optimal feature subset F for constructing the GMSEL model ′ ; Step 2: Use the improved generative adversarial network to learn the distribution of network traffic features after preprocessing in step 1, so as to expand the samples of network traffic features in small-scale categories; Step 2.1, data generation: Introduce a strategy that combines conditional generator with sample-based training in the generator network; the generator network regularly samples all categories in the discrete features during training, and restores the distribution of the original network traffic data during the test process, thereby generating the conditional distribution of the original network traffic data based on the hash value; ensure that all categories of discrete features are uniformly sampled during the training process, so as to generate network traffic samples that meet the distribution shown in formula (1); the conditional generator needs to learn the conditional distribution of the network traffic samples before preprocessing, as shown in formula (2); r~P g (raw|D i =k) (1) P g (raw|D i =k)=P(raw|D i =k) (2) Among them, r represents the network traffic sample generated by the generator network, P g represents the conditional distribution of the generated network service sample, P represents the conditional distribution of the original network traffic data sample, Di is the i-th discrete column of the original network traffic sample, k is the hash value of the i-th Di, and raw represents the original network traffic data; Step 2.2, data identification: The discriminator network receives the network traffic samples generated by the generator network and the real network traffic samples, and identifies the differences between the generated network traffic and the real network traffic by comparison; the discriminator network continuously optimizes to improve its ability to identify and generate network traffic, thereby driving the generator network to generate more realistic network traffic; Step 2.3, cycle stability: The generator network and the discriminator network are trained alternately, compete with each other, and continuously adjust parameters until a dynamic equilibrium state is reached; when the network traffic samples generated by the generator network are indistinguishable from the real network traffic samples in the eyes of the discriminator network, it is considered that the generator network and the discriminator network have reached a stable state, and the training ends here; the specific process is shown in formula (3); Among them, L(G,D) represents the cross entropy loss function, E represents the mathematical expectation, and p data Represents the distribution of original network traffic data, p z represents the distribution of generated network traffic, z represents the input noise of the generator network, G(z) represents the output of the generator network, and D(x) represents that the discriminator network judges the input data x as the original network traffic data; Change the ReLU activation function of the discriminator network and the generator network to the ELU activation function, as shown in Formula (4); That is, when the input is negative, the ELU produces a non-zero output, avoiding the situation where the neuron "dies" and becomes completely inactive; Among them, α is a hyperparameter with a value of 1; Step 3: Input the enhanced network traffic samples in step 2 into the network intrusion detection model based on multi-stack ensemble learning for training to obtain more accurate detection results; specifically, on the publicly available network traffic datasets CIRA, CICID, and NSL-KDD, perform preprocessing according to step 1, implement data enhancement according to step 2, and then put the enhanced network traffic data into the network intrusion detection model based on multi-stack ensemble learning to obtain better results in accuracy, precision, F1-score, FPR, and detection stability; Step 3.1, first layer stacking, basic model preliminary detection: The network traffic dataset X containing m samples is m ) is input into 5 pre-trained basic models for preliminary detection, and the output results of each basic model are integrated into a feature vector [r1, r2, r3, r4, r5] with a dimension of 5, which is fed to the second layer for final classification prediction; among them, the 5 pre-trained basic models include DT, RF, CatBoost, GBDT and MLP; the integration process of the basic model output in the first layer stack is shown in formula (5); [r1,r2,r3,r4,r5]=[DT(x j ),RF(x j ),CatBoost(x j ),GBDT(x j ),MLP(x j )](5)where x j represents the jth data in the network traffic data, r n represents the output of the nth basic model, j∈{1,2,3,…,m}, n∈{1,2,3,…,5}; Step 3.2, second layer stacking, meta-model LR classification prediction: input the feature vector R = [r1, r2, r3, r4, r5] obtained by the first layer stacking into the pre-trained meta-model logistic regression for classification, and obtain the final classification result y through the output of the meta-model logistic regression, as shown in formula (6); Among them, w is the training parameter and T is the transpose of w.