Network abnormal traffic detection method and system based on ensemble learning
Through the integrated learning method, the problem of high-dimensional feature redundancy and generalization capabilities in traditional network anomaly traffic detection is solved, and efficient and accurate network threat recognition is achieved.
Patent Information
- Application Number
- CN202510409725.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-02
AI Technical Summary
Traditional network anomaly traffic detection methods have problems such as high-dimensional feature redundancy, data sample imbalance and model generalization capabilities, which leads to low detection efficiency and poor accuracy, making it difficult to adapt to dynamic changes in complex network environments.
Using an integrated learning method, key traffic characteristics are screened out by integrating the feature importance analysis, Pearson correlation analysis and mutual information score of extreme gradient enhancement, and a stacked generalization model is constructed. Combined with the Optuna framework to optimize hyperparameters, a high-precision and strong generalization threat identification model is constructed.
It realizes efficient screening and dimensionality reduction, improves the accuracy and generalization capabilities of network threat identification, and can perform accurate abnormal traffic detection in complex network environments, with significantly better performance than traditional algorithms.
Smart Images

Figure CN120263473A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network abnormal traffic detection, and particularly to a network abnormal traffic detection method and system based on ensemble learning. Background Art
[0003] Traditional network security defense methods include: firewalls, intrusion detection systems, vulnerability scanning, etc. A firewall creates isolation between different networks and protects the network by blocking data flows. This method is used to resist conventional external attacks in a simple network architecture, but it is helpless against abnormal traffic inside the network and hidden threats in encrypted information; intrusion detection or prevention systems intercept attacks by analyzing abnormal network traffic behaviors in alarm log records, but this method cannot identify unknown attacks, relies on the update of the rule base, and cannot respond to unknown attacks in a timely manner; the method of performing vulnerability scanning on hosts and network devices through tool software faces problems such as lagging update of the vulnerability feature library, blind spots in traffic detection, and huge time consumption for global scanning.
[0004] Network abnormal traffic detection aims to identify patterns in network traffic that do not conform to expected behaviors, thereby ensuring network security. With the increasing complexity of networks and attack means, traditional signature- or rule-based detection methods have become difficult to cope with new and advanced threats. Therefore, researchers are actively exploring the use of machine learning and deep learning technologies to improve the performance and efficiency of network abnormal detection.
[0005] Although certain achievements have been made in the field of network abnormal traffic detection, with the continuous escalation of network threats in recent years, existing methods still have great room for improvement. In terms of feature extraction, traditional feature extraction methods cannot effectively handle high-dimensional features and redundant information, resulting in a large number of features in the extracted features that contribute little to abnormal detection, and there may be a high degree of correlation between features. These problems increase the computational complexity and may also cause problems such as overfitting of the model and low evaluation accuracy. In addition, the generalization ability of current traffic detection methods is poor, and they can only achieve good performance for specific data sets, lacking effective capture of the potential internal laws of the data and being unable to adapt to the dynamic changes of the network. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the present invention provides a network abnormal traffic detection method and system based on ensemble learning, which solves the technical problems of high-dimensional feature redundancy, data sample imbalance, and weak model generalization ability existing in traditional network abnormal traffic detection methods.
[0007] To solve the above technical problems, the present invention provides the following technical solution: A network abnormal traffic detection method based on ensemble learning, the method includes the following steps:
[0008] By integrating the feature importance analysis of extreme gradient boosting, Pearson correlation analysis, and mutual information scores, the traffic data in the network traffic data is screened and dimensionally reduced to form a final traffic feature set;
[0009] Using the extreme gradient boosting algorithm, categorical gradient boosting algorithm, and light gradient boosting machine as base learners, and the random forest algorithm as a meta-learner, and introducing the Optuna optimization framework for quickly finding the best hyperparameters in the high-dimensional hyperparameter space, a stacking generalization model for identifying abnormal traffic data in the network security threat situation recognition scenario is constructed;
[0010] The stacking generalization model is used to detect abnormal traffic data in the traffic feature set, identify threats, and thus determine whether the traffic data is abnormal.
[0011] Through the XPM feature selection method that integrates the feature importance analysis of extreme gradient boosting (XGBoost), Pearson correlation analysis, and mutual information score, multi-dimensional screening and dimensional reduction are performed on the network traffic data, redundant features are removed, and key threat-related features are retained.
[0012] Furthermore, based on the improved stacking generalization model (OPStacking), using the extreme gradient boosting algorithm (XGBoost), categorical gradient boosting algorithm (CatBoost), and light gradient boosting machine (LightGBM) as base learners, and random forest as a meta-learner, combined with the Optuna framework to automatically optimize the model hyperparameters, a threat recognition model with high precision and strong generalization ability is constructed; through the coordinated operation of the traffic data preprocessing module, feature selection module, and threat recognition module, the efficient detection and classification of abnormal traffic in a complex network environment are realized.
[0013] Furthermore, the specific process of forming the final traffic feature set includes the following steps:
[0014] Perform data preprocessing on the network traffic data F = {f1, f2,..., f n} to form multiple traffic feature subsets F1. The data preprocessing includes invalid feature removal, feature one-hot encoding, feature standardization, and label numericalization;
[0015] Based on the XGBoost feature screening model, quantify the feature importance of the traffic data in the traffic feature subset F1, arrange them in descending order, and retain the traffic features with high feature importance scores according to the preset standard to form a preliminarily screened traffic feature subset F2;
[0016] Perform Pearson correlation analysis on the traffic data in the traffic feature subset F2, set the Pearson traffic feature correlation threshold θ2, and at the same time calculate the mutual information score of the traffic features in the traffic feature subset F2;
[0017] If there are traffic features exceeding the Pearson traffic feature correlation threshold θ2, then compare the mutual information scores and eliminate the traffic features with low scores to form the final traffic feature set F3;
[0018] If no traffic feature exceeds the set Pearson traffic feature correlation threshold θ2, then directly form the traffic feature set F3. For the traffic feature f j There is:
[0019]
[0020] In the formula, |r jk | is the absolute value of the Pearson correlation coefficient r jk .
[0021] Furthermore, the formation of the preliminary screened traffic feature subset F2 is as follows:
[0022] Screen out the traffic feature subset F1 whose feature importance is higher than the XGBoost traffic feature importance threshold θ1. The expression is:
[0023] F2 = {f j ∈F | S XGB (f j ) > θ1}
[0024] In the formula, f j is the j-th traffic feature in the traffic feature subset F1; S XGB (f j ) is the XGBoost importance score of the traffic feature f j .
[0025] Furthermore, the performing of Pearson correlation analysis on the traffic data in the traffic feature subset F2 is as follows:
[0026] For any two traffic features f j and f k in the traffic feature subset F2, if the absolute value of their Pearson correlation coefficient r jk exceeds the Pearson traffic feature correlation threshold θ2, then retain the traffic feature with a higher mutual information score. The expression is:
[0027]
[0028] In the formula, I(f j, Y) is the traffic feature f j The mutual information score with the target variable Y; I(f k , Y) is the traffic feature f k The mutual information score with the target variable Y.
[0029] Furthermore, the formation of the final traffic feature set F3 is as follows:
[0030] For the filtered traffic feature subset F2, retain the traffic features with mutual information scores higher than the traffic feature mutual information score threshold θ3. The expression is:
[0031] F3 = {f j ∈ F2 | I(f j ; Y) > θ3}
[0032] In the formula, I(f j , Y) is the mutual information score of the traffic feature f j with the target variable Y.
[0033] Furthermore, the use of the stacked generalization model to detect abnormal traffic data in the traffic feature set, identify threats and thus judge whether the traffic data is abnormal. The specific process includes the following steps:
[0034] Let the optimization objective of the traffic feature set input into the threat recognition module be to maximize the comprehensive evaluation index, that is:
[0035] S = ω1Accuacy + ω2Precision + ω3Recall + ω4F1
[0036] In the formula, x i represents the network traffic data feature vector; y i is the actual threat label of the i-th sample; Accuracy, Precision, Recall, and F1 represent accuracy, precision, recall, and F1 score respectively; ω1, ω2, ω3, ω4 are the index weights, and ω1 + ω2 + ω3 + ω4 = 1;
[0037] Input the traffic feature set F3 into the base learner L j , and the base learner L1 iteratively constructs multiple decision trees to minimize the loss function The expression is:
[0038]
[0039] In the formula, l(·) is calculated using the logarithmic loss method; y is the actual threat label; is the predicted threat label; is the predicted threat label of the i-th sample; n is the total number of samples, that is, the number of samples included in the traffic feature set F3;
[0040] The parameters of the tree model are obtained by calculating the gradient of the loss function to get the predicted output That is:
[0041]
[0042] In the formula, η is the learning rate; f k is a single regression tree; k ∈ K is the number of iterations;
[0043] The base learner L2 is trained using a symmetric decision tree and its predicted output is obtained after K iterations That is:
[0044]
[0045] In the formula, a(x i , T) is the probability prediction value output after the sample i falls into the leaf node of the decision tree T;
[0046] For the base learner L3, its predicted output is:
[0047]
[0048] In the formula, ε is the learning rate of the base learner L3; h k (x i ) represents the predicted value of the k-th decision tree for the sample i;
[0049] The predicted results of the base learners are used as new features and merged with the original network traffic data feature vector x i to form a new network traffic feature vector
[0050] The new network traffic feature vector z i and the corresponding threat label y i are used to train the meta-learner M, and its loss function is:
[0051]
[0052] In the formula, l'(·) uses the logarithmic loss calculation method;
[0053] A surrogate model is constructed using a Gaussian process to approximate the objective function S. Based on the existing experimental result set to estimate the value of S under different hyperparameter combinations, and a new hyperparameter combination θ new is calculated through the expected improvement acquisition function EI(θ), and the expression is:
[0054]
[0055] Wherein, S best is the optimal optimization objective; θ (k) is the hyperparameter combination of the k-th experiment; S (k) is the objective function value of the k-th experiment; t is the number of experiments included in the experimental result set;
[0056] Save the hyperparameter combination θ new obtained from the experiment and the optimal optimization objective S best into the result set;
[0057] Repeat the experiment until the preset number of trials to obtain the final optimal hyperparameter combination θ * , and the stacked generalization model corresponding to the optimal hyperparameter combination θ * is the optimal training model. Use this model to predict the probability that sample i belongs to class c
[0058] After calculating the probability distribution of each class, the final predicted class is:
[0059]
[0060] Wherein, argmax is a function to calculate the c value when the probability is the largest, and is used to determine the most likely threat type of the traffic sample i.
[0061] Furthermore, the expression for predicting the probability that sample i belongs to class c using the stacked generalization model is:
[0062]
[0063] Wherein, f M,c (z i ; θ M ) is the prediction score of the meta-learner M for sample i on class c, and there are a total of C classes; f M,k (z i ; θ M ) is the prediction score of the meta-learner M for sample i on class k; exp(·) is the natural exponential function.
[0064] A system for implementing the above network abnormal traffic detection method includes:
[0065] A traffic data preprocessing module that preprocesses network traffic data, including data cleaning, feature pre-screening, encoding conversion, and standardization operations;
[0066] A feature selection module analyzes the importance, linear correlation, and non - linear dependence of traffic features based on feature importance analysis, Pearson correlation analysis, and mutual information scores, and selects the most valuable features to form a traffic feature set.
[0067] A threat recognition module fuses the Optuna optimization framework and the Stacking ensemble learning algorithm to detect abnormal traffic data in the traffic feature set, identify threats, and thus determine whether the traffic is abnormal.
[0068] By means of the above - mentioned technical solutions, the present invention provides a network abnormal traffic detection method and system based on ensemble learning, which at least has the following beneficial effects:
[0069] 1. In view of the problem of the explosion of feature dimensions of network security situation elements, the present invention proposes an XPM network traffic data feature selection method that combines extreme gradient boosting feature importance analysis, Pearson correlation analysis, and mutual information scores, efficiently selects important features, and realizes the automatic dimensionality reduction of high - dimensional data.
[0070] 2. In view of the problems of low efficiency, poor accuracy, and high requirements for the quality of training data in traditional network traffic anomaly recognition methods, the present invention proposes an improved Stacking threat recognition method that combines the Stacking ensemble algorithm and Optuna hyperparameter optimization. This method has high recognition efficiency, good accuracy, strong generalization ability, and is applicable to multi - sample data.
[0071] 3. The network abnormal traffic detection method proposed by the present invention has threat recognition accuracies of 83.26% and 99.96% on the UNSW - NB15 and CIC - IDS2017 data sets respectively. Its performance is significantly improved compared with traditional algorithms, and it can perform accurate abnormal traffic detection, providing a reliable reference for security decision - making in complex network environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The drawings described herein are used to provide a further understanding of the present application, form a part of the present application, and the schematic embodiments and descriptions thereof are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0073] Figure 1 is the flowchart of the network abnormal traffic detection method in the present invention;
[0074] Figure 2 is the flowchart of the network traffic feature selection method in the present invention;
[0075] Figure 3 is the network structure diagram of the stacking generalization model in the present invention;
[0076] Figure 4It is the curve graph of the hyperparameter optimization convergence of the present invention on UNSW-NB15;
[0077] Figure 5 It is the curve graph of the hyperparameter optimization convergence of the present invention on CIC-IDS2017;
[0078] Figure 6 It is the radar graph of the performance of the stacking generalization model and various machine learning algorithms of the present invention on UNSW-NB15;
[0079] Figure 7 It is the radar graph of the performance of the stacking generalization model and various machine learning algorithms of the present invention on CIC-IDS2017;
[0080] Figure 8 It is the bar graph of the influence degree ranking of the ablation module of the present invention. Detailed implementation manners
[0081] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Thereby, the implementation process of how the present application uses technical means to solve technical problems and achieve technical effects can be fully understood and implemented accordingly.
[0082] With the increasing complexity of the types and scales of network threats, traditional network abnormal traffic detection methods face severe challenges in the problems of high-dimensional feature redundancy, data sample imbalance, and weak model generalization ability. In view of the above problems, this embodiment proposes a network abnormal traffic detection method based on ensemble learning. By integrating the feature importance analysis of extreme gradient boosting, Pearson correlation analysis, and mutual information score methods, efficient screening and dimensionality reduction of network traffic features are realized. And using the extreme gradient boosting algorithm, categorical gradient boosting algorithm, and light gradient boosting machine as base learners, and the random forest as a meta-learner to construct a stacking generalization model, and introducing the Optuna framework to efficiently find the best hyperparameters of the model, so as to achieve high-precision identification of network threats. As Figure 1 shown, this method includes the following steps:
[0083] This embodiment proposes a network traffic feature selection method (XPM) that integrates XGBoost feature importance analysis, Pearson correlation analysis, and MI mutual information score, aiming to screen and reduce the dimensionality of traffic features in multiple dimensions, extract the traffic features that are most valuable for abnormal traffic identification, and improve the efficiency and accuracy of abnormal traffic data identification in network traffic data. Specifically: by integrating the feature importance analysis of extreme gradient boosting, Pearson correlation analysis, and mutual information score, the traffic data in the network traffic data is screened and the dimensionality is reduced to form a final traffic feature set.
[0084] The key link in detecting abnormal network traffic is to extract the most valuable features from complex and high-dimensional network data. When traditional traffic feature extraction methods face high-dimensional features, they show problems such as high redundancy in feature extraction, complex calculations, and no consideration of the correlation and dependence between features.
[0085] In this embodiment, the XGBoost algorithm is an ensemble learning algorithm based on gradient-boosted trees. It iteratively optimizes the loss function by constructing multiple decision trees and is good at efficiently modeling complex and high-dimensional traffic data. The contribution degree of each threat feature to the model performance is quantified by the gain when each decision tree node splits. Here, the gain represents the degree of optimization of the loss function after feature splitting. Finally, the feature importance score is obtained by weighted summation of the total gain, and the calculation formula is:
[0086]
[0087] In the formula, g i and h i are the first-order and second-order derivatives of the loss function respectively; L and R are the left and right child nodes after splitting; P is the parent node; λ is the regularization coefficient.
[0088] The Pearson correlation coefficient is used to analyze the linear correlation between two traffic features. Its value range is [-1, 1]. The closer the coefficient is to the boundary of the value range, the stronger the correlation between the features. By setting a coefficient threshold to judge whether the traffic features are highly redundant, pairs of features with high correlation are screened out. The calculation formula is:
[0089]
[0090] In the formula, x i and y i are feature values; and are feature means.
[0091] MI mutual information is used to evaluate the non-linear statistical dependence relationship between traffic features and the target variable and is good at capturing complex dependence relationships. By calculating the mutual information score between traffic data features and threat labels and combining the Pearson correlation coefficient r xy the most valuable features are preferably selected. The calculation formula for the mutual information score I(X; Y) is:
[0092]
[0093] In the formula, p(x, y) is the joint probability distribution; p(x) and p(y) are the marginal probability distributions.
[0094] The specific process of forming the final traffic feature set is asFigure 2 As shown, it includes the following steps:
[0095] Perform data preprocessing on the network traffic data F = {f1, f2,..., f n} to form multiple traffic feature subsets F1. The data preprocessing includes invalid feature removal, feature one-hot encoding, feature standardization, and label numericalization, so as to ensure that the traffic data can be effectively used by the feature extraction module.
[0096] Based on the XGBoost feature screening model, quantify the feature importance of the traffic data in the traffic feature subset F1, arrange them in descending order, and retain the traffic features with high feature importance scores according to the preset standard to form a preliminarily screened traffic feature subset F2; that is, screen out the traffic feature subset F1 whose feature importance is higher than the XGBoost traffic feature importance threshold θ1, and the expression is:
[0097] F2 = {f j ∈F|S XGB (f j ) > θ1}
[0098] In the formula, f j is the j-th traffic feature in the traffic feature subset F1; S XGB (f j ) is the XGBoost importance score of the traffic feature f j .
[0099] Perform Pearson correlation analysis on the traffic data in the traffic feature subset F2, construct a heat map of the correlation coefficient matrix and set the Pearson traffic feature correlation threshold θ2, and at the same time calculate the mutual information score of the traffic features in the traffic feature subset F2; that is, for any two traffic features f j and f k in the traffic feature subset F2, if the absolute value of its Pearson correlation coefficient r jk exceeds the Pearson traffic feature correlation threshold θ2, then retain the traffic feature with a higher mutual information score, and the expression is:
[0100]
[0101] In the formula, I(f j , Y) is the mutual information score of the traffic feature f j and the target variable Y; I(f k , Y) is the mutual information score of the traffic feature f k and the target variable Y.
[0102] If there are traffic features whose Pearson traffic feature correlation threshold θ2 is exceeded, then the mutual information scores are compared, and the traffic features with low scores are removed to form the final traffic feature set F3. That is, for the filtered traffic feature subset F2, the traffic features with mutual information scores higher than the traffic feature mutual information score threshold θ3 are retained. The expression is:
[0103] F3 = {f j ∈ F2 | I(f j ; Y) > θ3}
[0104] If no traffic feature exceeds the set Pearson traffic feature correlation threshold θ2, then the traffic feature set F3 is directly formed. That is, for the traffic feature f j there is:
[0105]
[0106] In the formula, |r jk | is the absolute value of the Pearson correlation coefficient r jk
[0107] In this embodiment, through phased feature screening, potential linear and non-linear relationships in traffic features can be identified, which not only retains features with high contribution degrees but also avoids information overlap and noise interference. The traffic feature set extracted by this method has good efficiency and robustness.
[0108] In the prior art, Stacking is an ensemble learning strategy. Its principle is to construct multiple base learners to train data, and use the outputs of the base learners as new features to input into the meta-learner for re-learning, aiming to improve the overall performance of the model by integrating the advantages of multiple base learners. In the scenario of network security threat situation recognition, the Stacking method uses multiple different base learners to learn threat features in network traffic. However, in the traditional Stacking method, the setting of hyperparameters of the base and meta-learners usually relies on experience and it is difficult to achieve the optimal configuration, resulting in limitations in the efficiency and accuracy of threat recognition. Moreover, when facing complex network traffic data with high-dimensional features, the generalization ability of the model is insufficient and overfitting is likely to occur.
[0109] To address the above problems, this embodiment proposes an improved stacked generalization model based on Optuna hyperparameter optimization: OPStacking. In this embodiment, the Extreme Gradient Boosting algorithm, the Categorical Gradient Boosting algorithm, and the LightGBM are used as base learners, and the Random Forest algorithm is used as the meta-learner. An Optuna optimization framework for quickly finding the best hyperparameters in the high-dimensional hyperparameter space is introduced to construct a stacked generalization model for abnormal traffic data recognition in the network security threat situation recognition scenario, thereby achieving efficient and accurate network threat recognition. The model structure is as shown in Figure 3 shown. In this embodiment, XGBoost, CatBoost, and LightGBM are selected as base learners, and Random Forest is used as the meta-learner.
[0110] XGBoost has powerful parallel computing and modeling capabilities and can efficiently capture the complex non-linear relationship between abnormal traffic and threats when processing large-scale network traffic data. CatBoost is a machine learning algorithm based on symmetric decision trees, which is good at dealing with classification problems, can automatically process categorical features, and has good robustness to overfitting. LightGBM is a distributed gradient boosting framework that uses optimization methods such as the histogram algorithm, can quickly process feature data, and further improves the threat recognition efficiency and performance based on network traffic through techniques such as gradient unilateral sampling and exclusive feature bundling, saving a large amount of memory resources and computing time. Random Forest is a classic ensemble learning algorithm that generates a differentiated training set for each tree through Bootstrap resampling. At the same time, only the optimal threat features are selected from the random subset when splitting nodes, and it can automatically adjust the class weights to eliminate the influence of data imbalance and better utilize the prediction data output by the base learners for learning. Optuna is a hyperparameter optimization framework based on Sequential Model-Based Optimization (SMBO), improved from Bayesian optimization, with intelligent sampling strategies and efficient search algorithms, and can quickly find the parameters most suitable for the model in the high-dimensional hyperparameter space.
[0111] In this embodiment, a stacked generalization model is used to detect abnormal traffic data in the traffic feature set, identify threats, and thus determine whether the traffic data is abnormal. Among them, the base learners L1, L2, and L3 are XGBoost, CatBoost, and LightGBM respectively, and the meta-learner M is RandomForest. For the base learner L j , its hyperparameter set is The hyperparameter set of the meta-learner M is θ M , and the hyperparameter space optimized by Optuna is denoted as The specific process of the corresponding OPStacking threat recognition method is as follows:
[0112] Let the optimization objective of the traffic feature set in the input threat recognition module be to maximize the comprehensive evaluation index, that is:
[0113] S = ω1Accuracy + ω2Precision + ω3Recall + ω4F1
[0114] In the formula, x i represents the network traffic data feature vector; y i is the actual threat label of the i-th sample; Accuracy, Precision, Recall, and F1 represent accuracy, precision, recall, and F1 score respectively; ω1, ω2, ω3, and ω4 are the index weights, and ω1 + ω2 + ω3 + ω4 = 1.
[0115] Input the traffic feature set F3 into the base learner L j , and the base learner L1 iteratively constructs multiple decision trees to minimize the loss function The expression is:
[0116]
[0117] In the formula, l(·) uses the logarithmic loss calculation method; y is the actual threat label; is the predicted threat label; is the predicted threat label of the i-th sample; n is the total number of samples, that is, the number of samples contained in the traffic feature set F3.
[0118] Obtain the parameters of the tree model by calculating the gradient of the loss function, so as to obtain the predicted output That is:
[0119]
[0120] In the formula, η is the learning rate; f k is a single regression tree; k ∈ K is the number of iterations.
[0121] The base learner L2 is trained using a symmetric decision tree, and its predicted output is obtained after K iterations, that is:
[0122]
[0123] In the formula, a(x i ,T) is the probability prediction value output after the sample i falls into the leaf node of the decision tree T. Similarly, for the base learner L3, its predicted output is:
[0124]
[0125] where ε is the learning rate of the base learner L3; h k (x i ) represents the predicted value of the k-th decision tree for sample i.
[0126] Use the prediction results of the base learners as new features, and combine them with the original network traffic data feature vector x i to form a new network traffic feature vector
[0127] Use the new network traffic feature vector z i and the corresponding threat label y i to train the meta-learner M, and its loss function is:
[0128]
[0129] where l'(·) uses the logarithmic loss calculation method.
[0130] Construct a surrogate model using Gaussian processes to approximate the objective function S, and estimate the S value under different hyperparameter combinations according to the existing experimental result set to calculate the new hyperparameter combination θ new through the expected improvement acquisition function EI(θ), and the expression is:
[0131]
[0132] where S best is the best optimization objective; θ (k) is the hyperparameter combination of the k-th experiment; S (k) is the objective function value of the k-th experiment; the superscript t is the number of experiments included in the experimental result set.
[0133] Save the obtained hyperparameter combination θ new and the best optimization objective S best to the result set.
[0134] Repeat the experiment until the preset number of trials to obtain the final best hyperparameter combination θ * θ * The corresponding OPStacking model is the optimal training model, and the probability that the sample i belongs to the class c is predicted using this model The expression is:
[0135]
[0136] where f M,c (z i ; θ M) is the prediction score of the meta-learner M for sample i in class c, and there are C classes in total; f M,k (z i ; θ M ) is the prediction score of the meta-learner M for sample i in class k; exp(·) is the natural exponential function;
[0137] After calculating the probability distribution of each class, the final predicted class is:
[0138]
[0139] In the formula, argmax is a function that calculates the c value when the probability is the largest, and is used to determine the most likely threat type of traffic sample i.
[0140] In this embodiment, by integrating the feature importance analysis of extreme gradient boosting, Pearson correlation analysis, and mutual information score, the efficient screening and dimensionality reduction of network traffic features are realized. And using the extreme gradient boosting algorithm, category gradient boosting algorithm, and light gradient boosting machine as the base learners, and the random forest as the meta-learner to construct a stacking generalization model, and introducing the Optuna framework to efficiently find the best hyperparameters of the model, so as to achieve high-precision identification of network threats.
[0141] This embodiment also proposes a network abnormal traffic detection system based on ensemble learning. This system consists of three parts: a traffic data preprocessing module, a feature selection module, and a threat identification module. Aiming at the problem that the existing network abnormal traffic detection methods are difficult to efficiently and accurately capture abnormal traffic features, it can effectively screen high-dimensional network traffic features and effectively identify the traffic samples of the network system. As Figure 1 shown, the functions of each module are as follows:
[0142] The traffic data preprocessing module preprocesses the network traffic data, including operations such as data cleaning, feature pre-screening, encoding conversion, and standardization.
[0143] The feature selection module analyzes the importance, linear correlation, and non-linear dependence of traffic features based on feature importance analysis, Pearson correlation analysis, and mutual information score, and screens out the most valuable features to form a traffic feature set.
[0144] The threat identification module integrates the Optuna optimization framework and the Stacking ensemble learning algorithm to detect the abnormal traffic data in the traffic feature set, identify threats, and thus judge whether the traffic is abnormal.
[0145] In this embodiment, through experiments on two datasets, the XPM feature selection method successfully reduced the 48-dimensional traffic features of the UNSW-NB15 dataset to 32, and the 78 features of the CIC-IDS2017 dataset to 40, greatly reducing the redundancy of high-dimensional network traffic data features and significantly saving the computational cost of the abnormal traffic recognition model. The threat recognition accuracy of the OPStacking threat recognition method on the UNSW-NB15 and CIC-IDS2017 datasets reached 83.26% and 99.96% respectively, and other performance indicators were also significantly improved compared with traditional algorithms, enabling accurate detection of abnormal traffic. The specific experimental process and result analysis are as follows:
[0146] 1.1 Experimental Environment Configuration and Datasets
[0147] This experiment is based on the Windows 11 operating system, with the basic language interpreter being Python 3.12, the integrated development platform being the professional version of PyCharm 2023.2.1, and the training and evaluation of machine learning models mainly based on the scikit-learn 1.5.2 framework, using CUDA118 for auxiliary computing to save the time cost of model training. The specific environment information is shown in Table 1.
[0148] Table 1 Experimental Environment Configuration
[0149]
[0150] The experimental datasets used are the UNSW-NB15 dataset and the CIC-IDS2017 dataset. The UNSW-NB15 dataset is user network traffic data created by the University of New South Wales in 2015 through software such as IXIA PerfectStorm and Metasploit to simulate a real network environment.
[0151] In this experiment, the official "UNSW_NB15_training-set.csv" and "UNSW_NB15_testing-set.csv" files were used, which contain a total of 48 network traffic features, 10 types of sample labels, and 257,673 samples. The number of samples corresponding to each label is shown in Table 2. After merging and preprocessing them, the training set and test set were re-divided. This dataset overcomes the limitations of traditional datasets such as the outdated threat types of KDD99 and the insufficient complexity of network topologies, and is widely used in the field of network anomaly detection.
[0152] Table 2 Information of UNSW-NB15 Experimental Dataset
[0153]
[0154] The CIC-IDS2017 dataset is a network traffic dataset created by the Canadian Institute of Cybersecurity in 2017 by simulating a real network environment with various tools. This dataset collected traffic for five days from Monday to Friday and gathered more than 8 million network connection records. The dataset obtained rich features by collecting network traffic, and the labels focused on threats such as web attacks, penetration attacks, and denial of service, which could better verify the generalization ability of the model. For this experiment, the data file "Wednesday-workingHours.pcap_ISCX.csv" provided officially for machine learning model training was selected. It contains a total of 78 network features, 4 types of labels, and 692,703 samples. The specific information is shown in Table 3.
[0155] Table 3 Information of the Wednesday-workingHours Experimental Dataset
[0156]
[0157] 1.2, Performance Evaluation Metrics
[0158] In this embodiment, the accuracy rate A, precision rate P, recall rate R, and F1 score are used to evaluate the performance of the XPM-OPStacking model on two datasets. The accuracy rate represents the proportion of correctly predicted threats overall; the precision rate represents the proportion of true positive samples among the samples predicted as positive threats; the recall rate represents the proportion of correctly detected actual positive samples; the F1 score represents the harmonic mean of the precision rate and the recall rate. Through these four metrics, a comprehensive, objective, and comprehensive evaluation of the performance of the XPM-OPStacking model in identifying anomalies based on network traffic data can be achieved.
[0159] 1.3, Data Preprocessing and Feature Engineering
[0160] (1) Data Preprocessing
[0161] For the UNSW-NB15 dataset, first merge the original datasets into a total dataset, delete the id and label columns that are not involved in training, and at the same time eliminate the following features: srcip, sport, dstip, dsport, stime, ltime, ct_flw_http_mthd, is_ftp_login, ct_ftp_cmd, proto, and service. The first six features represent the source IP address, source port number, destination IP address, destination port number, sample start time, and sample end time. These information are of no value for the training of the XPM-OPStacking model. The latter five features have a large number of missing values, null values, and non-numerical features. If one-hot encoding is performed on them, the situation of dimensional explosion will be faced, so they are discarded. Perform one-hot encoding on the non-numerical state column to convert it into a numerical feature, then merge the encoded data with the original features, and convert the attack_cat column into a numerical value. The converted dataset has 47-dimensional features. Finally, standardize the features and divide the training set and test set according to the ratio of 8:2. For the CIC-IDS2017 dataset, since it has outliers and missing values, it is processed. Replace infinite values with missing values, fill the missing values with the median, and divide the training set and test set according to the ratio of 8:2.
[0162] (2) Feature engineering
[0163] Use the XPM method to perform feature selection on the two datasets. For the UNSW-NB15 dataset, use the XGBoost algorithm to screen the feature importance: remove the features that contribute less to the prediction results and may introduce noise. At the same time, restricting the number of features is beneficial to significantly reducing the computational burden and accelerating the model training speed. After considering the above factors, eliminate the features with an importance score of 0. Calculate the correlation and mutual information scores of the remaining features. The closer the correlation coefficient is to 1, the higher the correlation between the two features. The higher the mutual information score, the greater the dependence relationship between the feature and the label.
[0164] Remove the feature with a lower mutual information score while the absolute value of the correlation coefficient between each pair of features is greater than 0.95: dloss, sloss, state_FIN, ct_dst_ltm, ct_srv_src, spkts, dpkts, ct_dst_src_ltm. The finally retained features are: dur, sbytes, dbytes, rate, sttl, dttl, sload, dload, sinpkt, dinpkt, sjit, djit, swin, stcpb, dtcpb, tcprtt, synack, ackdat, smean, dmean, trans_depth, response_body_len, ct_state_ttl, ct_src_dport_ltm, ct_dst_sport_ltm, ct_src_ltm, ct_srv_dst, is_sm_ips_ports, state_CON, state_ECO, state_INT, state_RST, a total of 32 features.
[0165] Similarly, for the CIC-IDS2017 dataset, using the XGBoost algorithm to screen for feature importance, the ten features with a score of 0, namely Fwd Avg Packets / Bulk, CWE Flag Count, ECE Flag Count, FwdAvg Bytes / Bulk, BwdAvg Bulk Rate, RST Flag Count, Bwd PSH Flags, Fwd URG Flags, Bwd URG Flags, BwdAvg Bytes / Bulk, Bwd Avg Packets / Bulk, FwdAvg Bulk Rate, are removed. From the remaining features, features with an absolute correlation coefficient greater than 0.95 and a lower mutual information score in each pair of features are removed. The removed features are SubflowFwdPackets, Fwd Packets / s, Packet Length Std, act_data_pkt_fwd, Fwd IAT Total, BwdHeader Length, Subflow Bwd Packets, Fwd Packet Length Mean, Flow IAT Std, TotalBackward Packets, Fwd Header Length, Idle Max, Fwd IAT Std, Fwd Packet LengthStd, Average Packet Size, Bwd Packet Length Max, Total Length ofFwd Packets, SYNFlag Count, Idle Min, Total Fwd Packets, Fwd IAT Max, Idle Mean, Subflow BwdBytes, Avg Bwd Segment Size, Fwd Header Length.1, Bwd Packet Length Std, and a total of 40 features remaining in the official dataset are retained.
[0166] After the above feature engineering, the 48 features of the UNSW-NB15 dataset are successfully reduced to 32, and the 78 features of the CIC-IDS2017 dataset are reduced to 40, greatly reducing the redundancy of high-dimensional network traffic data features and significantly saving the computational cost of the abnormal traffic recognition model.
[0167] 1.4. Analysis of Experimental Results
[0168] (1) Threat Recognition Results
[0169] The data after feature extraction is input into the OPStacking threat recognition module for training. The Optuna framework is used to optimize the hyperparameters of the stacking model. The CATBoost algorithm comes with regularization methods for ordered boosting and adversarial validation. To prevent overfitting, the default sample sampling ratio and feature sampling ratio are used. Keeping the parameters of the meta-learner relatively simple is beneficial to the overall stability and generalizability of the model architecture. Therefore, the scope of parameter optimization targets the base learners of the model. The hyperparameter configurations obtained through parameter optimization are shown in Table 4. The number of optimization rounds for UNSW-NB15 and CIC-IDS2017 is set to 30. The Stacking model is trained using the optimized hyperparameters. The hyperparameter optimization convergence curves on the two datasets are as Figure 4 , 5 shows.
[0170] Table 4 Hyperparameter configurations optimized by Optuna
[0171]
[0172]
[0173] It can be seen that the optimization process on the UNSW-NB15 dataset is not monotonically increasing. The fluctuations in the blue broken line indicate that during the process of finding the optimal solution, different trial parameters may lead to large fluctuations in the target value. This is because the optimization problem is relatively complex and there are multiple local optimal solutions. The red broken line finally tends to be stable, indicating that as the number of trials increases, the algorithm gradually finds a relatively stable optimal solution. In the CIC-IDS2017 dataset, the convergence curves of the first few experiments fluctuate violently, but reach a relatively high level and tend to be stable after 10 optimizations, indicating that a relatively good combination of hyperparameters is quickly found during the optimization process. The convergence curve of the model reflects the effectiveness of the Optuna optimization framework in finding the optimal hyperparameters for this threat recognition model.
[0174] To verify the effectiveness of the OPStacking threat recognition model in threat recognition, it is compared with multiple machine learning algorithms. The datasets for training and validation of all algorithms are preprocessed in the same way. The recognition results are shown in Table 5.
[0175] Table 5 Comparison of threat recognition performance of different algorithms
[0176]
[0177]
[0178] The performance of the model on the two datasets is presented in the form of a radar chart, as Figure 6 and Figure 7As shown in the figure, from the results of the performance comparison, it can be seen that the OPStacking method performs better than the sum of other algorithms in identifying threat types on both datasets. On the UNSW-NB15 dataset, the recognition accuracy reaches 83.26%, which is 10 percentage points higher than the logistic regression algorithm and also improved compared to other algorithms. The performance of other metrics is also comprehensively better than that of traditional machine learning algorithms. On the CIC-IDS2017 dataset, the accuracy of the OPStacking model reaches 99.96%, and other performance metrics are all above 99%, indicating that the model has strong generalization ability. This is because the XPM feature selection method is used to effectively eliminate high-dimensional redundant features while retaining key threat-related features, reducing noise interference and the risk of overfitting. At the same time, XGBoost, CatBoost, and LightGBM are used as the base learners, effectively complementing the model's ability to capture linear and non-linear complex traffic patterns. The introduced Optuna framework automatically searches for the optimal hyperparameter combination, enabling the threat recognition model to achieve the best performance balance in complex network traffic data.
[0179] (2) Ablation experiment
[0180] To verify the impact of each module in the constructed threat recognition model on the overall performance of the model, an ablation experiment was designed. Using the Stacking model as the benchmark model, the XGBoost, CatBoost, and LightGBM components in the base learner layer and the meta-learner layer were removed respectively to verify the necessity of the base learner integration; the RandomForest algorithm in the meta-learner layer was changed to logistic regression, gradient boosting decision tree (GBDT), and support vector machine (SVM) to verify the necessity of the meta-learner. The results of the ablation experiment are shown in Table 6. The experimental dataset was selected as CIC-IDS2017 to show the performance of the model when facing complex network features.
[0181] Table 6 Comparison of ablation experiment results
[0182]
[0183] It can be seen from the experimental results that changing specific modules in the model will lead to varying degrees of performance degradation in terms of accuracy, precision, recall, and F1 score. Since the F1 score can show the balance of the accuracy and integrity of the ablation module through a single metric, the degree of reduction in the F1 score is used to quantify the impact of each module.
[0184] Removing the XGBoost component has the highest impact on the model, reducing the F1 score by 0.15%, indicating that this base learner component is a key component of the model. The removal of the base learner results in a 0.01%-0.15% decrease in the F1 score, indicating that the heterogeneity of the three base learners is complementary to the learning of network traffic characteristics. As the meta-learner, the impact of the Random Forest is 0.05% higher than that of GBDT, 0.07% higher than that of Logistic Regression, and 0.06% higher than that of SVM, indicating that the Random Forest model is more suitable for feature recombination and capturing high-order features, and at the same time, the non-linear meta-learner is more suitable for this model. Figure 8 Taking the absolute value of the amplitude of the F1 score reduction in each ablation experiment as the influence score, the ranking of the influence of the ablation module on the complete model is shown. The results of the ablation experiment indicate that the ensemble learning model constructed in this paper provides a clear direction for future optimization: focusing on maintaining the XGBoost base learner and the ForestRandom meta-learner, and paying attention to their advantages in dealing with non-linear features, and secondly optimizing other base learners.
[0185] In summary, with the rapid development of the network and artificial intelligence, the network environment is becoming increasingly complex and severe, and it is of great value and significance to quickly and effectively detect network traffic. Traditional network abnormal traffic detection methods have problems such as high-dimensional feature redundancy and poor model generalization. Therefore, this embodiment proposes an XPM-OPStacking ensemble machine learning model for network abnormal traffic detection. XPM extracts key traffic features of high value, and OPStacking accurately identifies network threats, thereby detecting abnormal traffic. The experimental results show that the XPM-OPStacking method is superior to traditional machine learning models in various evaluation indicators and has strong generalization ability.
[0186] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. Therefore, this application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0187] Each embodiment in this specification is described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the above embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0188] The above embodiments have introduced the present invention in detail. Specific examples are used in this article to elaborate on the principles and embodiments of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A network abnormal traffic detection method based on ensemble learning, characterized in that, The method includes the following steps: By integrating the feature importance analysis of Extreme Gradient Boosting, Pearson correlation analysis, and mutual information scores, the traffic data in the network traffic data is screened and dimensionality-reduced to form a final traffic feature set; Using the Extreme Gradient Boosting algorithm, the Categorical Gradient Boosting algorithm, and the LightGBM as the base learners, the Random Forest algorithm as the meta-learner, and introducing the Optuna optimization framework for quickly finding the best hyperparameters in the high-dimensional hyperparameter space, a stacking generalization model for identifying abnormal traffic data in the network security threat situation recognition scenario is constructed; The stacking generalization model is used to detect the abnormal traffic data in the traffic feature set, identify threats, and thus determine whether the traffic data is abnormal.
2. The network abnormal traffic detection method according to claim 1, characterized in that, The specific process of forming the final traffic feature set includes the following steps: Perform data preprocessing on network traffic data F = {f1, f2,..., f n} to form multiple traffic feature subsets F1. The data preprocessing includes invalid feature removal, feature one-hot encoding, feature standardization, and label numericalization; Based on the XGBoost feature screening model, the feature importance of the traffic data in the traffic feature subset F1 is quantified, sorted in descending order, and the traffic features with high feature importance scores are retained according to the preset criteria to form a preliminarily screened traffic feature subset F2; Perform Pearson correlation analysis on the traffic data in the traffic feature subset F2, set the Pearson traffic feature correlation threshold θ2, and calculate the mutual information scores of the traffic features in the traffic feature subset F2 at the same time; If there are traffic features exceeding the Pearson traffic feature correlation threshold θ2, then compare the mutual information scores, and eliminate the traffic features with low scores to form the final traffic feature set F3; If no traffic feature exceeds the set Pearson traffic feature correlation threshold θ2, a traffic feature set F3 is directly formed. For the traffic feature f j there is: where |r jk | is the absolute value of the Pearson correlation coefficient r jk .
3. The network abnormal traffic detection method according to claim 2, wherein The formation of the preliminarily screened traffic feature subset F2 is: Screen out the traffic feature subset F1 with feature importance higher than the XGBoost traffic feature importance threshold θ1, and the expression is: F2 = {f j ∈ F | S XGB (f j ) > θ1} where f j is the j-th traffic feature in the traffic feature subset F1; S XGB (f j ) is the XGBoost importance score of the traffic feature f j .
4. The network abnormal traffic detection method according to claim 2, wherein The Pearson correlation analysis of the traffic data in the traffic feature subset F2 is: For any two traffic features f j and f k in the traffic feature subset F2, if the absolute value of their Pearson correlation coefficient r jk exceeds the Pearson traffic feature correlation threshold θ2, then retain the traffic feature with a higher mutual information score. The expression is as follows: If where, I(f j , Y) is the mutual information score between the traffic feature f j and the target variable Y; I(f k , Y) is the mutual information score between the traffic feature f k and the target variable Y.
5. The network abnormal traffic detection method according to claim 2, wherein The formation of the final traffic feature set F3 is: For the screened traffic feature subset F2, retain the traffic features with mutual information scores higher than the traffic feature mutual information score threshold θ3, and the expression is: F3 = {f j ∈ F2 | I(f j ; Y) > θ3} where I(f j , Y) is the mutual information score between the flow feature f j and the target variable Y.
6. The network abnormal traffic detection method according to claim 1, characterized in that The process of using the stacking generalization model to detect the abnormal traffic data in the traffic feature set, identify threats, and thus determine whether the traffic data is abnormal specifically includes the following steps: Let the optimization objective of the traffic feature set in the input threat recognition module be to maximize the comprehensive evaluation index, that is: S = ω1Accuracy + ω2Precision + ω3Recall + ω4F1 where x i represents the network traffic data feature vector; y i is the actual threat label of the i-th sample; Accuracy, Precision, Recall, and F1 represent accuracy, precision, recall, and F1-score respectively; ω1, ω2, ω3, and ω4 are the index weights respectively, and ω1 + ω2 + ω3 + ω4 = 1; Input traffic feature set F3 to the base learner L j , the base learner L1 iteratively constructs multiple decision trees to minimize the loss function accordingly The expression is: where l(·) is calculated using logarithmic loss; y is the actual threat label; is the predicted threat label; is the predicted threat label of the i-th sample; n is the total number of samples, that is, the number of samples contained in the traffic feature set F3; Obtain the parameters of the tree model by calculating the gradient of the loss function to get the predicted output That is: where η is the learning rate; f k is a single regression tree; k ∈ K is the number of iterations; The base learner L2 is trained using a symmetric decision tree and its predicted output is obtained after K iterations That is: where \(a(x i , T)\) is the probability prediction value output after the sample \(i\) falls into the leaf node of the decision tree \(T\); For the base learner L3, its predicted output is as follows: where ε is the learning rate of the base learner L3; h k (x i ) represents the predicted value of the k-th decision tree for the i-th sample; Use the prediction results of the base learner as new features and combine them with the original network traffic data feature vector x i to form a new network traffic feature vector Use the new network traffic feature vector z i and the corresponding threat label y i Train the meta-learner M, whose loss function is: In the formula, l'(·) is calculated using the logarithmic loss calculation method; Constructing a surrogate model using Gaussian process Approximating the objective function S based on the existing experimental result set To estimate the value of S under different hyperparameter combinations, and calculate the new hyperparameter combination θ through the expected improvement acquisition function EI(θ) new The expression is as follows: where S best is the optimal optimization objective; θ (k) is the hyperparameter combination of the k-th experiment; S (k) is the objective function value of the k-th experiment; t is the number of experiments included in the experimental result set; Save the hyperparameter combination θ obtained from the experiment new and the best optimization objective S best to the result set; Repeat the experiment until the preset number of trials is reached to obtain the final optimal hyperparameter combination θ * , the optimal hyperparameter combination θ * The corresponding stacked generalization model is the optimal training model. Use this model to predict the probability that sample i belongs to class c After calculating the probability distributions of each category, the final predicted category is as follows: In the formula, argmax is a function that calculates the c value when the probability is the largest, and is used to determine the most likely threat type of the traffic sample i.
7. The network abnormal traffic detection method according to claim 6, characterized in that The probability that the sample i belongs to the category c predicted by the stacked generalization model has the following expression: where, f M,c (z i ; θ M ) is the prediction score of the meta-learner M for sample i on class c, and there are C classes in total; f M,k (z i ; θ M ) is the prediction score of the meta-learner M for sample i on class k; exp(·) is the natural exponential function.
8. A system for implementing the network abnormal traffic detection method according to any one of claims 1-7 above, characterized in that, It includes: A traffic data preprocessing module that preprocesses network traffic data, including data cleaning, feature pre-screening, encoding conversion, and standardization operations; A feature selection module that analyzes the importance, linear correlation, and non-linear dependence of traffic features based on feature importance analysis, Pearson correlation analysis, and mutual information scores, and screens out the most valuable features to form a traffic feature set; A threat identification module that integrates the Optuna optimization framework and the Stacking ensemble learning algorithm to detect the abnormal traffic data in the traffic feature set, identify threats, and thus determine whether the traffic is abnormal.
Citation Information
Patent Citations
Network abnormal flow detection method, model and system
CN112784881A
Encrypted malicious traffic detection method and system based on multi-feature selection stacking
CN118353724A
Dynamic prediction and optimization control method for power grid line loss driven by deep learning
CN118469352A
VPN traffic identification method and system based on multi-model fusion
CN119675905A
Cited By
Multi-stage cross-domain adaptive feature selection method
CN121598218A