Abnormal network traffic detection method based on deep multi-stack ensemble learning
By adopting a deep multi-stack integrated learning method in malicious traffic detection, combining RGB image feature representation and multi-stack integration strategy, and integrating multiple deep learning models, the shortcomings of traffic feature representation and integration strategies in the existing technology are solved, and the detection accuracy and robustness are significantly improved.
Patent Information
- Application Number
- CN202510230395.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
AI Technical Summary
Existing malicious traffic detection methods have shortcomings in traffic feature representation and integration strategies, resulting in loss of edge information and difficulty in dealing with complex timing features and diverse attack types in a single model.
Using a deep multi-stack integrated learning method, by splitting, pruning and filling network traffic, different layers of information are stored in different channels of RGB images, and a multi-stack integration strategy is introduced to integrate deep multi-stack integrated deep learning models such as CNN, LSTM, BiLSTM, TCN and BiTCN to build the first and second layer meta learners of deep multi-stack integration.
It significantly improves the accuracy and robustness of malicious traffic detection, can capture the multi-dimensional characteristics of network traffic more comprehensively, and enhances the ability to identify abnormal behaviors in large-scale network traffic.
Smart Images

Figure CN120074932A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of abnormal network traffic detection, and relates to an abnormal network traffic detection method based on deep multi-stack ensemble learning. Background Art
[0002] With the rapid development of the Internet, network traffic has grown explosively, resulting in network intrusion behaviors such as DoS attacks, network scanning attacks, and traffic hijacking becoming serious threats. These attacks can not only cause service interruptions and data leaks, but also endanger personal privacy and national security. Therefore, timely and effectively detecting network intrusion behaviors has become a key task in the field of network security. There are obvious differences in characteristics between malicious traffic and normal traffic. Therefore, by accurately detecting malicious traffic in large-scale network traffic, potential network attacks can be effectively discovered, and network security can be guaranteed.
[0003] To cope with network attacks, researchers have proposed various malicious traffic detection methods. The simple and easy-to-implement port-based method has problems such as a high false alarm rate and a lack of adaptability. Therefore, scholars have turned to using machine learning models such as decision trees and SVMs to automatically learn patterns from network traffic, but there are still deficiencies in relying on manual experience and being difficult to capture complex traffic characteristics. In contrast, deep learning methods can automatically learn complex features and show better effects in malicious traffic detection. Therefore, they have become an important technology to solve this problem.
[0004] Currently, many malicious traffic detection methods use a single deep learning model for detection. However, with the increasing complexity of the network environment, a single model cannot effectively handle various types of traffic. Ensemble learning can significantly improve the detection effect of malicious traffic by combining the advantages of multiple models. The malicious traffic detection method based on ensemble learning uses different ensemble strategies to fuse multiple single models, which can improve the detection ability.
[0005] However, most of the existing ensemble learning methods convert traffic in PCAP format into grayscale images. Although this method can represent some traffic characteristics, it will cause the loss of edge information due to the convolution kernel being unable to fully cover the image edge. In addition, most of the existing malicious traffic detection methods based on ensemble learning adopt a simple voting strategy and cannot make full use of the complex relationships between models.
[0006] To address the deficiencies of existing methods in traffic feature representation and integration strategies, the present invention proposes a malicious traffic detection method based on deep multi-stack ensemble learning. By splitting, pruning, and padding network traffic, different layer information is effectively stored in different channels of an RGB image, capturing network traffic features more comprehensively. Then, a multi-stack ensemble strategy is introduced, constructing five deep learning base models such as CNN, LSTM, BiLSTM, TCN, and BiTCN as the first layer of the deep multi-stack ensemble to fully utilize the advantages of different models to improve the detection accuracy of malicious traffic. The bidirectional convolution feature and temporal modeling ability of the BiTCN model are used as the second-layer meta-learner to enhance the complexity and diversity of the model, reducing the risk of overfitting. The method proposed by the present invention optimizes traffic feature representation and integration strategies, gives full play to the advantages of different deep learning models, improves the accuracy and robustness of malicious traffic detection, and significantly enhances the ability to identify abnormal behaviors in large-scale network traffic. Summary of the Invention
[0007] Existing malicious traffic detection methods often use grayscale image representation, resulting in the loss of edge information and the introduction of invalid data by padding. In addition, traditional malicious traffic detection methods often rely on a single model and are difficult to handle complex temporal features and diverse attack types in network traffic. To address these problems, the present invention proposes a method for detecting abnormal network traffic based on a deep multi-stack ensemble learning architecture.
[0008] The present invention provides a method for detecting abnormal network traffic based on deep multi-stack ensemble learning, including:
[0009] Step 1, obtaining the original traffic file, splitting the network traffic into multiple data according to sessions, extracting all layer and application layer information, removing empty traffic and duplicate traffic, intercepting and complementing the split traffic data to generate a byte sequence of the traffic, and marking the type of the traffic;
[0010] Step 2, performing normalization processing on different types of PCAP files, placing all layer information of the PCAP in the R channel and B channel of the RGB image, and placing the application layer information in the G channel to achieve RGB image feature representation of the network traffic;
[0011] Step 3, using uniform random sampling to divide the generated RGB image into a training set and a test set, inputting the training set into a deep multi-layer stacked ensemble learning framework to train multiple base models and construct a high-level meta-model, using the test set for verification to obtain an abnormal network traffic detection model, completing abnormal traffic detection and outputting the detection result.
[0012] In the first aspect, the specific steps for obtaining the RGB feature representation of network traffic in the above Step 1 are as follows:
[0013] Step 1.1, capture network traffic and save it in the.pcap format;
[0014] Step 1.2, split the obtained network traffic into multiple session records by taking a session-based splitting method;
[0015] Step 1.3, divide each session into two parts: all layers and the application layer, and store them in different folders respectively;
[0016] Step 1.4, delete the address information of the traffic, and fill this position with randomly generated addresses to ensure that the training result is only related to the traffic content and avoid the address interfering with traffic classification;
[0017] Step 1.5, traverse all traffic data and delete the blank traffic and duplicate traffic among them;
[0018] Step 1.6, set the length of the traffic sequence to 784 and select the first 784 bytes of the flow and session. If the total length of the traffic sequence exceeds this length, then intercept the first 784 bytes. If the traffic length is insufficient, fill the insufficient part with '0' to obtain the byte sequence X = (x 0 , x 1 ,..., x T ) for each session and flow;
[0019] Step 1.7, label the generated network traffic byte sequence and set the category label corresponding to each type of traffic.
[0020] In the second aspect, the specific steps of the above step 2 are as follows:
[0021] Step 2.1, perform normalization processing on different types of PCAP files;
[0022] Step 2.2, read the information files of all layers and the application layer in binary mode and convert them into hexadecimal form;
[0023] Step 2.3, use formula (1) to convert the network traffic represented in hexadecimal into a uint8 array to map it to the interval [0, 255], and then place all layer information of the PCAP in the R channel and B channel of the RGB image, and place the application layer information in the G channel;
[0024] [Value uint8 = X × 16 1 + Y × 16 0 Formula (1) where X is the tens digit after being converted into a hexadecimal number, and Y is the units digit after being converted into a hexadecimal number
[0025] Step 2.4, use formula (2) to stack the R, G, and B channel matrices along the third dimension through np.dstack to form a three-dimensional RGB array rgb_array;
[0026] rgb_array = np.dstack((r_matix, g_matix, b_matix))[:, :, :3] Formula (2)
[0027] Among them, r_matrix is the uint8 data of the R channel, g_matrix is the uint8 data of the R channel, and b_matrix is the uint8 data of the B channel
[0028] Step 2.5, use the PIL library to convert the NumPy array into an image object to realize the RGB image feature representation of network traffic.
[0029] Thirdly, the specific steps of the above step 3 are as follows:
[0030] Step 3.1, use the method of uniform random sampling to divide the generated RGB image into ten parts, where nine parts are used as the training data set and one part is used as the test data set;
[0031] Step 3.2, in the first layer stacking process, select five malicious traffic detection models including CNN (which can automatically extract complex features in network traffic through multiple convolutional and pooling operations, and has strong feature learning ability and sensitivity to local features), LSTM (which can effectively capture long-term dependencies and temporal features in network traffic, has high accuracy and low false alarm rate when identifying potential attack patterns), BiLSTM (which can capture both past and future information in network traffic, learn the spatio-temporal features of traffic more comprehensively, and effectively identify complex patterns and long-term dependencies), TCN (which can effectively capture the temporal features in network traffic, avoid the problem of gradient vanishing or explosion, and at the same time process long sequence data through causal convolution and dilated convolution structures, and has high computational efficiency to meet the needs of large-scale traffic analysis), and BiTCN (which can capture both past and future information in network traffic, improve the feature extraction ability through a bidirectional temporal convolution structure, and maintain high computational performance, so as to better understand and analyze dynamic network traffic) as the base models, and then input the network traffic training set Z = (z 1 , z 2 , …, z n ) and the test set T = (t 1 , t 2 , …, t m ) of the transformed image into each base model respectively, and obtain the output value p i of each model's prediction on the training set and the output value q of the prediction on the test seti , where p i and q i are one-dimensional column vectors with n and m numerical values, and i represents the serial number of different base models;
[0032] Step 3.3, concatenate the prediction outputs of the first layer into vectors X = (p 1 , p 2 , p 3 , p 4 , p 5 ) and Y = (q 1 , q 2 , q 3 , q 4 , q 5 ), initialize the meta-learner to adjust its input size to the number of base models, input X again as the training set and validation set into the meta-learner BiTCN for training, and use Y as the test set of the second-layer meta-model BiTCN for prediction. Finally, take the prediction result of the meta-learner as the final result and return it.
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0034] 1. By splitting, pruning, and filling network traffic, the present invention innovatively stores different layer information effectively in different channels of RGB images. Specifically, all layer information is ingeniously placed in the R and G channels, while the application layer information is precisely placed in the B channel. This unique design avoids the defect that traditional grayscale image representation methods are prone to losing edge information, thus enabling more comprehensive capture of the multi-dimensional features of network traffic; this technological breakthrough provides a more accurate feature representation for malicious traffic detection, significantly improving the accuracy and efficiency of detection.
[0035] 2. Aiming at the limitations of a single model in traditional ensemble learning methods when dealing with multiple traffic types, the present invention adopts a multi-stack ensemble strategy, integrating five deep learning models, namely CNN, LSTM, BiLSTM, TCN, and BiTCN, as the first layer of deep multi-stack ensemble to comprehensively capture traffic information, thereby improving the detection accuracy and the overall generalization ability of the model. More advanced is that the present invention also introduces the BiTCN model as the second-layer meta-learner of the deep multi-stack ensemble. This innovative design can capture both past and future information in the network traffic feature sequence, thus further improving the recognition ability of abnormal behaviors in large-scale network traffic.
[0036] 3. The present invention deeply integrates a deep learning model with an ensemble learning method. This combination not only fully exploits the advantages of the deep learning model in processing complex data but also utilizes the potential of the ensemble learning method in enhancing the stability and accuracy of the model. This innovative integration strategy has enabled the present invention to achieve remarkable technological breakthroughs in the field of malicious traffic detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is the overall flowchart of an abnormal network traffic detection method DMSE based on deep multi-stack ensemble learning.
[0038] Figure 2 is the deep multi-stack ensemble framework diagram.
[0039] Figure 3 is the overall framework of the abnormal network traffic detection method DMSE based on deep multi-stack ensemble learning.
[0040] Figure 4 are the detection results of the DMSE model of the invention and other ensemble methods (including: hard-voting, soft-voting, adaptive boosting ensemble adaboost, and bagging algorithm bagging) on the USTC-TFC2016 dataset.
[0041] Figure 5 are the detection results of the DMSE model of the invention and other ensemble methods (including: hard-voting, soft-voting, adaptive boosting ensemble adaboost, and bagging algorithm bagging) on the CTU dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The present invention will be further described below in conjunction with the drawings and embodiments. It should be noted that the described embodiments are only intended to facilitate the understanding of the present invention and do not impose any limitations on it.
[0043] In view of the deficiencies of the existing malicious traffic detection methods in traffic feature representation and integration strategies, the present invention proposes a detection method based on deep multi-stack ensemble learning. By efficiently storing different layers of traffic information in each channel of an RGB image and introducing a multi-stack integration strategy to combine multiple deep learning models, the accuracy and robustness of malicious traffic detection are significantly improved, and the ability to identify abnormal behaviors in large-scale network traffic is enhanced.
[0044] As Figure 1 shown, an abnormal network traffic detection method based on deep multi-stack ensemble learning proposed by the present invention includes:
[0045] Step 201 obtains the original traffic file, splits the network traffic into multiple data in the form of sessions, extracts all layer and application layer information, removes empty traffic and duplicate traffic, intercepts and complements the split traffic data to generate a byte sequence of the traffic, and marks the type of the traffic;
[0046] The purpose of implementing traffic splitting in the present invention is that the original traffic file is stored in the.pcap format which cannot be directly trained by a neural network, and a large amount of blank traffic and duplicate traffic existing in the traffic file will interfere with the learning of the neural network, resulting in insufficient learning features; splitting the network traffic in the form of flows and sessions can reduce the training scope of the model, and removing duplicate traffic and blank traffic can also reduce the error of the model, thereby enhancing the recognition accuracy of the model. In addition, the reason for implementing the interception and complementation of network traffic in the present invention is that for network traffic, the content that affects the traffic type judgment is often in the first part of the traffic. If all the traffic is used to train the neural network, the training efficiency will be reduced. Selecting the first 784 bytes of the traffic sequence as the interception length not only retains the characteristics of the network traffic, but also ensures that the size of the traffic sequence input into the model is the same.
[0047] Step 2011 captures network traffic and saves it in the.pcap format;
[0048] Step 2012 splits the obtained network traffic into multiple session records by adopting a session-based splitting method;
[0049] Step 2013 divides each session into two parts: all layers and the application layer, and stores them in different folders respectively;
[0050] Step 2014 deletes the address information of the traffic, and complements this position with a randomly generated address to ensure that the training result is only related to the traffic content and avoid the address interfering with the traffic classification;
[0051] Step 2015 traverses all the traffic data and deletes the blank traffic and duplicate traffic therein;
[0052] Step 2016 sets the length of the traffic sequence to 784 and selects the first 784 bytes of the flow and session. If the total length of the traffic sequence exceeds this length, the first 784 bytes are intercepted. If the traffic length is insufficient, the insufficient part is complemented with '0' to obtain the byte sequence X=(x 0 ,x 1 ,...,x T ) of each session and flow;
[0053] Step 2017 marks the generated byte sequence of the network traffic and sets the category label corresponding to each type of traffic.
[0054] Step 202 performs normalization on different types of PCAP files, places all layer information of the PCAP in the R channel and B channel of the RGB image, and places the application layer information in the G channel, to achieve the RGB image feature representation of network traffic.
[0055] Step 2021 performs normalization on different types of PCAP files;
[0056] Step 2022 reads the information files of all layers and the application layer in binary mode and converts them into hexadecimal form;
[0057] Step 2023 converts the network traffic represented in hexadecimal into a uint8 array to map it to the [0, 255] interval, and then places all layer information of the PCAP in the R channel and B channel of the RGB image, and places the application layer information in the G channel;
[0058] Step 2024 stacks the R, G, and B channel matrices along the third dimension through np.dstack to form a three-dimensional RGB array;
[0059] Step 2025 uses the PIL library to convert the NumPy array into an image object to achieve the RGB image feature representation of network traffic.
[0060] Step 203 uses uniform random sampling to divide the generated RGB image into a training set and a test set, inputs the training set into a deep multi-layer stacked ensemble learning framework to train multiple base models and construct a high-level meta-model, and uses the test set for verification to obtain an abnormal network traffic detection model, complete the abnormal traffic detection and output the detection result.
[0061] Step 2031 uses the method of uniform random sampling to divide the generated RGB image into ten parts, with nine parts as the training data set and one part as the test data set;
[0062] Step 2032 selects five malicious traffic detection models, including CNN (which can automatically extract complex features in network traffic through multi-layer convolution and pooling operations, and has strong feature learning ability and sensitivity to local features), LSTM (which can effectively capture long-term dependencies and temporal features in network traffic, has high accuracy and low false alarm rate when identifying potential attack patterns), BiLSTM (which can capture both past and future information in network traffic, learn the spatio-temporal features of traffic more comprehensively, and effectively identify complex patterns and long-term dependencies), TCN (which can effectively capture the temporal features in network traffic, avoid the problem of gradient vanishing or explosion, and at the same time process long sequence data through causal convolution and dilated convolution structures, and has high computational efficiency to meet the needs of large-scale traffic analysis), and BiTCN (which can capture both past and future information in network traffic, improve the feature extraction ability through a bidirectional temporal convolution structure, and maintain high computational performance, so as to better understand and analyze dynamically changing network traffic), as the base models. Then, the network traffic training set Z = (z 1 , z 2 , …, z n ) and the test set T = (t 1 , t 2 , …, t m ) of the transformed images are respectively input into each base model to obtain the output values p i of each model for prediction on the training set and the output values q i for prediction on the test set, where p i and q i are one-dimensional column vectors with n and m values, and i represents the serial number of different base models;
[0063] Step 2033 concatenates the prediction outputs of the first layer into vectors X = (p 1 , p 2 , p 3 , p 4 , p 5 ) and Y = (q 1 , q 2 , q 3 , q 4 , q 5 ). Initialize the meta-learner to adjust its input size to the number of base models. Input X again as the training set and validation set into the meta-learner BiTCN for training, and use Y as the test set of the second-layer meta-model BiTCN for prediction. Finally, take the prediction result of the meta-learner as the final result and return it.
[0064] The present invention mainly focuses on the detection of abnormal network traffic, proposes an abnormal network traffic detection method DMSE based on deep multi-stack ensemble learning, and selects the USTC-TFC2016 dataset and the CTU dataset for effect testing. Among them, the USTC-TFC2016 dataset includes 10 types of abnormal traffic and 1 type of normal traffic collected in the real network environment from 2011 to 2015; the CTU dataset includes 10 types of abnormal traffic and 1 type of normal traffic collected from 2016 to 2019. The present invention compares the proposed DMSE model with other hard-voting, soft-voting, adaboost, and bagging, and illustrates the detection ability of the model by calculating the average detection efficiency (including: accuracy, false positive rate FPR, true positive rate TPR, and F1 value F1-measure).
[0065] Figure 4 shows the detection effects of five ensemble models on the USTC-TFC2016 dataset. From Figure 4 it can be seen that the DMSE ensemble method proposed by the present invention has the highest detection accuracy Accuracy for network traffic, while the detection accuracy of bagging is the worst, and the TPR and FPR of the adaboost ensemble strategy are not very high. Compared with other ensemble strategies, the DMSE ensemble strategy has a higher accuracy, verifying that the DMSE model can make more full use of the advantages of each model to improve the detection accuracy.
[0066] Figure 5 shows the detection effects of five ensemble models on the CTU dataset. From Figure 5 it can be seen that the DMSE model proposed by the present invention has the highest detection accuracy for network traffic, which can reach 99.19%; as the concealment of abnormal network traffic becomes higher and higher, the DMSE model proposed by the present invention still has a good detection effect on abnormal network traffic. From Figure 5 it can be seen that the DMSE model has improved in terms of Accuracy, TPR, and F1-measure. Compared with the hard_voting ensemble model, the average Accuracy has increased by about 2.0%, and compared with the bagging ensemble model, the F1-measure has increased by about 4%. Generally speaking, on the CTU dataset, compared with the hard_voting, soft_voting, adaboost, and bagging models, the accuracy of the DMSE model has increased by 2.0%, 1.7%, 1.3%, and 3.8% respectively, and the accuracy of the F1-measure has increased by 3.2%, 2.9%, 3.2%, and 3.9% respectively.
Claims
1. A method for detecting abnormal network traffic based on deep multi-stack ensemble learning, which features the following steps: Step 1: Get the original traffic file and segment the network traffic into multiple data according to the session method, extract all layer and application layer information and remove empty traffic and duplicate traffic, intercept and complete the segmented traffic data to generate a byte sequence of the traffic, and mark the type of traffic; Step 2: Normalize different types of PCAP files, put all layer information of PCAP in the R channel and B channel of RGB image, and put application layer information in the G channel, so as to realize RGB image feature representation of network traffic; Step 3: Use uniform random sampling to divide the generated RGB images into training sets and test sets, and input the training set into a deep multi-layer stacked ensemble learning framework to train multiple base models and build a high-level meta-model. Use the test set for verification to obtain an abnormal network traffic detection model, complete abnormal traffic detection, and output the detection results.
2. The method according to claim 1, characterized in that The specific implementation of step 1 includes the following steps: Step 1.1, capture network traffic and save it in .pcap format; Step 1.2, dividing the acquired network traffic into multiple session records by taking the session as the unit; Step 1.3, divide each session into two parts: all layers and application layer, and store them in different folders; Step 1.4, delete the address information of the traffic, and fill the position with a randomly generated address to ensure that the training result is only related to the traffic content, avoiding the interference of the address on the traffic classification; Step 1.5, traverse all traffic data and delete blank traffic and duplicate traffic; Step 1.6, set the length of the traffic sequence to 784 and select the first 784 bytes of the flow and session. If the total length of the traffic sequence exceeds this length, the first 784 bytes are intercepted. If the traffic length is insufficient, the insufficient part is filled with '0' to obtain the byte sequence X = (x0, x1, ..., x T ); Step 1.7, mark the generated network traffic byte sequence and set the type label corresponding to each type of traffic.
3. The method according to claim 1, characterized in that The specific implementation of step 2 includes the following steps: Step 2.1, normalize different types of PCAP files; Step 2.2, read the information files of all layers and application layers in binary mode and convert them into hexadecimal form; Step 2.3, convert the network traffic represented in hexadecimal into a uint8 array to map it to the interval [0, 255], and then put all the layer information of PCAP in the R channel and B channel of the RGB image, and put the application layer information in the G channel; Step 2.4, stack the R, G, and B three-channel matrices along the third dimension using np.dstack to form a three-dimensional RGB array; In step 2.5, use the PIL library to convert the NumPy array into an image object to implement the RGB image feature representation of network traffic.
4. The method according to claim 1, characterized in that: The specific implementation of step 3 includes the following steps: Step 3.1, use uniform random sampling method to divide the generated RGB image into ten parts, nine of which are used as training data sets and one as test data set; Step 3.2: In the first layer stacking process, five deep learning models including CNN (convolutional neural network), LSTM (long short-term memory network), BiLSTM (bidirectional long short-term memory network), TCN (temporal convolutional network) and BiTCN (bidirectional temporal convolutional network) are selected as base models, and trained on the same network traffic dataset to obtain the output p of each base model. i and q i (i=1,2,…,5), where p i and q i is the one-dimensional column vector output produced by the basis model prediction; In step 3.3, the prediction output of the first layer is concatenated into vectors X = (p1, p2, p3, p4, p5) and Y = (q1, q2, q3, q4, q5), and X is again input into the meta-learner BiTCN as the training set and validation set for training, and Y is used as the test set of the second-layer meta-model BiTCN for prediction to obtain the final malicious traffic detection result.