Network flow classification method, system and device and storage medium
By adding self-attention and convolutional attention modules after LSTM and CNN networks, the information loss and local optimal solution problems caused by network traffic data resizing are solved, feature extraction capabilities are enhanced, and the accuracy of network flow classification and minority class recognition capabilities are improved.
Patent Information
- Application Number
- CN202510466503.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, network traffic data needs to be resized to adapt to the input layer of the CNN network, resulting in information loss, and the LSTM model is prone to local optimal solutions, ignoring important network traffic characteristics, affecting classification accuracy.
A new self-attention module was added after the LSTM model to enhance the time and statistical feature extraction ability; a new convolutional attention module with channel and spatial attention mechanism was added after the CNN network to enhance the spatial feature extraction ability, and a network flow data was processed through the self-attention module and the convolutional attention module to form fusion features for classification.
Effectively avoid data loss, enhance attention to important features, improve feature representation, improve classification accuracy, especially the ability to identify a few categories.
Smart Images

Figure CN120281671A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of encrypted network traffic classification, and in particular, to a network traffic classification method, system, device and storage medium. Background Art
[0002] With the expansion and wide adoption of Internet technology, the diversity of transmitted data is also increasing, including text, images, audio, video and other forms of data. At the same time, network security threats such as malware, cyberattacks and data breaches are also escalating. To address these threats, network traffic encryption technology has been widely adopted to protect information transmission, among which encrypted flow classification is the key to effectively monitor and manage the network. However, with the continuous development of network technology, the accurate classification ability of traditional traffic classification methods for encrypted network traffic is becoming increasingly limited.
[0003] In recent years, machine learning technology has been continuously improved and the computing cost has been continuously reduced, and it has been widely used in network traffic classification. Researchers apply machine learning to extract traffic features and use algorithms for classification, aiming to improve accuracy and efficiency. However, traditional machine learning methods rely on manually selected features, usually based on subjective judgment. The lack of a standardized feature selection process leads to inconsistent and suboptimal classification performance.
[0004] With the rapid development of CNNs, deep learning-based methods have become the dominant methods for various applications including network traffic classification. Different from traditional machine learning methods, deep learning technology can automatically extract important hidden features and representations from raw input data, thus enabling direct learning from raw traffic data.
[0005] Existing deep learning methods for encrypted network traffic classification usually focus on exploiting spatial or temporal features, using CNN to extract spatial features and using LSTM to extract temporal features. Although these methods improve classification accuracy, CNN is designed to process input data of fixed size and structure. Therefore, the original network traffic must be resized through cropping, padding or other transformations to fit the input layer of CNN. This preprocessing, especially when it involves a large amount of cropping or padding, often results in serious information loss, thus hindering the ability to analyze the overall structure of network traffic. In addition, due to the problem that the LSTM model is prone to local optimal solutions, some important network traffic characteristics are ignored, which has a negative impact on subsequent classification. Summary of the Invention
[0006] In view of the deficiencies in the prior art, the present invention provides a network traffic classification method, system, device and storage medium, which solves the problems in the prior art that due to the need to resize network traffic data to adapt to the input layer of the CNN network, serious information loss will occur, and the LSTM model is prone to local optimal solutions, resulting in the neglect of some important network traffic characteristics and having a negative impact on subsequent classification.
[0007] According to an embodiment of the present invention, a network traffic classification method includes:
[0008] Obtain network traffic data, and import the network traffic data into the LSTM model and the CNN network respectively to obtain the first feature and the second feature;
[0009] Process the first feature using the self-attention module to obtain the first concatenated feature;
[0010] Import the second feature into the convolutional attention module for processing to obtain the second concatenated feature;
[0011] Concatenate the first concatenated feature and the second concatenated feature to obtain a fused feature, and then pass the fused feature through the prediction module for label prediction to obtain the corresponding network traffic category.
[0012] Preferably, after obtaining the network traffic data, the network traffic data needs to be processed, and the processing method includes:
[0013] Convert the formats of all data in the network traffic data into the same format, and then slice the network traffic data to obtain multiple sessions;
[0014] Extract the statistical feature information in each session, and delete the sessions with unnecessary information in the statistical feature information;
[0015] Normalize the sizes of all sessions to the same number of bytes, and then import all sessions into the LSTM model and the CNN network.
[0016] Preferably, the convolutional attention module includes a channel attention module and a spatial attention module;
[0017] The method of importing the second feature into the convolutional attention module for processing includes:
[0018] Import the second feature into the channel attention module to extract the first global feature information, and then concatenate the first global feature information with the second feature to obtain the first intermediate feature;
[0019] Import the first intermediate feature into the spatial attention module to extract the second global feature information, and then concatenate the second global feature information with the first intermediate feature to obtain the second intermediate feature;
[0020] Process the second intermediate feature through a fully connected layer to obtain a second concatenated feature.
[0021] Preferably, the method for extracting the first global feature information includes:
[0022] Perform global average pooling and global max pooling on the second feature respectively to obtain corresponding first global average information and first global max information;
[0023] Use a shared multi-layer perceptron to perform channel modeling on the first global average information and the first global max information respectively to generate corresponding first channel weight vector and second channel weight vector;
[0024] Perform modulo addition on the first channel weight vector and the second channel weight vector and then process with a sigmoid activation function to obtain the first global feature information.
[0025] Preferably, the method for extracting the second global feature information includes:
[0026] Perform global average pooling and global max pooling on the second feature respectively to obtain corresponding second global average information and second global max information;
[0027] Perform modulo addition on the second global average information and the second global max information and then process with a sigmoid activation function to obtain the first global feature information.
[0028] Preferably, the shared multi-layer perceptron includes a rectified linear unit and two convolutional layers, and the two convolutional layers are respectively connected to the input end and the output end of the rectified linear unit.
[0029] Preferably, the prediction module includes two fully connected layers and a Dropout module, and the two fully connected layers are respectively connected to the input end and the output end of the Dropout module.
[0030] On the other hand, according to an embodiment of the present invention, there is also provided a network flow classification system, which uses the above-mentioned network flow classification method, including:
[0031] A network flow interception module, which is used to intercept network flow data and preprocess the network flow data;
[0032] A processing module, which is used to perform feature extraction and feature fusion on the network flow data using an LSTM model, a CNN network, a self-attention module, and a convolutional attention module to obtain a fused feature;
[0033] A classification module, which is used to perform prediction classification on the network flow data according to the fused feature to obtain corresponding network flow categories.
[0034] On the other hand, according to an embodiment of the present invention, there is also provided a computer, including at least one processor and a memory, where the memory stores a computer program, and the computer program is configured to be executed by the processor to implement the above-mentioned network flow classification method.
[0035] On the other hand, according to an embodiment of the present invention, there is also provided a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and the computer program can be executed by one or more processors to implement the above-mentioned network flow classification method.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] In the present invention, a convolutional attention module with both channel attention mechanism and spatial attention mechanism is newly added after the CNN network to enhance the spatial feature extraction ability of the CNN network and avoid data loss. At the same time, a self-attention module is newly added after the LSTM model to enhance the extraction ability of the LSTM model for the temporal features and statistical characteristics of data and avoid the problem of local optimal solution, so that the entire classification model can more effectively focus on important features, strengthen key information, improve feature representation, and enhance the recognition ability for minority classes. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of the network flow classification method according to an embodiment of the present invention.
[0039] Figure 2 It is a model network architecture diagram of the present invention according to an embodiment of the present invention.
[0040] Figure 3 It is a schematic diagram of the dot product attention mechanism according to an embodiment of the present invention.
[0041] Figure 4 It is an architecture diagram of the convolutional attention module according to an embodiment of the present invention.
[0042] Figure 5 It is an architecture diagram of the channel attention module according to an embodiment of the present invention.
[0043] Figure 6 It is an architecture diagram of the spatial attention module according to an embodiment of the present invention.
[0044] Figure 7 It is a comparison diagram of the accuracy rates of the model of the present invention and other models according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The technical solutions in the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0046] AsFigure 1 As shown in the figure, an embodiment of the present invention proposes a network traffic classification method, including:
[0047] Obtain network traffic data, and import the network traffic data into an LSTM model and a CNN network respectively to obtain a first feature and a second feature;
[0048] Capture a segment of network traffic data, which captures traffic from various applications, including regular (non-VPN) and VPN traffic, and the dataset has the same format as the ISCX dataset. Among them, six types of regular traffic and six types of VPN traffic are selected, as shown in Table 1.
[0049] Table 1 Network Traffic Data Table
[0050]
[0051] The data preprocessing process is as follows:
[0052] (1) Unify the format of the dataset into the pcap format to ensure that the file formats of different network traffic types are consistent, facilitating subsequent traffic slicing and chunking.
[0053] (2) According to the description provided by the ISCX dataset, classify the traffic into 12 files, including email, chat, streaming, FT, VoIP, P2P, VPN_email, VPN_chat, VPN_streaming, VPN_file transfer, VPN_VoIP, and VPN_P2P.
[0054] (3) Use the SplitCap tool to slice the traffic data into multiple sessions. A session usually refers to a set of continuous data streams transmitted between two specific IP addresses and port numbers.
[0055] (4) Extract statistical feature information from the session traffic obtained by slicing the traffic data.
[0056] (5) Delete unnecessary information such as IP addresses, MAC addresses, version numbers, etc. in the session traffic generated during the traffic data slicing process.
[0057] (6) Normalize the flow size to 784 bytes.
[0058] As Figure 2 shown, import all the preprocessed sessions into a CNN network and an LSTM model respectively.
[0059] Process the first feature using a self-attention module to obtain a first concatenated feature;
[0060] In the LSTM branch, the preprocessed session is input into the LSTM module for feature extraction to obtain the first feature. Then, the self-attention mechanism is applied to the first feature output by the LSTM to fully enhance the extraction of statistical features and generate the first concatenated feature. A self-attention module is added after the LSTM model to enhance the LSTM model's ability to extract temporal features and statistical characteristics of the data, avoiding the problem of local optimal solutions.
[0061] The self-attention mechanism, also known as internal attention, is mainly used to process text and time series data. It determines the weights of each position by calculating the similarity between different positions in the input sequence, allowing the model to consider the relationships between all positions simultaneously and capture long-term dependencies within the sequence. The main steps of the self-attention mechanism are: (1) generating queries, keys, and values; (2) calculating attention scores; (3) weight normalization; (4) outputting the weighted sum. Q, K, and V are obtained by the following formulas respectively:
[0062] Q = XW Q
[0063] K = XW K
[0064] V = XW v
[0065] In the formula, X is the input, Q, K, and V are the query, key, and value respectively, and WQ, WK, and Wv are the corresponding weight matrices.
[0066] A common method used to calculate attention scores in the self-attention mechanism is dot-product attention, as Figure 3 shown. Figure 3 In it, the first MatMul represents a matrix multiplication operation, where Q and K are matrix dot-multiplied to calculate the attention scores. Scale refers to a scaling operation, where the attention scores are usually scaled by the square root of the k vector dimension. Mask is an optional component that performs a masking operation, setting the scores of certain positions to large negative numbers to ensure that the model rarely or does not pay attention to these positions. Then, Softmax is applied to normalize the scaled scores. Finally, in the second MatMul operation, the output is obtained by performing the dot-product of the values (V) and the attention weights, and the resulting output is shown in the following formula:
[0067]
[0068] In the formula, V is the original input, represents the square root of the k vector dimension.
[0069] The second feature is imported into the convolutional attention module for processing to obtain the second concatenated feature;
[0070] In the CNN branch, spatial features are extracted from the data in the session through two convolutional layers of the CNN to obtain the second feature. Then, the convolutional attention module is applied to the second feature output by the CNN network to further enhance the spatial feature extraction process. By adding a convolutional attention module with both channel attention mechanism and spatial attention mechanism after the CNN network, the spatial feature extraction ability of the CNN network is enhanced, and data loss is avoided.
[0071] The structure of the convolutional attention module is as Figure 4 shown, consisting of two main parts: the channel attention module as Figure 5 shown and the spatial attention module as Figure 6 shown.
[0072] The second feature extracted by the CNN network is imported into the channel attention module to extract the first global feature information, and then the first global feature information is concatenated with the second feature to obtain the first intermediate feature;
[0073] The first intermediate feature is imported into the spatial attention module to extract the second global feature information, and then the second global feature information is concatenated with the first intermediate feature to obtain the second intermediate feature;
[0074] The second intermediate feature is processed through a fully connected layer to obtain the second concatenated feature.
[0075] In the channel attention module, global average pooling and global max pooling are applied to the second feature to capture the first global average information and the first global max information of each channel. After that, a shared multi-layer perceptron (MLP) is used to model the non-linear dependence relationship between channels, generating the important weight vectors (corresponding first channel weight vector and second channel weight vector) of each channel. This shared multi-layer perceptron includes a rectified linear unit and two convolutional layers, and the two convolutional layers are connected to the input end and the output end of the rectified linear unit respectively. Finally, to reduce the number of parameters and improve the computational efficiency, a convolutional layer is used instead of the traditional fully connected layer in the MLP to obtain the first global feature information. The channel attention module can be expressed as the following formula:
[0076] M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))
[0077] In the formula, M c is the channel weight vector of the feature map F, σ is the sigmoid function, and MLP is the shared multi-layer perceptron. AvgPool and MaxPool represent average pooling and max pooling respectively.
[0078] In the spatial attention module, global average pooling and global max pooling are respectively performed on the second feature to obtain the corresponding second global average information and second global max information;
[0079] The second global average information and the second global max information are subjected to modulo addition and then processed using the sigmoid activation function to obtain the first global feature information.
[0080] The spatial attention mechanism can be expressed as the following formula:
[0081] M s (F) = σ(f([AvgPool(F), MaxPool(F)]))
[0082] In the formula, M s is the spatial weight vector of the feature map F, σ is the sigmoid function, and f is the convolution operation.
[0083] The first concatenated feature and the second concatenated feature are concatenated to obtain a fused feature, and then the fused feature is passed through the prediction module for label prediction to obtain the corresponding network flow category.
[0084] Through the above LSTM branch self-attention module and the convolutional attention module of the CNN branch, the model can more effectively focus on important features, strengthen key information, improve feature representation, and enhance the recognition ability of minority classes
[0085] The statistical features (the first concatenated feature) extracted by the LSTM branch are concatenated with the spatial features (the second concatenated feature) extracted by the CNN branch to form a fused feature set that combines the original network flow spatial features and statistical features, compensating for the loss of the original network flow data. Then it is imported into the prediction module for prediction classification. The prediction module includes two fully connected layers and a Dropout module, and the two fully connected layers are respectively connected to the input end and the output end of the Dropout module.
[0086] The loss function used by the prediction module is the cross-entropy loss function, which measures the difference between the actual label and the model prediction, as shown in the following formula:
[0087]
[0088] In the formula, N is the number of samples, C is the total number of categories, y i,c indicates whether sample i belongs to class C (belongs to 1, does not belong to 0), is the predicted probability that sample i belongs to class C, and L c is the cross-entropy loss value.
[0089] To verify the performance of the model of the present invention, the present invention uses the F1 score as a performance indicator and compares it with other prediction models:
[0090] To evaluate the impact of adding an attention module on the classification of minority-class traffic, the model of the present invention was compared with traditional methods (including the focal loss function (FL), SMOTE oversampling, and undersampling). The experimental results are shown in Table 2.
[0091] Table 2 Comparison of the model of the present invention with conventional methods
[0092]
[0093] As shown in Table 2, the model of the present invention is superior to traditional methods in terms of overall performance. The best results were achieved in terms of accuracy, recall, and F1-score, demonstrating the effectiveness and feasibility of the model of the present invention.
[0094] Table 3 Accuracy of each model on non-vpn traffic
[0095]
[0096] Table 4 Accuracy of each model on VPN traffic
[0097]
[0098] As shown in Table 3 and Table 4, the classification performance of the present invention was compared with FL, SMOTE oversampling, and undersampling under 12 specific traffic types. For traffic types with fewer samples (such as chat, email, P2P, VPN_email), the present invention is generally superior to the other three methods. It also has significant improvements in the classification of streaming media traffic, while the other methods have poor classification effects on streaming media traffic. These results indicate that adding an attention mechanism helps to alleviate the impact of data imbalance during the training process. To further evaluate the effect of adding an attention mechanism to the model of the present invention, we conducted additional experiments. The comparison results of the model of the present invention, the CNN_LSTM model, the CNN_LSTM_SA model (only adding self-attention mechanism), and the CNN_CBAM_LSTM model (only adding CBAM attention mechanism) using the same dataset under the same experimental environment are as Figure 7 shown.
[0099] As Figure 7 shown in the accuracy graph, the model of the present invention is superior to CNN_LSTM, CNN_LSTM_SA, and CNN_CBAM_LSTM in terms of classification accuracy. This result confirms that the model of the present invention not only achieves excellent performance on the training set but also proves the effectiveness of introducing an attention mechanism.
[0100] To verify the impact of the attention mechanism and the feature fusion module in the present invention, we conducted ablation experiments, and the results are shown in Table 5.
[0101] Table 5 Contributions of the attention mechanism and the feature fusion module (%)
[0102]
[0103] The experimental results show that adding the self-attention mechanism to the LSTM branch can improve the ability of the LSTM branch to extract statistical features. In the CNN branch, adding CBAM improves the extraction of the original traffic features. The present invention significantly improves the classification performance of encrypted network traffic by fusing the features of CNN_LSTM_SA and CNN_CBAM_LSTM.
[0104] On the other hand, an embodiment of the present invention also provides a network flow classification system that uses the above network flow classification method, including:
[0105] A network flow interception module for intercepting network flow data and preprocessing the network flow data;
[0106] A processing module for extracting and fusing features of the network flow data using an LSTM model, a CNN network, a self-attention module, and a convolutional attention module to obtain fused features;
[0107] A classification module for predicting and classifying the network flow data based on the fused features to obtain the corresponding network flow categories.
[0108] On the other hand, an embodiment of the present invention also provides a computer including at least one processor and a memory, where the memory stores a computer program, and the computer program is configured to be executed by the processor to implement the above network flow classification method.
[0109] On the other hand, an embodiment of the present invention also provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and the computer program can be executed by one or more processors to implement the above network flow classification method.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A network flow classification method, characterized in that: Including: Obtain network flow data, and import the network flow data into the LSTM model and the CNN network respectively to obtain the first feature and the second feature; Process the first feature using a self-attention module to obtain a first concatenated feature; Import the second feature into a convolutional attention module for processing to obtain a second concatenated feature; Concatenate the first concatenated feature and the second concatenated feature to obtain a fused feature, and then pass the fused feature through a prediction module for label prediction to obtain the corresponding network flow category.
2. A network flow classification method according to claim 1, wherein: After obtaining the network flow data, the network flow data needs to be processed, and the processing method includes: Convert the formats of all data in the network flow data into the same format, and then slice the network flow data to obtain multiple sessions; Extract the statistical feature information in each session, and delete the sessions with unnecessary information in the statistical feature information; Normalize the sizes of all sessions to the same number of bytes, and then import all sessions into the LSTM model and the CNN network.
3. A network flow classification method according to claim 1, wherein: The convolutional attention module includes a channel attention module and a spatial attention module; The method of importing the second feature into the convolutional attention module for processing includes: Import the second feature into the channel attention module to extract the first global feature information, and then concatenate the first global feature information with the second feature to obtain a first intermediate feature; Import the first intermediate feature into the spatial attention module to extract the second global feature information, and then concatenate the second global feature information with the first intermediate feature to obtain a second intermediate feature; Process the second intermediate feature through a fully connected layer to obtain a second concatenated feature.
4. A network flow classification method according to claim 3, wherein: The method for extracting the first global feature information includes: Perform global average pooling and global max pooling on the second feature respectively to obtain the corresponding first global average information and first global max information; Use a shared multi-layer perceptron to perform channel modeling on the first global average information and the first global max information respectively to generate the corresponding first channel weight vector and second channel weight vector; Perform modulo addition on the first channel weight vector and the second channel weight vector and then process them using a sigmoid activation function to obtain the first global feature information.
5. A network flow classification method according to claim 3, wherein: The method for extracting the second global feature information includes: Perform global average pooling and global max pooling on the second feature respectively to obtain the corresponding second global average information and second global max information; Perform modulo addition on the second global average information and the second global max information and then process them using a sigmoid activation function to obtain the first global feature information.
6. A network flow classification method according to claim 1, wherein: The shared multi-layer perceptron includes a rectified linear unit and two convolutional layers, and the two convolutional layers are respectively connected to the input end and the output end of the rectified linear unit.
7. A network flow classification method according to claim 1, wherein: The prediction module includes two fully connected layers and a Dropout module, and the two fully connected layers are respectively connected to the input end and the output end of the Dropout module.
8. A network flow classification system, characterized in that: The system uses a network flow classification method according to any one of claims 1-7, including: A network flow intercepting module, which is used to intercept network flow data and preprocess the network flow data; A processing module, which is used to extract and fuse features of the network flow data by using an LSTM model, a CNN network, a self-attention module and a convolutional attention module to obtain fused features; A classification module, which is used to predict and classify the network flow data according to the fused features to obtain corresponding network flow categories.
9. A computer, characterized in that: It includes at least one processor and a memory, and the memory stores a computer program, and the computer program is configured to be executed by the processor to implement a network flow classification method according to any one of claims 1-7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and the computer program can be executed by one or more processors to implement a network flow classification method according to any one of claims 1-7.