Encrypted traffic identification scheme based on improved CNN and ET-Bert in public testing environment

Through the improved fusion method of CNN and ET-Bert models, combined with local and global feature extraction, the problem of poor generalization of encrypted traffic recognition in the crowd test environment is solved, and higher classification accuracy and generalization capabilities are achieved, effectively preventing malicious attacks.

CN120075100APending Publication Date: 2025-05-30GUILIN UNIV OF ELECTRONIC TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510217962.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has poor generalization of the identification of encrypted traffic in the crowd test environment and cannot effectively adapt to new encryption strategies and complex environments.

Method used

Using the fusion method of the improved CNN (Pro-CNN) and ET-Bert model, the local patterns and features of network traffic are captured through CNN, and ET-Bert captures global dependencies and complex contextual relationships, combining feature extraction and fusion to realize multi-classification tasks of encrypted traffic.

Benefits of technology

It improves the accuracy and generalization ability of encrypted traffic classification, can more effectively identify complex encrypted traffic, prevent malicious attacks, and protect the security of software systems and user data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention relates to the field of network traffic analysis, deep learning and public testing environment protection, in particular to an encrypted traffic identification method based on an improved CNN (Pro-CNN) and ET-Bert in a public testing environment. According to the method, the CNN is combined for local feature extraction and the Transform architecture in the ET-Bert is combined for global dependency capture, so that higher classification accuracy and stronger generalization ability are provided, complex encrypted traffic classification tasks are effectively handled, hostile attacks are detected and prevented in time, and the security of a software system and user data is protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of network traffic analysis, deep learning, and crowdsourcing environmental protection, and specifically to a method for identifying encrypted traffic based on improved CNN and ET-Bert in a crowdsourcing environment. Background Art

[0002] Crowdsourced Testing is a software testing method that utilizes a large number of users or workers on the Internet to identify and report errors, vulnerabilities, or problems in software products. Compared with traditional software testing, crowdsourced testing can provide broader test coverage, faster feedback speed, and lower costs.

[0003] However, problems such as data leakage, privacy security, intellectual property protection, and malicious attacks may occur during the crowdsourcing process. Especially the malicious attack problem, which is an important security issue that may be encountered during the crowdsourcing process. In crowdsourcing, due to the diverse backgrounds and motivations of participants, there are some who may take advantage of the opportunities in crowdsourcing to launch malicious attacks on software systems. This may cause the interruption of crowdsourcing services, lead to the theft or leakage of sensitive data, and more seriously, cause the collapse of the crowdsourcing system, resulting in serious property losses. Therefore, it is very necessary to detect malicious attacks by timely discovering abnormal traffic generated during possible malicious attack processes in the crowdsourcing environment, so as to quickly detect and respond to potential malicious attacks. Thus, avoiding the risks brought by malicious attacks and protecting the security of the crowdsourcing software system and user data is very necessary.

[0004] Existing traffic classification methods mainly include classification methods based on preset rules (such as fixed port numbers), which are gradually becoming ineffective due to the fact that many popular applications allow users to customize communication ports; deep packet inspection (DPI) classification methods based on load characteristics cannot identify cryptographic protocols; classification methods based on statistical features require a large amount of statistical work; traditional machine learning classification methods (such as SVM, RF) require manual extraction of traffic features; deep learning classification methods (such as DNN, CNN, RNN) cannot well adapt to new environments or invisible encryption policies due to the rapid development of encryption technology. Existing methods using large models (such as ET-Bert) to learn deep contextualized datagram-level representations of traffic from large-scale unlabeled encrypted traffic data are a solution idea. Summary of the Invention

[0005] The present invention proposes an encrypted traffic recognition method that fuses improved CNN (Pro-CNN) and ET-Bert. This solution aims to solve the problem of poor generalization of traditional deep learning methods in specific encrypted traffic classification (especially crowdsourcing environment traffic). By leveraging the excellent performance of CNN in processing spatial structure data (such as images), and treating network traffic data as a two-dimensional data structure (such as a traffic feature matrix), CNN can effectively capture local patterns and features in the data, such as traffic patterns within a specific time window. At the same time, the self-attention mechanism of the Transformer architecture in the ET-Bert model is used to capture global dependencies in the traffic sequence and understand the complex context relationships in network traffic, thereby improving the traffic classification accuracy of the ET-Bert model.

[0006] The object of the present invention is achieved through the following technical solutions: Step 1: First, obtain the effective payload information in the encrypted network traffic in the crowdsourcing environment; Step 2: Process the effective payload information in the network traffic to make it conform to the inputs of the improved CNN (Pro-CNN) and ET-Bert models respectively. Specifically: For the Pro-CNN model, convert the payload data into a two-dimensional grayscale image (such as a 28x28 grayscale image); for the ET-Bert model, first split the network traffic into multiple packet sequences, each sequence representing a flow, then convert the packets into byte streams, and further convert them into hexadecimal strings. Then, serialize the packet sequences represented by the hexadecimal strings into text sequences, and then convert the text sequences into Token sequences, where each Token represents a byte or a combination of multiple bytes. At the same time, add position encoding (Position Embedding) and segment encoding (Segment Embedding) to each Token, and input the processed Token sequence into the ET-BERT model for feature extraction; Step 3: Feature extraction of the Pro-CNN model and the ET-BERT model. For the Pro-CNN model, input the two-dimensional grayscale image obtained in Step 2 into the Pro-CNN model for feature extraction. Among them, the Pro-CNN model uses multiple convolutional layers to extract the local spatial features of the network traffic, and at the same time uses multiple corresponding pooling operations to reduce the dimension of the features, so as to improve the robustness of the features. The specific structure of the Pro-CNN model is as Figure 1As shown, the features output by the model are determined by the last fully connected layer, and its format is (BatchSize, 1568 (32x7x7)); Feature extraction of the ET-BERT model. For the ET-BERT model, the Token sequence processed in Step 2 is input into the ET-BERT model for feature extraction. The ET-BERT model, like BERT, is composed of multiple Transformer bidirectional encoders. Thanks to the self-attention mechanism of the Transformer architecture, it can capture the global dependencies in the sequence and can efficiently process long sequence data. Therefore, ET-BERT can be used to understand the complex context relationships in network traffic and thus better capture the global features in network traffic. Its specific structure is as Figure 2 shown, so the features output by the model are determined by the last Transformer encoder, and its format is (BatchSize, 768); Step 4: Feature fusion. The output features (BatchSize, 1568) of the last fully connected layer of the Ptr-CNN model obtained in Step 3 and the output features (BatchSize, 768) of the last Transformer encoder of the ET-BERT model obtained in Step 4 are concatenated in the feature dimension to form fused features (BatchSize, 2336 (1568 + 768)); Step 5: Traffic classification. The fused features of the encrypted network traffic in the crowdsourcing testing environment obtained in Step 5 are used to implement the multi-classification task of the encrypted network traffic in the crowdsourcing testing environment by adding a Softmax activation function; Description of the Drawings

[0007] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0008] Figure 1 It is a schematic structural diagram of the Ptr-CNN model of the present invention. Figure 2 It is a schematic structural diagram of the ET-BERT model of the present invention. Figure 3 It is a flow chart of the solution of the present invention Specific Implementation

[0009] The following clearly and completely describes the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments, which does not constitute a limitation to the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0010] Please refer to Figure 1 , the PtrCNN structure starts with two alternating convolutional layers and pooling layers. The C1 layer is the first convolutional layer. Regarding the input as an image, it is first padded so that the size of the convolutional output image after convolution is equal to the size of the input image. Then, 16×3×3 filters - filters (3D convolutional kernels) are selected and added to the bias term. Finally, 32@28×28 feature images are obtained through the ReLU activation function. The S2 layer is the first pooling layer. The output image of the C1 layer is processed by max sampling. The pooling window is set to 2×2, and 32@14×14 feature images are obtained. The C3 layer is the second convolutional layer. The C3 layer uses 32 5×5 convolutional kernels, and the data processing method is the same as that of the C1 layer. The resulting feature images are 32@14×14. The S4 layer is the second pooling layer, obtaining 32@7×7 feature images, and finally the features of the traffic data are output through a fully connected layer (BatchSize, 1568 (32×7×7)). The formula for updating the weight parameters of each layer in backpropagation is: where w (l) : represents the weight matrix of the l-th layer, Δw (l) is the gradient of the loss function with respect to the weight w (l) , α: learning rate, used to control the step size of weight update. λ is an adjustable regularization parameter, m represents the number of training samples; the formula for updating the bias parameters of each layer in backpropagation is: where b (l) represents the bias value of the l-th layer, Δb (l) is the gradient of the loss function with respect to the bias b (l) .

[0011] Please refer to Figure 2, The ET-BERT model extracts effective features through a pre-trained Transformer model (i.e., the BERT model) for the classification task of encrypted traffic. Data preprocessing includes collecting raw network traffic data, cleaning the data, extracting the payload, and converting it into a Token sequence in text format. After the model input undergoes position encoding and segment encoding, feature extraction is performed through multiple Transformer encoder layers, and each encoder layer contains a multi-head self-attention mechanism and a feed-forward neural network. The features output by the last Transformer encoder are pooled to obtain a feature vector with a shape of (BatchSize, 768).

[0012] Performing the classification of encrypted traffic in the crowdsourcing testing environment is to detect and prevent malicious attacks in a timely manner and protect the security of software systems and user data. The proposed method of fusing the improved CNN (Pro-CNN) and ET-Bert models provides higher classification accuracy and stronger generalization ability by combining the local feature extraction of CNN and the global dependency capture of ET-Bert, effectively dealing with complex encrypted traffic classification tasks. To achieve the above object, the technical solution of the present invention is realized through the following steps: Step 1: First, obtain the effective payload information (payload) in the encrypted network traffic in the crowdsourcing testing environment; Step 2: Process the effective payload information in the network traffic so that it respectively conforms to the inputs of the improved CNN (Pro-CNN) and ET-Bert models. Specifically: For the Pro-CNN model, convert the payload data into a two-dimensional grayscale image (such as a 28x28 grayscale image); for the ET-Bert model, first split the network traffic into multiple packet sequences, each sequence representing a flow, then convert the packets into byte streams, and further convert them into hexadecimal strings. Then, serialize the packet sequences represented by the hexadecimal strings into text sequences, and then convert the text sequences into Token sequences. Each Token represents a byte or a combination of multiple bytes. At the same time, add position encoding (Position Embedding) and segment encoding (Segment Embedding) to each Token. Input the processed Token sequence into the ET-BERT model for feature extraction; Step 3: Feature extraction of the Pro-CNN model and the ET-BERT model. For the Pro-CNN model, the two-dimensional grayscale value image obtained in Step 2 is input into the Pro-CNN model for feature extraction. The Pro-CNN model uses multiple convolutional layers to extract the local spatial features of network traffic, and at the same time uses multiple corresponding pooling operations to reduce the dimension of the features, so as to improve the robustness of the features. The specific structure of the Pro-CNN model is as shown in Figure 1 . Therefore, the features output by this model are determined by the last fully connected layer, and its format is (BatchSize, 1568 (32x7x7)); Feature extraction of the ET-BERT model. For the ET-BERT model, the Token sequence processed in Step 2 is input into the ET-BERT model for feature extraction. The ET-BERT model, like BERT, is composed of multiple Transformer bidirectional encoders. Thanks to the self-attention mechanism of the Transformer architecture, it can capture the global dependencies in the sequence and can efficiently process long sequence data. Therefore, ET-BERT can be used to understand the complex context relationships in network traffic, so as to better capture the global features in network traffic. The specific structure is as shown in Figure 2 . Therefore, the features output by this model are determined by the last Transformer encoder, and its format is (BatchSize, 768); Step 4: Feature fusion. The output features (BatchSize, 1568) of the last fully connected layer of the Ptr-CNN model obtained in Step 3 and the output features (BatchSize, 768) of the last Transformer encoder of the ET-BERT model obtained in Step 4 are concatenated in the feature dimension to form fused features (BatchSize, 2336 (1568 + 768)); Step 5: Traffic classification. The fused features of the encrypted network traffic in the crowdsourcing environment obtained in Step 5 are used to implement the multi-classification task of the encrypted network traffic in the crowdsourcing environment by adding a Softmax activation function.

Claims

1. An encrypted traffic identification method based on improved CNN and ET-Bert in a crowd testing environment, characterized in that: The techniques include: The first step is to obtain the effective load information (payload) in the encrypted network traffic in the crowd testing environment.

2. The encrypted traffic identification solution based on improved CNN and ET-Bert in a crowd testing environment according to claim 1 is characterized in that: The process of step 1 is specifically as follows: The payload information in the network traffic is processed to make it conform to the input of the improved CNN (Pro-CNN) and ET-Bert models respectively. Specifically: for the Pro-CNN model, the payload data is converted into a two-dimensional grayscale image (e.g., a 28x28 grayscale image); and for the ET-Bert model, the network traffic is first divided into multiple data packet sequences, each sequence represents a flow, and then the data packets are converted into byte streams, and further converted into hexadecimal strings, and then the data packets represented by the hexadecimal strings are serialized into text sequences, and then the text sequences are converted into Token sequences, each Token represents a combination of one byte or multiple bytes, and each Token is added with position encoding (PositionEmbedding) and segment encoding (Segment Embedding) The processed Token sequence is input into the ET-BERT model for feature extraction.

3. The encrypted traffic identification solution based on improved CNN and ET-Bert in a crowd testing environment according to claim 1 is characterized in that: The process of step 2 is specifically as follows: Feature extraction of Pro-CNN model and ET-BERT model. For the Pro-CNN model, the two-dimensional grayscale image obtained in step 2 is input into the Pro-CNN model for feature extraction, where the Pro-CNN model uses multiple convolutional layers to extract local spatial features of network traffic, and uses multiple corresponding pooling operations to reduce the dimension of the features to improve the robustness of the features. The specific structure of the Pro-CNN model is shown in Figure 1, so the features output by the model are determined by the last fully connected layer, and its format is (BatchSize, 1568 (32x7x7)); feature extraction of the ET-BERT model. For the ET-BERT model, the Token sequence processed in step 2 is input into the ET-BERT model for feature extraction. The ET-BERT model, like BERT, is composed of multiple Transformer bidirectional encoders. Thanks to the self-attention mechanism of the Transformer architecture, it can capture the global dependencies in the sequence and efficiently process long sequence data. Therefore, ET-BERT can be used to understand the complex contextual relationships in network traffic, thereby better capturing the global features in network traffic. Its specific structure is shown in Figure 2, so the features output by the model are determined by the last Transformer encoder, and its format is (BatchSize, 768).

4. The encrypted traffic identification solution based on improved CNN and ET-Bert in a crowd testing environment according to claim 1 is characterized in that: The process of step 3 is specifically as follows: Feature fusion. The output features of the last fully connected layer of the Ptr-CNN model obtained in step 3 (BatchSize, 1568) and the output features of the last Transformer encoder of the ET-BERT model obtained in step 4 (BatchSize, 768) are concatenated in the feature dimension to form a fused feature (BatchSize, 2336 (1568 + 768)).

5. The encrypted traffic identification solution based on improved CNN and ET-Bert in a crowd testing environment according to claim 1 is characterized in that: The process of step 4 is specifically as follows: Traffic classification. The fusion features of the encrypted network traffic in the crowd-testing environment obtained in step 5 are used to implement the multi-classification task of the encrypted network traffic in the crowd-testing environment by adding the Softmax activation function.