Packet loss resistant encrypted traffic classification method and system
By using mask reconstruction self-supervised learning method in encrypted traffic classification, the interference problem of packet loss on encrypted traffic recognition is solved, and the robustness and accuracy of the classification are improved.
Patent Information
- Application Number
- CN202510150624.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-30
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional encrypted traffic classification methods have problems such as insufficient robustness and reduced classification accuracy in packet loss caused by abnormal network environment and device performance limitations.
The anti-packet loss encrypted traffic classification method based on mask reconstruction is adopted. By performing packet loss detection, byte block embedding conversion, partial byte block mask and joint pre-trained stream encoder and stream decoder on network traffic, and finally, a fine-tuned prediction model is used for encrypted traffic classification.
It improves the robustness of encrypted traffic classification, effectively deals with the data sequence disorder caused by packet loss, and significantly improves classification performance and task adaptability.
Smart Images

Figure CN120145134A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security and encrypted traffic analysis, and particularly relates to a method and system for classifying encrypted traffic resistant to packet loss. Background Art
[0002] During the process of network data transmission, the state of the network link through which the data flows is usually not real-time controllable, thus inevitably resulting in the phenomenon of packet loss. Packet loss in each link of the network link may limit the integrity of traffic data, thereby reducing the practicality of traditional encrypted traffic classification methods. To address this problem, existing research mainly focuses on how to more effectively use the existing packet information for classification. These methods can be summarized into three types: methods based on packet loss rate perception, methods based on packet recovery, and methods based on packet classification. Among them, the methods based on packet loss rate perception additionally consider the current network state, the methods based on packet recovery infer the lost data through the existing data, and the methods based on packet classification only use the existing data for classification.
[0003] The methods based on packet loss rate perception attempt to fuse the current network state (such as packet loss rate) as a feature into other features, enabling the model to make adaptive adjustments according to the current network conditions. Due to the negative impacts of time-dependent factors such as latency and packet loss in the real network environment on the intrusion detection system, researchers proposed a multi-factor intrusion detection system, which uses relative spectral scaling feature selection for feature collection and uses the proposed cross-layer adaptive convolutional neural network to classify these features. The experimental results show that this method has higher classification accuracy and lower false alarm rate. This method takes the network condition as one of the features and can adapt to different network environments compared with other methods. However, the lost packets will still cause changes in the flow distribution and reduction of available information, and it is still difficult to achieve good classification results based on only the existing data.
[0004] The main idea of the methods based on packet recovery is to identify the position of packet loss and fill or restore the numerical values with a certain strategy. Currently, there are three packet loss mitigation strategies: zero-value filling, default-value filling, and adjacent-value filling. When the detection model identifies a loss, the three predefined filling values are respectively used as the size of the lost packet, thereby reducing the impact of packet loss on the encrypted traffic data. The advantage of the methods based on lost packet recovery is that they can recover additional information to a certain extent, and their disadvantage is that the fixed numerical filling does not conform to the real data distribution, and there is a large gap in the classification effect compared with the classification based on the original traffic data.
[0005] The method based on packet classification believes that packet loss does not change the structure of a single packet. Therefore, the existing packet information can be utilized to mitigate the impact of packet loss through packet-level classification. The advantage of this method is that the classification performance is almost independent of the network environment, and the model has stronger robustness. However, the information contained in a single packet is too little, and the model effect is less than satisfactory. In addition, under the trend of full traffic encryption, the visible information in the packet is even less, and the obfuscated byte data will further reduce the classification performance of this method. Therefore, there is an urgent need to propose a technical solution for encrypted traffic classification that resists packet loss. Summary of the Invention
[0006] The object of the present invention is to propose an encrypted traffic classification scheme for resisting packet loss based on masked reconstruction self-supervised learning, aiming to solve the interference problem of packet loss caused by abnormal network environment and device performance limitations on encrypted traffic recognition, thereby effectively improving the robustness of encrypted traffic classification and overcoming the negative impact of packet loss on the integrity and classification accuracy of streaming data.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] An encrypted traffic classification method for resisting packet loss, comprising the following steps:
[0009] 1) Perform packet loss detection on network traffic and extract flow byte vectors;
[0010] 2) Divide the flow byte vectors into byte blocks, perform byte block embedding transformation to generate an original flow byte block embedding sequence;
[0011] 3) Perform partial byte block masking on the original flow byte block embedding sequence, and use the remaining unmasked byte blocks to jointly pre-train the flow encoder and the flow decoder, where the flow encoder extracts the deep features of the flow byte block embedding sequence, and the flow decoder reconstructs the flow byte block embedding sequence according to the extracted deep features;
[0012] 4) Combine the pre-trained flow encoder and the classifier to form a prediction model, fine-tune the prediction model using the labeled traffic for a specific classification task, and use the fine-tuned prediction model to classify encrypted traffic.
[0013] Further, before performing packet loss detection in step 1), the original network traffic is segmented into flows according to the five-tuple <source IP, destination IP, source port, destination port, protocol>, and each flow corresponds to a traffic type.
[0014] Further, the method for packet loss detection in step 1) is: recover the out-of-order packets caused by sequence offset by calculating the ACK and SEQ fields of TCP packet by packet, remove the redundant packets caused by sequence repetition, and mark the positions of the lost packets.
[0015] Further, the method for extracting the flow byte vector in step 1) is: extract the headers of each data packet and the first several bytes of the payload, and splice them as the flow byte vector; wherein, delete the Ethernet headers of each data packet, and randomize the IP addresses and port numbers.
[0016] Further, in step 2), byte block embedding transformation is performed through a byte block embedder, and the steps include: map each byte block through a trainable linear transformation layer to obtain the embedding representation of the byte block; then add absolute position encoding to make the embedding representation of the byte block contain position information, and obtain the original flow byte block embedding sequence.
[0017] Further, the method for performing partial byte block masking on the flow byte block embedding sequence in step 3) includes byte-level masking and packet-level masking; byte-level masking randomly masks byte blocks to learn context relationships and packet-level features; packet-level masking randomly masks the entire data packet to simulate packet loss and recover the features of the lost data packet.
[0018] Further, in step 3), the flow encoder and the flow decoder have the same structure, both consisting of multiple Transformer blocks, and each Transformer block consists of a multi-head self-attention layer, a feed-forward network layer, and two layer normalization algorithms.
[0019] Further, in step 3), the mean squared error loss is used to calculate the reconstruction loss between the reconstructed flow byte block embedding sequence and the original flow byte block embedding sequence.
[0020] Further, in step 4), the labeled traffic used for fine-tuning is all available byte blocks in the flow byte block embedding sequence except for the lost data packets, and the cross-entropy loss is used to calculate the difference between the predicted label and the true label.
[0021] An anti-packet-loss encrypted traffic classification system, comprising:
[0022] A traffic characterization module, used for detecting packet loss in network traffic and extracting flow byte vectors;
[0023] A byte embedding module, used for dividing the flow byte vector into byte blocks, performing byte block embedding transformation, and generating an original flow byte block embedding sequence;
[0024] A mask pre-training module, used for performing partial byte block masking on the original flow byte block embedding sequence, and jointly pre-training the flow encoder and the flow decoder using the remaining unmasked byte blocks, wherein the flow encoder extracts deep features of the flow byte block embedding sequence, and the flow decoder reconstructs the flow byte block embedding sequence according to the extracted deep features;
[0025] A task fine-tuning module for fine-tuning a prediction model composed of a pre-trained flow encoder and a classifier using labeled traffic for a specific classification task to adapt to classifying encrypted traffic for a specific classification task.
[0026] The technical effects achieved by the present invention:
[0027] 1. The present invention adopts a masking mechanism. Compared with the traditional fixed-value filling method based on empirical data, it can more naturally simulate the impact of packet loss on the flow sequence and restore the change in the flow data distribution caused by packet loss through a reconstruction mechanism. This method improves the robustness of encrypted traffic classification and effectively deals with the disorder of the data sequence caused by packet loss.
[0028] 2. The present invention is based on pre-training technology. Different from the method of initializing parameters randomly in other methods, it can capture and learn the general representation of encrypted traffic during the learning stage of large-scale unlabeled data. This method enables the model to have a certain ability to understand and process traffic data before fine-tuning for specific tasks. Therefore, in terms of classification performance, it is significantly better than the other two types of methods and has stronger task adaptability and performance.
[0029] 3. The present invention designs two masking pre-training tasks for traffic data, namely byte-level masking and packet-level masking. This design is different from common pre-training methods. By forcing the model to restore the original flow byte sequence at the packet level and flow level with multiple granularities, it enhances the model's ability to represent and discriminate the characteristics of encrypted traffic. Through this multi-level learning, the model can more effectively adapt to the change in data distribution brought by the packet loss scenario, reduce the interference of packet loss on the model's feature extraction, and achieve robust recognition in various network environments. Description of the Drawings
[0030] Figure 1 It is an architecture diagram of an anti-packet-loss encrypted traffic classification method in an embodiment.
[0031] Figure 2 It is a processing flow chart of an anti-packet-loss encrypted traffic classification system in an embodiment.
[0032] Figure 3 It is a test performance curve graph under different packet loss environments. Detailed Embodiments
[0033] To make the technical features, advantages, or technical effects in the above technical solutions of the present invention more obvious and understandable, the following will be described in detail in conjunction with the accompanying drawings.
[0034] The embodiment of the present invention specifically provides an anti-packet-loss encrypted traffic classification method and system. The architecture of this method is as Figure 1As shown in the figure. For the changes in the encrypted flow sequence caused by packet loss, this method performs flow byte embedding characterization based on TCP semantic enhancement, restores the out-of-order phenomenon caused by sequence offset, repetition, etc. through packet reordering, and further locates the packet loss position. Based on the masked autoencoder, this method proposes two masked pre-training tasks: byte-level masking and packet-level masking. Byte-level masking is used to learn the context relationship between byte blocks, and packet-level masking is used to simulate the packet loss process. Through the "masking-reconstruction" mechanism, the model can learn the distribution characteristics of missing data and perform pre-training on large-scale unlabeled encrypted traffic data, so as to extract robust encrypted traffic representations. This system performs the following steps of this method through four modules: a traffic representation module, a byte embedding module, a masked pre-training module, and a task fine-tuning module, as shown in Figure 2 the figure.
[0035] S1: Traffic Representation
[0036] Traffic representation analyzes the original traffic and extracts features through three steps: flow division, packet loss detection, and byte representation.
[0037] S1-1: First, the original network traffic is sliced into flows according to the <source IP, destination IP, source port, destination port, protocol> five-tuple. Each flow corresponds to a specific traffic type (i.e., label), and these flows constitute the basic unit of the classification task.
[0038] S1-2: Perform packet loss detection on the divided flows, and use the TCP semantic enhancement mechanism to solve the problem of flow sequence changes caused by packet loss. For any router in the network, packet loss at different positions in the transmission link may cause different packet sequence change phenomena. When packet loss occurs upstream of the router, the TCP retransmission mechanism may cause the router to receive out-of-order packets, resulting in sequence offset; when packet loss occurs downstream of the router, the upstream router may receive duplicate packets, resulting in packet sequence repetition; when the router itself experiences packet loss, the packets may not be recoverable, resulting in packet sequence holes. For packets that are out-of-order due to sequence offset, the original packet sequence is restored by calculating the ACK and SEQ fields of TCP packet by packet; for redundant packets due to sequence repetition, these packets are removed by de-duplication using the ACK field; for packets that cannot be restored due to sequence holes, the position of packet loss is retained in the flow to obtain the processed TCP semantic enhanced flow. It should be noted that due to the existence of the cumulative acknowledgment mechanism, it is very difficult to accurately calculate the loss of multiple consecutive packets. Therefore, this method regards each located packet loss position as only a single packet loss and restores the overall distribution of the lost data in subsequent steps.
[0039] S1-3: For the packets that have passed the packet loss detection and undergone TCP semantic enhancement, the method extracts the first H original bytes of the packet header and the payload respectively, and concatenates them as the flow byte vector.
[0040] The packet header contains information related to protocol control and interaction details, while the payload carries the transmission of actual service data. These two types of fields describe the communication characteristics of the packet from different perspectives of the network layer, transport layer, and application layer, and together constitute a valid representation of the packet. Although the payload cannot be understood by humans due to encryption, the pseudo-randomness of the encryption algorithm still makes it possible to learn the differences in the payload byte distribution. In addition, to avoid the privacy leakage of plaintext fields such as IP addresses and ports in the packet header and the resulting model memorization problem, that is, the model classifies packets based on memorizing IP addresses rather than learning the potential traffic representation, this method deletes the Ethernet header of each packet and randomizes the IP address and port number. Finally, the combined header bytes and payload bytes of the first k packets in the flow form a flow byte vector of length k×2H. Formally, for a network flow x, its flow byte vector V x can be defined as:
[0041] V x = [packet 1 , packet 2 , …, packet k
[0042]
[0043] where, if the i-th packet is identified as a lost packet, then packet i = 0, that is, it is replaced with all-zero elements.
[0044] S2: Byte Embedding
[0045] Based on the flow byte vector representation obtained in S1, the flow byte vector is divided into byte blocks with a higher degree of information aggregation, and then further embedded into the implicit feature space, including two steps: byte block generation and byte block embedding.
[0046] S2-1: Considering that the information contained in a single byte is limited and not very representative, in this step, consecutive n bytes in the flow byte vector V x are combined into byte blocks, and then converted into a sequence of byte blocks of the same size.
[0047] Specifically, a stream byte vector of length k×2H is divided into (k×2H) / / N byte blocks (" / / " represents integer division, that is, the fractional part of the result is discarded and only the integer part is retained), and each byte block contains N bytes. The sequence of byte blocks can significantly reduce the length of the model input sequence and reduce the computational overhead while aggregating features.
[0048] S2-2: The main purpose of byte block embedding is to convert the extracted sequence of byte blocks into an input form acceptable to the flow autoencoder, which is completed by a byte block embedder.
[0049] First, each byte block is mapped to a D-dimensional vector space using a trainable linear transformation layer:
[0050]
[0051] where, represents the stream byte block embedding sequence, represents the linear mapping layer.
[0052] Then, in order to preserve the packet order and arrival relationship in the encrypted stream, triangle-based absolute position encoding will also be added to the byte block embedding. After the transformation, the stream byte block embedding sequence can be expressed as:
[0053]
[0054] S3: Masked Pre-training
[0055] Context learning and packet loss simulation are carried out by designing two traffic-specific masked pre-training tasks, and a large-scale unlabeled encrypted traffic is used to pre-train the flow autoencoder to extract traffic features.
[0056] S3-1: To effectively utilize a large amount of unlabeled data to learn the encoding and reconstruction of flow representations, two types of pre-training tasks are designed: byte-level masking and packet-level masking. Among them, byte-level masking is used for context relationship learning of byte blocks and packet-level feature extraction, while packet-level masking aims to simulate packet loss by randomly masking packets and then recover the missing data based on context information.
[0057] Byte-level masking randomly masks byte blocks in the stream, thus forcing the model to mine context byte information and packet-level features. Specifically, a random masking mechanism is used to learn latent features by masking a higher proportion of byte blocks. The purpose of this design is to reduce redundant information between byte blocks, increase the challenge of the reconstruction task, and thus prompt the model to learn the global features of the traffic.
[0058] Packet-level masking enhances the model's understanding of the relationships between packets from the perspective of flow-level feature extraction by randomly masking entire packets. By simulating packet loss and using a reconstruction mechanism to learn the features of the lost packets, the model learns how to adapt to the data distribution changes caused by sequence holes.
[0059] The above two types of masking tasks act together on the byte-block embeddings. After byte-level masking with a ratio of rb and packet-level masking with a ratio of rp, the remaining unmasked byte-blocks (with a ratio of 1 - rb - rp) are fed into the flow decoder for the pre-training process.
[0060] S3-2: Pre-train the flow autoencoder-flow decoder using a large amount of unlabeled encrypted traffic
[0061] The goal of the flow encoder is to learn flow-level latent features from the sequence of flow byte-block embeddings. The flow encoder consists of L stacked Transformer blocks, each block composed of a multi-head self-attention layer (MSA), a feed-forward network layer (FFN), and two layer normalization operations. The multi-head self-attention layer uses multiple attention heads to model the flow sequence from different perspectives, thereby capturing complex associations and dependencies in the flow sequence. Given a sequence of flow byte-block embeddings P x , its multi-head self-attention calculation process can be formally described as follows:
[0062]
[0063] MSA(P x ) = Concat(head 1 , head 2 , …, head h )W O
[0064] where head i = Attention(P x W i Q , P x W i K , P x W i V ), i ∈ [1, h]
[0065] Among them, represent the weight matrices of the query vector Q, key vector K, and value vector V in the attention mechanism respectively, h represents the number of attention heads, and D k = D / / h represents the dimension of a single attention head. Represents the weight matrix of the linear transformation layer after concatenating multiple attention heads. The output of the multi-head self-attention layer undergoes residual connection and layer normalization to alleviate the problems of vanishing gradients and exploding gradients. Subsequently, the feed-forward network layer further extracts rich features in the flow sequence through the following non-linear transformation:
[0066] FFN(P x ) = GELU(P x W 1 + b 1 )W 2 + b 2
[0067] Where W 1 , b 1 , W 2 , b 2 Represent the weight matrix and bias vector respectively, and GELU(.) represents the Gaussian Error Linear Units (GELU) activation function. The output results of the feed-forward network layer and the multi-head self-attention layer are added together to form the output of the first Transformer block. Therefore, for a flow encoder with L layers, the feature extraction process of the l-th layer where l ∈ [1, L] can be expressed as:
[0068]
[0069] Through the feature learning of multiple Transformer blocks, the flow encoder can extract the deep features of the input sequence P x and generate the byte block encoding e x , providing support for the subsequent flow decoder to restore the flow byte block embedding sequence.
[0070] The main task of the flow decoder is to reconstruct the original flow byte block embedding sequence based on the deep features extracted by the flow encoder. Similar to the flow encoder, the flow decoder is also stacked by L' Transformer blocks with the same structure. The input of the flow decoder includes the output of the last layer of the flow encoder (i.e., the byte block encoding e x ) and the masked byte block placeholder. After the operations of multiple multi-head self-attention layers and feed-forward network layers, the flow decoder finally generates the reconstructed flow byte block embedding sequence. For the flow x, the calculation process of its reconstructed flow byte block embedding sequence P x can be expressed as the following formula:
[0071]
[0072] Where represents the output of L' Transformer blocks, W rec and brec They respectively represent the linear transformation weights and bias vectors that convert the output results into the same dimension as the original stream byte block embedding sequence.
[0073] The flow encoder and the flow decoder are jointly trained so that the flow encoder can generate more representative traffic features, and at the same time, the decoder can more accurately reconstruct the original encrypted stream. The main purpose of the masked pre-training is to accurately reconstruct these masked byte blocks. Therefore, this method uses the Mean Square Error (MSE) loss to calculate the reconstructed stream byte block embedding sequence and the original stream byte block embedding sequence The reconstruction loss between them is as follows:
[0074]
[0075] By minimizing the MSE, the byte block embedder, the flow encoder, and the flow decoder can be optimized simultaneously, and finally, a robust representation of the encrypted traffic can be extracted.
[0076] S4: Task fine-tuning
[0077] After masked pre-training, the flow decoder is replaced with an MLP classifier, and a new prediction model is composed of the flow autoencoder and the classifier. The labeled traffic is used to fine-tune this model so that it can adapt to a specific encrypted traffic classification task to obtain the prediction distributions of different classes. This step does not use the masking mechanism, but uses all available byte blocks in the flow byte block embedding sequence except for the lost data packets as the input. Among them, the class with the highest prediction probability will be used as the prediction label of the current stream The goal of fine-tuning is to minimize the loss of the model on the classification task. Therefore, the Cross Entropy (CE) loss is used to calculate the difference between the prediction label and the true label y:
[0078]
[0079] By minimizing the cross-entropy loss, the method can not only learn a robust traffic feature representation but also optimize for specific classification tasks. The fine-tuned flow encoder can well adapt to various downstream classification tasks.
[0080] Experimental tests:
[0081] The method of the present invention (NetMAE) and 11 existing methods are used for performance testing on 4 test sets. The test results are shown in Table 1 and Figure 3 .
[0082] Table 1 Comparative experimental results in a non-packet-loss environment
[0083]
[0084] From Table 1 and Figure 3 the test results shown, it can be seen that the encryption traffic classification and recognition performance of the method of the present invention on four datasets is significantly better than the existing 11 methods.
[0085] Although the present invention has been disclosed above by way of examples, it is not intended to limit the present invention. Any appropriate modifications or equivalent replacements made by those of ordinary skill in the art to the technical solutions of the present invention shall be covered within the protection scope of the present invention. The protection scope of the present invention shall be subject to what is defined by the claims.
Claims
1. A packet loss resistant encrypted traffic classification method, characterized in that: The following steps are involved: 1) Perform packet loss detection on network traffic and extract stream byte vectors; 2) Divide the stream byte vector into byte blocks, perform byte block embedding conversion, and generate an original stream byte block embedding sequence; 3) performing partial byte block masking on the original stream byte block embedding sequence, and using the remaining unmasked byte blocks to jointly pre-train the stream encoder and the stream decoder, wherein the stream encoder extracts deep features of the stream byte block embedding sequence, and the stream decoder reconstructs the stream byte block embedding sequence according to the extracted deep features; 4) The pre-trained stream encoder and classifier are combined into a prediction model, the prediction model is fine-tuned using the labeled traffic for the specific classification task, and the encrypted traffic is classified using the fine-tuned prediction model.
2. The method according to claim 1, characterized in that Before performing packet loss detection in step 1), the original network traffic is divided into flows according to the five-tuple <source IP, destination IP, source port, destination port, protocol>, and each flow corresponds to a traffic type.
3. The method according to claim 1, characterized in that The method for packet loss detection in step 1) is: recovering out-of-order data packets caused by sequence offset by calculating the ACK and SEQ fields of TCP packet by packet, deduplicating redundant data packets caused by sequence duplication, and marking the position of lost data packets.
4. The method according to claim 1, characterized in that The method for extracting the stream byte vector in step 1) is: extracting the header of each data packet and the first several bytes of the payload, and splicing them as the stream byte vector; wherein, the Ethernet header of each data packet is deleted, and the IP address and port number are randomized.
5. The method according to claim 1, characterized in that In step 2), byte block embedding conversion is performed through a byte block embedder, and the steps include: mapping each byte block through a trainable linear transformation layer to obtain an embedding representation of the byte block; then adding absolute position encoding so that the embedding representation of the byte block contains position information, and obtaining the original stream byte block embedding sequence.
6. The method according to claim 1, characterized in that The method of partially masking the stream byte block embedding sequence in step 3) includes byte-level masking and packet-level masking; the byte-level masking learns contextual relationships and packet-level features by randomly masking byte blocks; the packet-level masking simulates packet loss and recovers lost packet features by randomly masking the entire packet.
7. The method according to claim 1, characterized in that In step 3), the stream encoder and stream decoder have the same structure, both consisting of multiple Transformer blocks. Each Transformer block consists of a multi-head self-attention layer, a feedforward network layer, and two layer normalization algorithms.
8. The method according to claim 1, characterized in that In step 3), the mean square error loss is used to calculate the reconstruction loss between the reconstructed stream byte block embedding sequence and the original stream byte block embedding sequence.
9. The method according to claim 1, characterized in that The annotated traffic used for fine-tuning in step 4) is all available byte blocks in the stream byte block embedding sequence except for the lost data packets, and the cross entropy loss is used to calculate the difference between the predicted label and the true label.
10. An encrypted traffic classification system for anti-packet loss, implementing the method according to any one of claims 1 to 9, characterized in that: include: Traffic characterization module, used to detect packet loss on network traffic and extract flow byte vectors; A byte embedding module is used to divide the stream byte vector into byte blocks, perform byte block embedding conversion, and generate an original stream byte block embedding sequence; a mask pre-training module, for performing partial byte block masking on the original stream byte block embedding sequence, and using the remaining unmasked byte blocks to jointly pre-train the stream encoder and the stream decoder, wherein the stream encoder extracts deep features of the stream byte block embedding sequence, and the stream decoder reconstructs the stream byte block embedding sequence according to the extracted deep features; The task fine-tuning module is used to fine-tune the prediction model composed of the pre-trained stream encoder and classifier using the labeled traffic for the specific classification task, so as to adapt to the classification of the encrypted traffic of the specific classification task.