Encryption Network Traffic Analysis Method, Device and Electronic Device Based on Packet Length Sequence

By using a packet-length sequence-based method in encrypted network traffic analysis, the PS-LM model is pre-trained and fine-tuned, which solves the problem of insufficient robustness and generalization in the existing technology, and achieves higher analysis accuracy and robustness, adapting to the optimization needs of different downstream tasks.

CN119830119BActive Publication Date: 2025-06-17NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510312774.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-17
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The prior art is difficult to adapt to the adjustment and focus on optimization of different downstream tasks in encrypted network traffic analysis, resulting in weak robustness and insufficient generalization.

Method used

Using an encrypted network traffic analysis method based on packet-long sequence, the PS-LM model is pre-trained and fine-tuned, combined with the word segmentation embedding module and the PS-LM module, the classification header is used to classify downstream tasks.

Benefits of technology

It achieves higher analysis accuracy and stronger robustness, can stably output correct results under common network flow perturbations, and adapt to the optimization needs of different downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830119B_ABST
    Figure CN119830119B_ABST
Patent Text Reader

Abstract

The present invention discloses a network traffic analysis method, specifically a method, device, and electronic device for encrypted network traffic analysis based on packet length sequences. The method includes: pre-training a PS-LM model to be trained based on a first training dataset to obtain a PS-LM model, where the PS-LM model includes a token embedding module and a PS-LM module. The token embedding module is used to obtain the embedding of the token sequence corresponding to the network flow, and the PS-LM module is used to process the embedding of the token sequence to obtain the output vector of each token in the token sequence; performing supervised fine-tuning on the network traffic analysis model based on a second training dataset to obtain a trained network traffic analysis model, where the network traffic analysis model includes a PS-LM model and a classification head; analyzing the network flow to be processed based on the trained network traffic analysis model to obtain an analysis result. The method provided by the present invention can achieve better network traffic analysis effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to a network traffic analysis method, and more particularly, to an encrypted network traffic analysis method, device and electronic device based on packet length sequences. Background Art

[0002] The application traffic recognition technology was originally based on port numbers. However, dynamic port numbers and port misuse have made the accuracy of this method lower and lower. Subsequently, Deep Packet Inspection (DPI) was used to identify traffic by recognizing the characteristic words in the data packets. With the emergence of encrypted traffic, the applicability of the deep packet inspection method based on plaintext information has been greatly reduced. Currently, the application traffic recognition methods based on machine learning and deep learning are the mainstream.

[0003] The method based on packet length sequences mainly relies on the behavior information of encrypted traffic to achieve analysis. The effectiveness of this type of method is not limited by the amount of plaintext information retained in the encryption protocol. Therefore, in the field of encrypted traffic analysis, it is a more general means to analyze network behavior based on packet length sequences.

[0004] In the field of encrypted traffic analysis, there are usually different downstream tasks. Although the existing methods can achieve acceptable accuracy, it is difficult to make adaptive adjustments and focus optimizations for different downstream tasks, and there are problems of weak robustness and insufficient generalization. Summary of the Invention

[0005] The technical problem to be solved by the present invention is that the above-mentioned methods commonly used in the prior art are difficult to make adaptive adjustments and focus optimizations for different downstream tasks, and there are problems of weak robustness and insufficient generalization. To solve the above problems, the present invention provides an encrypted network traffic analysis method, device and electronic device based on packet length sequences.

[0006] The content of the present invention includes:

[0007] In a first aspect, an embodiment of the present invention provides an encrypted network traffic analysis method based on packet length sequences, including:

[0008] Pre-training a PS-LM model to be trained based on a first training data set to obtain a PS-LM model. The PS-LM model includes a word segmentation embedding module and a PS-LM module connected in sequence. The word segmentation embedding module is used to process the network flow to obtain the embedding of the corresponding word element sequence of the network flow, and the PS-LM module is used to process the embedding of the word element sequence to obtain the output vector of each word element in the word element sequence;

[0009] Perform supervised fine-tuning on the network traffic analysis model based on the second training dataset to obtain a trained network traffic analysis model. The network traffic analysis model includes the PS-LM model and a classification head, and the classification head is used to perform classification processing corresponding to downstream tasks;

[0010] Analyze the network flow to be processed based on the trained network traffic analysis model to obtain an analysis result.

[0011] Optionally, the analyzing the network flow to be processed based on the trained network traffic analysis model to obtain an analysis result includes:

[0012] Obtain the embedding of each token in the token sequence corresponding to the network flow to be processed based on the token embedding module, and a special identifier is added at the beginning of the token sequence;

[0013] Input the embedding of each token in the token sequence into the PS-LM module for processing to obtain the output vector of each token in the token sequence, and the output vector of the special identifier is used to represent the token sequence;

[0014] Input the output vector of the special identifier into the classification head for classification processing corresponding to downstream tasks to obtain an analysis result.

[0015] Optionally, the token embedding module includes a tokenizer and an embedding layer. The obtaining the embedding of each token in the token sequence corresponding to the network flow to be processed based on the token embedding module, and a special identifier is added at the beginning of the token sequence includes:

[0016] Input the network flow to be processed into the tokenizer for tokenization processing, and add the special identifier at the beginning of the sequence to obtain the token sequence based on the packet length value;

[0017] Use the embedding layer to encode the token sequence to obtain the embedding of each token in the token sequence, and the embedding of the token includes the position information of the token.

[0018] Optionally, the second training dataset includes a first-stage dataset and a second-stage dataset. The performing supervised fine-tuning on the network traffic analysis model based on the second training dataset to obtain a trained network traffic analysis model includes:

[0019] Connect the classification head after the PS-LM module of the PS-LM model to obtain the network traffic analysis model;

[0020] Freeze the network parameters of the PS-LM model in the network traffic analysis model, and perform gradient propagation on the classification head based on the first-stage dataset to obtain a first-stage trained model;

[0021] Thaw the network parameters of the PS-LM model in the first-stage trained model, and perform full-parameter fine-tuning on the first-stage trained model based on the second-stage dataset to obtain the trained network traffic analysis model.

[0022] Optionally, analyzing the to-be-processed network flow based on the trained network traffic analysis model to obtain an analysis result, including:

[0023] Calculating the sample robustness based on the probabilities of the perturbation samples falling into different category regions;

[0024] Calculating the model robustness based on the distribution of the recognition results of the trained network traffic analysis model on the test set under noise perturbation;

[0025] When the sample robustness and the model robustness meet the preset conditions, analyzing the to-be-processed network flow based on the trained network traffic analysis model to obtain an analysis result.

[0026] Optionally, the calculating the sample robustness based on the probabilities of the perturbation samples falling into different category regions includes:

[0027] Determining the probability of the perturbation sample falling into the first category region as the first probability, and determining the probability of the perturbation sample falling into the second category region as the second probability, where the first category region is the category region with the largest confidence probability, and the second category region is the category region with the second largest confidence probability;

[0028] Calculating the lower bound of the first probability and the upper bound of the second probability using the binomial proportion confidence interval;

[0029] Determining the sample robustness based on the difference between the lower bound of the first probability and the upper bound of the second probability.

[0030] Optionally, the model robustness is:

[0031] ;

[0032] where is determined based on the PA curve:

[0033] ;

[0034] where is the indicator function, is the label of the th sample, is the label of the th sample determined by the first probability, and , is the number of samples, is the coordinate value in the PA curve, is logical OR, is the lower bound of the first probability corresponding to the is the upper bound of the second probability corresponding to the

[0035] In a second aspect, an embodiment of the present invention provides an encrypted network traffic analysis model based on a packet length sequence, including:

[0036] A pre-training module for pre-training a to-be-trained PS-LM model based on a first training dataset to obtain a PS-LM model. The PS-LM model includes a token embedding module and a PS-LM module connected in sequence. The token embedding module is used to process a network flow to obtain an embedding of a token sequence corresponding to the network flow, and the PS-LM module is used to process the embedding of the token sequence to obtain an output vector of each token in the token sequence;

[0037] A fine-tuning module for performing supervised fine-tuning on the network traffic analysis model based on a second training dataset to obtain a trained network traffic analysis model. The network traffic analysis model includes the PS-LM model and a classification head, and the classification head is used to perform classification processing corresponding to a downstream task;

[0038] An analysis module for analyzing a to-be-processed network flow based on the trained network traffic analysis model to obtain an analysis result.

[0039] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a program stored in the memory and executable on the processor; the processor is used to read the program in the memory to implement the steps in the encrypted network traffic analysis method based on a packet length sequence as described in the first aspect.

[0040] In a fourth aspect, an embodiment of the present invention provides a readable storage medium for storing a program, and the program, when executed by a processor, implements the steps in the encrypted network traffic analysis method based on a packet length sequence as described in the first aspect.

[0041] In the embodiments of the present application, by pre-training the PS-LM model on a large amount of unlabeled data, the PS-model module can automatically learn the relevant knowledge of network traffic data. After adding classification heads corresponding to different downstream tasks to the PS-LM model, network training is carried out through fine-tuning to obtain a trained network traffic analysis model for implementing specific analysis tasks. Compared with the existing methods, the method provided by the embodiments of the present invention can not only achieve higher accuracy, but also has stronger robustness under common network flow perturbations. At the same time, the method provided by the embodiments of the present invention can train different classification heads in the fine-tuning stage for different downstream tasks, so as to focus on optimization for different downstream tasks and obtain better analysis effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Attached Figure 1 is a flowchart of the encrypted network traffic analysis method based on packet length sequence provided by the embodiments of the present invention;

[0043] Attached Figure 2a is an encrypted traffic analysis framework based on the pre-trained PS-LM model provided by the embodiments of the present invention;

[0044] Attached Figure 2b is Figure 2a an enlarged schematic diagram of the token embedding part in

[0045] Attached Figure 2c is Figure 2a an enlarged schematic diagram of the PS-LM module part in

[0046] Attached Figure 2d is Figure 2a a schematic diagram of the method of the model fine-tuning part in

[0047] Attached Figure 3 is a related schematic diagram of the sample robustness provided by the embodiments of the present invention;

[0048] Attached Figure 4 is a related schematic diagram of the model robustness provided by the embodiments of the present invention;

[0049] Attached Figure 5 is a schematic diagram of the encrypted network traffic analysis device based on packet length sequence provided by the embodiments of the present invention;

[0050] Attached Figure 6 is a schematic diagram of the structure of the electronic device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] In the embodiments of the present application, the term "and / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In the embodiments of the present application, the term "multiple" refers to two or more, and other quantifiers are similar. The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same type, and do not limit the number of objects. For example, the first object can be one or multiple.

[0052] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0054] The embodiments of the present application provide a method, device, and electronic device for encrypted network traffic analysis based on packet length sequences, aiming to improve the robustness and generalization of the network traffic analysis model.

[0055] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the method for encrypted network traffic analysis based on packet length sequences provided by the embodiments of the present invention. The method specifically includes the following steps:

[0056] Step 101: Pre-train the PS-LM model to be trained based on the first training dataset to obtain the PS-LM model. The PS-LM model includes a word segmentation embedding module and a PS-LM module connected in sequence. The word segmentation embedding module is used to process the network flow to obtain the embedding of the word element sequence corresponding to the network flow, and the PS-LM module is used to process the embedding of the word element sequence to obtain the output vector of each word element in the word element sequence.

[0057] Step 102: Perform supervised fine-tuning on the network traffic analysis model based on the second training dataset to obtain a trained network traffic analysis model. The network traffic analysis model includes the PS-LM model and a classification head, and the classification head is used to perform classification processing corresponding to downstream tasks.

[0058] Step 103: Analyze the network flow to be processed based on the trained network traffic analysis model to obtain an analysis result.

[0059] Please refer to Figure 2a , and the embodiment of the present invention also provides an encrypted traffic analysis framework based on a pre-trained PS-LM model, which specifically includes a word segmentation and embedding part as shown in Figure 2b , a PS-LM module part as shown in Figure 2c , and a model fine-tuning part as shown in Figure 2d .

[0060] In step 101, first pre-train the PS-LM model to be trained. Among them, the PS-LM model includes a word segmentation and embedding module and a PS-LM module connected in sequence. The PS-LM module is used to process the token sequence to obtain an output vector for each token in the token sequence. A token can also be called a tag. In Natural Language Processing (NLP), a token usually refers to the smallest unit after word segmentation, such as a word, a sub-word, or even a character, and a token sequence is a sequence formed by arranging these tags in a certain order.

[0061] The PS-LM module is used to process the token sequence to obtain an output vector for each token in the token sequence, and its specific structure is not limited here. In the encrypted traffic analysis task, it is more inclined to understand the semantic information of the data and has less demand for the generation task. Among the models based on the Transformer structure, the Encoder-Only type of model represented by BERT adopts a complete multi-head attention mechanism, in which each token can simultaneously pay attention to the information before and after, so it is more suitable for understanding tasks. Therefore, the PS-LM module in this embodiment selects Encoder-Only as the basic model architecture. As a specific embodiment, the PS-LM module is stacked by the Encoder structure of the Transformer.

[0062] As an alternative implementation, the pre-training of the PS-LM model to be trained adopts the denoising autoencoder (DAE) method. Its basic principle is to mask the word to be predicted, and then predict the original value of the masked word based on other unmasked words provided by the context.

[0063] During training, a masking operation needs to be performed on 15% of the tokens in the input sequence. Specifically, the masking operation for the selected 15% of the tokens also includes a third part: directly replacing 80% of them with [MASK], directly replacing 10% with new tokens, and leaving the remaining 10% unchanged. The masked language model (MLM) takes [MASK] as noise, and through the self-encoding training method, bidirectional semantic context information can be obtained. Since the encrypted traffic analysis task does not require tasks such as question answering (QA) and natural language inference (NLI), there is no need to add the NSL task (a variant of the next sentence prediction (NSP)) during the pre-training process.

[0064] In this embodiment, the first training dataset contains a large amount of training data, so that the pre-training process of the PS-LM model is carried out on the packet length sequences of massive network traffic. The purpose is to let the PS-LM model learn the underlying structure and general patterns of network traffic through self-supervised methods, so as to obtain a general and meaningful representation of network traffic. The pre-trained PS-LM model cannot solve problems end-to-end. For specific encrypted traffic analysis tasks, other network layers need to be added after the PS-LM module.

[0065] Optionally, in some embodiments, the second training dataset includes a first-stage dataset and a second-stage dataset, and step 102 includes:

[0066] Connect the classification head after the PS-LM module of the PS-LM model to obtain the network traffic analysis model;

[0067] Freeze the network parameters of the PS-LM model in the network traffic analysis model, and perform gradient propagation on the classification head based on the first-stage dataset to obtain the first-stage training model;

[0068] Unfreeze the network parameters of the PS-LM model in the first-stage training model, and perform full-parameter fine-tuning on the first-stage training model based on the second-stage dataset to obtain the trained network traffic analysis model.

[0069] As Figure 2d shown, in this embodiment, a classification head is connected after the PS-LM module. For different downstream tasks, the corresponding classification heads are different. Optionally, the downstream tasks include encrypted malicious traffic detection, encrypted proxy traffic classification, tunnel website access identification, or application traffic identification.

[0070] Connecting different classification heads for different downstream tasks can make the trained network traffic analysis model more adaptable to the downstream task. For the classified downstream tasks, the classification head is usually a single-layer Multilayer Perceptron (MLP) plus a Softmax layer.

[0071] The PS-LM model has been pre-trained on a large-scale second training dataset, and its parameters usually have converged and reached a relatively stable state. If all model parameters are greatly adjusted directly on a small-scale data, it is easy to cause overfitting. In this embodiment, during the fine-tuning process, it is carried out in the following two stages:

[0072] The first stage: First, freeze the network parameters of the PS-LM model in the network traffic analysis model, and only perform gradient propagation on the classification head based on the first-stage dataset, so that the parameters of the classification head have a preliminary effect, and a first-stage training model is obtained.

[0073] The second stage: Unfreeze the network parameters of the PS-LM model in the first-stage training model, and perform full-parameter fine-tuning based on the second-stage dataset at a low learning rate, so that the model has a better effect, and a trained network traffic analysis model is obtained.

[0074] During the fine-tuning process, the loss function is:

[0075] ;

[0076] Among them, is used to represent the true label, is used to represent the predicted probability, is used to represent the label set, is used to represent the number of samples.

[0077] After obtaining the trained network traffic analysis model, the trained network traffic analysis model can be used for network flow analysis. In specific implementation, different network traffic analysis models are trained for different downstream tasks. The trained network traffic analysis models corresponding to different downstream tasks have the same structure, but different model parameters, so that the model can achieve better results on the downstream task.

[0078] Optionally, in some embodiments, step 103 includes:

[0079] Based on the token embedding module, obtain the embedding of each token in the token sequence corresponding to the network flow to be processed, and a special identifier is added at the start of the token sequence;

[0080] Input the embedding of each token in the token sequence into the PS-LM module for processing to obtain the output vector of each token in the token sequence, and the output vector of the special identifier is used to represent the token sequence;

[0081] Input the output vector of the special identifier into the classification head for classification processing corresponding to the downstream task to obtain an analysis result.

[0082] Please refer to Figure 2a and Figure 2b . Optionally, in some embodiments, the token embedding module specifically includes a tokenizer for tokenization and an embedding layer. The obtaining the embedding of each token in the token sequence corresponding to the network flow to be processed based on the token embedding module, with a special identifier added at the start of the token sequence, includes: inputting the network flow to be processed into the tokenizer for tokenization processing, and adding the special identifier at the start of the sequence to obtain the token sequence based on the packet length value; using the embedding layer to encode the token sequence to obtain the embedding of each token in the token sequence, and the embedding of the token includes the position information of the token.

[0083] It should be understood that the absolute value of the packet length is the load length of the data packet transmission layer, and the sign of the packet length represents the direction of the data packet. In this embodiment, it is stipulated that the uplink packet length is positive and the downlink packet length is negative, and the packet length is directly used as a token, thereby realizing the tokenization of the network flow. In specific implementation, a special identifier ([CLS]) is added before the token sequence of the packet length sequence segmentation result, that is, the token sequence, to represent the start of the token sequence. The output vector of the special identifier token input is used to represent the entire token sequence. It should be understood that in order to identify the end of the token sequence, [SEP] is added at the end position of the token sequence.

[0084] In some embodiments, since the PS-LM module requires a fixed-length input, special tokens, namely [PAD], need to be filled in for the token sequence with insufficient length. These padding values will be masked by [MASK] in the subsequent model to make them not affect the representation result, and specific details are not limited here.

[0085] The purpose of embedding is to map each token into a Euclidean space, so as to represent its semantic information in the form of a fixed-length vector for easy processing by machines. In this embodiment, the representation of the packet length value Token is directly achieved by adding an embedding layer. In network traffic analysis, the packet position information also contains information helpful for analysis. Packets with the same length have different effects on network flows at different positions. Therefore, in some embodiments, position encoding is set for each token in the token sequence, aiming to add position information to the representation. Specifically, the position encoding (PositionEmbedding) is also implemented through an embedding layer with learnable parameters, which is not specifically limited here.

[0086] Specifically, the embedding of each token in the token sequence is input into the PS-LM module of the trained network traffic analysis model for processing to obtain the output vector of each token. The output vector of [CLS] is used for the representation of the entire sequence. Therefore, the output vector of [CLS] is input into the classification head of the trained network traffic analysis model for classification processing corresponding to the downstream task to obtain the analysis result.

[0087] There are many network noises in the real network, such as packet loss, packet retransmission, and packet out-of-order. These network noises will all have a certain impact on the packet length sequence. Therefore, when generating the detection result, the network traffic analysis model also needs to analyze the noise intensity it can resist, that is, robustness. The essence of robustness is whether the model can stably output the correct result when the sample is subject to a certain perturbation.

[0088] Optionally, in some embodiments, step 103 includes:

[0089] Calculating the sample robustness based on the probability that the perturbed sample falls into different category regions;

[0090] Calculating the model robustness based on the distribution of the recognition results of the trained network traffic analysis model on the test set under noise perturbation;

[0091] When the sample robustness and the model robustness meet the preset conditions, analyzing the network flow to be processed based on the trained network traffic analysis model to obtain the analysis result.

[0092] The network traffic analysis model essentially divides the sample space into different regions, with each region corresponding to a category. For a trained network traffic analysis model, the spatial structure of its class regions is usually complex. Some samples are far from the boundary of the class region to which they belong, while some samples are close to the boundary of the class region to which they belong. Mathematically, robustness is the size of the sample neighborhood that can stably output. By calculating the distribution of different classes within the sample neighborhood, that is, the probability that the perturbed samples fall into different class regions, the robustness can be measured.

[0093] Optionally, in some embodiments, calculating the sample robustness based on the probability that the perturbed samples fall into different class regions includes:

[0094] Determine the probability that the perturbed samples fall into the first class region as the first probability, and determine the probability that the perturbed samples fall into the second class region as the second probability. The first class region is the class region with the largest confidence probability, and the second class region is the class region with the second largest confidence probability;

[0095] Use the binomial proportion confidence interval to calculate the lower bound of the first probability and the upper bound of the second probability;

[0096] Determine the sample robustness based on the difference between the lower bound of the first probability and the upper bound of the second probability.

[0097] The probability that the network traffic analysis model determines the sample as class is . Among them, set the class with the largest confidence probability (i.e., the first class region) as , and the class with the second largest confidence probability as (i.e., the second class region). The confidence probability corresponding to the first class region (i.e., the first probability) is , and the confidence probability corresponding to the second class region (i.e., the second probability) is . As long as , it means that the correct result can be output under perturbation within the neighborhood, and The greater the difference between , the more stable the output. Therefore, in this embodiment, the robustness of the sample is measured by .

[0098] Please refer to Figure 3 , in some embodiments, the Monte Carlo method is used to estimate and . To ensure the reliability of the calculation results, the binomial proportion confidence interval is used to calculate the lower bound of and the upper bound of respectively. The sample robustness It can be used to measure the robustness of a sample under specified noise. When the sample robustness meets the preset conditions, it is considered that the sample robustness meets the requirements.

[0099] As a specific embodiment, under specified noise, the calculation method of sample robustness is as follows:

[0100] Set the count vector ; for a certain sample, add specified distribution noise to the sample to form a sample containing noise and input it into the model to obtain the recognition result , that is where . When the number of samples is , perform the above steps times, and then perform statistics on all to obtain and . Among them, and are the indices of the two largest values in the count vector counts, where the count value of is , the count value of is . Calculate the lower bound of the category probability and the upper bound of the category respectively according to the binomial proportion confidence interval algorithm, and finally return , specifically, ; ; ; is the confidence level.

[0101] Since the robustness of different samples is different, the measurement of the model robustness should be based on the distribution of sample robustness. Since it is meaningless to consider the robustness of samples mispredicted by the model, in this embodiment, the measurement of the model robustness also takes into account the accuracy rate.

[0102] The PA curve can show the distribution of the recognition results of the model on the test set under a certain noise perturbation. The PA curve equation is as follows:

[0103] ;

[0104] where is the indicator function, which is 1 if the input variable is true; otherwise it is 0, is the The label of a sample, is the label determined by the first probability for the sample, i.e., class A.

[0105] As Figure 4 shown, the PA curve can simultaneously represent the accuracy and robustness of the model: if the curve is to the right, it indicates strong robustness; if the curve is above, it indicates high accuracy. Specifically, the PA curve is used to characterize the proportion of test samples that can be correctly predicted under the condition of , and its abscissa is , and the ordinate is the prediction accuracy (Accuracy(y)).

[0106] Furthermore, in this embodiment, the model robustness PA-aera is proposed to quantitatively analyze the robustness of the model. PA-aera is used to characterize the area enclosed by the PA curve and the coordinate axes and can be calculated in the following way:

[0107] ;

[0108] where is a point on the PA curve, and .

[0109] In this embodiment, when the model robustness PA-aera meets the preset conditions, it is considered that the model robustness meets the requirements. Through the above method, the quantitative index of robustness under the influence of network noise is considered. By measuring the sample robustness and model robustness, the robustness is quantitatively analyzed from different perspectives, and the performance of the trained network traffic analysis model can be better evaluated. Through the above method, when the performance of the trained network traffic analysis model meets the preset conditions, using the trained network traffic analysis model for network flow analysis can further improve the accuracy of the analysis results.

[0110] In the embodiment of the present application, by pre-training the PS-LM model on a large amount of unlabeled data, the PS-LM model can automatically learn the relevant knowledge of network traffic data. After adding classification heads corresponding to different downstream tasks to the PS-LM model, network training is carried out by fine-tuning to obtain a trained network traffic analysis model for realizing specific analysis tasks. Compared with the existing methods, the method provided by the embodiment of the present invention can not only achieve higher accuracy, but also has stronger robustness under common network flow perturbations. In addition, the method provided by the embodiment of the present invention requires less labeled data to achieve the same accuracy, which proves its better generalization.

[0111] Please refer to Figure 5, an embodiment of the present invention further provides an encrypted network traffic analysis model 500 based on packet length sequences, including:

[0112] A pre-training module 501, configured to pre-train a to-be-trained PS-LM model based on a first training dataset to obtain a PS-LM model. The PS-LM model includes a token embedding module and a PS-LM module connected in sequence. The token embedding module is configured to process a network flow to obtain an embedding of the token sequence corresponding to the network flow, and the PS-LM module is configured to process the embedding of the token sequence to obtain an output vector of each token in the token sequence;

[0113] A fine-tuning module 502, configured to perform supervised fine-tuning on the network traffic analysis model based on a second training dataset to obtain a trained network traffic analysis model. The network traffic analysis model includes the PS-LM model and a classification head, and the classification head is configured to perform classification processing corresponding to a downstream task;

[0114] An analysis module 503, configured to analyze a to-be-processed network flow based on the trained network traffic analysis model to obtain an analysis result.

[0115] Optionally, the analysis module 503 includes:

[0116] A token embedding module, configured to obtain an embedding of each token in the token sequence corresponding to the to-be-processed network flow, and a special identifier is added at the start of the token sequence;

[0117] A PS-LM module, configured to process the embedding of each token in the token sequence to obtain an output vector of each token in the token sequence, and the output vector of the special identifier is used to represent the token sequence;

[0118] A classification head, configured to perform classification processing corresponding to a downstream task on the input of the output vector of the special identifier to obtain an analysis result.

[0119] Optionally, the token embedding module:

[0120] A tokenizer, configured to perform tokenization processing on the input of the to-be-processed network flow and add the special identifier at the start of the sequence to obtain the token sequence based on packet length values;

[0121] An embedding layer, configured to encode the token sequence to obtain an embedding of each token in the token sequence, and the embedding of the token includes the position information of the token.

[0122] Optionally, the second training dataset includes a first-stage dataset and a second-stage dataset, and the fine-tuning module 502 includes:

[0123] A building unit, configured to connect the classification head after the PS-LM module of the PS-LM model to obtain the network traffic analysis model;

[0124] A first training unit, configured to freeze the network parameters of the PS-LM model in the network traffic analysis model, and perform gradient propagation on the classification head based on the first-phase dataset to obtain a first-phase training model;

[0125] A second training unit, configured to unfreeze the network parameters of the PS-LM model in the first-phase training model, and perform full-parameter fine-tuning on the first-phase training model based on the second-phase dataset to obtain the trained network traffic analysis model.

[0126] Optionally, the analysis module 503 is specifically configured to:

[0127] Calculate the sample robustness based on the probabilities of the perturbed samples falling into different category regions;

[0128] Calculate the model robustness based on the distribution of the recognition results of the trained network traffic analysis model on the test set under noise perturbation;

[0129] When the sample robustness and the model robustness meet the preset conditions, analyze the to-be-processed network flow based on the trained network traffic analysis model to obtain an analysis result.

[0130] Optionally, the calculating the sample robustness based on the probabilities of the perturbed samples falling into different category regions includes:

[0131] Determine the probability of the perturbed sample falling into the first category region as the first probability, and determine the probability of the perturbed sample falling into the second category region as the second probability, where the first category region is the category region with the largest confidence probability, and the second category region is the category region with the second largest confidence probability;

[0132] Calculate the lower bound of the first probability and the upper bound of the second probability using the binomial proportion confidence interval;

[0133] Determine the sample robustness based on the difference between the lower bound of the first probability and the upper bound of the second probability.

[0134] Optionally, the model robustness is:

[0135] ;

[0136] wherein, Determined based on the PA curve:

[0137] ;

[0138] Among them, is an indicator function, is the label of the th sample, is the label determined by the first probability for the th sample, and , is the number of samples, is the coordinate value in the PA curve, is logical OR, is the lower bound of the first probability corresponding to the th sample, is the upper bound of the second probability corresponding to the th sample.

[0139] The encryption network traffic analysis model 500 based on the packet length sequence provided by the embodiment of the present application can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0140] It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0141] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc., which can store program codes.

[0142] Such as Figure 6As shown in the figure, an embodiment of the present application provides an electronic device 600, including: a memory 602, a processor 601, and a program stored on the memory 602 and executable on the processor 601; the processor 601 is configured to read the program in the memory 602 to implement the steps in the above-mentioned encrypted network traffic analysis based on the packet length sequence.

[0143] An embodiment of the present application further provides a readable storage medium, on which a program is stored. When the program is executed by a processor, it implements each process of the above-mentioned embodiment of the encrypted network traffic analysis based on the packet length sequence and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memories (such as compact disks (CD), digital versatile disks (DVD), Blu-ray discs (BD), high-definition versatile discs (HVD), etc.), and semiconductor memories (such as read-only memories (ROM), erasable programmable read-only memories (EPROM), electrically erasable programmable read-only memories (EEPROM), non-volatile memories (NAND FLASH), solid-state disks (SSD), etc.).

[0144] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0145] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0146] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A method for analyzing encrypted network traffic based on packet length sequence, characterized in that: include: Pre-training a PS-LM model to be trained based on the first training data set to obtain a PS-LM model, wherein the PS-LM model includes a word segmentation embedding module and a PS-LM module connected in sequence, wherein the word segmentation embedding module is used to process the network flow to obtain the embedding of a word unit sequence corresponding to the network flow, and the PS-LM module is used to process the embedding of the word unit sequence to obtain an output vector of each word unit in the word unit sequence; Performing supervised fine-tuning on the network traffic analysis model based on the second training data set to obtain a trained network traffic analysis model, wherein the network traffic analysis model includes the PS-LM model and a classification head, and the classification head is used to perform classification processing corresponding to the downstream task; Analyze the network flow to be processed based on the trained network traffic analysis model to obtain an analysis result; The step of analyzing the network flow to be processed based on the trained network traffic analysis model to obtain the analysis result includes: The sample robustness is calculated based on the probability that the perturbed sample falls into different category areas; Calculate the model robustness based on the distribution of recognition results of the trained network traffic analysis model on the test set under noise disturbance; When the sample robustness and the model robustness meet preset conditions, the network flow to be processed is analyzed based on the trained network traffic analysis model to obtain an analysis result; The calculating of sample robustness based on the probability of the perturbed sample falling into different category areas includes: The probability that the perturbed sample falls into the first category area is determined as the first probability, and the probability that the perturbed sample falls into the second category area is determined as the second probability, wherein the first category area is the category area with the largest confidence probability, and the second category area is the category area with the second largest confidence probability; Calculating a lower bound of the first probability and an upper bound of the second probability using a binomial proportion confidence interval; Determining the sample robustness based on a difference between a lower bound of the first probability and an upper bound of the second probability; Among them, the model robustness for: ; in, Determined based on the PA curve: ; in, is the indicator function, For the The labels of samples, For the The labels of samples are determined by the first probability, and , is the number of samples, is the coordinate value in the PA curve, For logical AND, For the The lower bound of the first probability corresponding to samples, For the The upper bound of the second probability corresponding to samples.

2. The method according to claim 1, characterized in that: The analyzing the network flow to be processed based on the trained network traffic analysis model to obtain the analysis result includes: Obtaining the embedding of each word in the word-unit sequence corresponding to the network flow to be processed based on the word segmentation embedding module, wherein a special mark is added at the beginning of the word-unit sequence; Input the embedding of each word in the word-gram sequence into the PS-LM module for processing to obtain an output vector of each word in the word-gram sequence, wherein the specially marked output vector is used to represent the word-gram sequence; The output vector of the special mark is input into the classification head to perform classification processing corresponding to the downstream task to obtain the analysis result.

3. The method according to claim 2, characterized in that: The word segmentation embedding module includes a word segmenter and an embedding layer. The embedding of each word in the word-unit sequence corresponding to the network flow to be processed is obtained based on the word segmentation embedding module, and a special mark is added at the beginning of the word-unit sequence, including: Inputting the network flow to be processed into the word segmenter for word segmentation processing, and adding the special identifier at the beginning of the sequence to obtain the word element sequence based on the packet length value; The word-gram sequence is encoded using the embedding layer to obtain an embedding of each word-gram in the word-gram sequence, wherein the embedding of the word-gram includes position information of the word-gram.

4. The method according to claim 1, characterized in that: The second training data set includes a first-stage data set and a second-stage data set, and the network traffic analysis model is fine-tuned in a supervised manner based on the second training data set to obtain a trained network traffic analysis model, including: Connecting the classification head after the PS-LM module of the PS-LM model to obtain the network traffic analysis model; Freeze the network parameters of the PS-LM model in the network traffic analysis model, and perform gradient propagation on the classification head based on the first-stage data set to obtain a first-stage training model; Unfreeze the network parameters of the PS-LM model in the first-stage training model, fine-tune all parameters of the first-stage training model based on the second-stage data set, and obtain the trained network traffic analysis model.

5. An encrypted network traffic analysis model based on packet length sequence, characterized in that: include: A pre-training module, used for pre-training a PS-LM model to be trained based on a first training data set to obtain a PS-LM model, wherein the PS-LM model comprises a word segmentation embedding module and a PS-LM module connected in sequence, wherein the word segmentation embedding module is used for processing a network flow to obtain an embedding of a word unit sequence corresponding to the network flow, and the PS-LM module is used for processing the embedding of the word unit sequence to obtain an output vector of each word unit in the word unit sequence; A fine-tuning module, used to perform supervised fine-tuning on the network traffic analysis model based on the second training data set to obtain a trained network traffic analysis model, wherein the network traffic analysis model includes the PS-LM model and a classification head, and the classification head is used to perform classification processing corresponding to the downstream task; An analysis module is used to analyze the network flow to be processed based on the trained network traffic analysis model to obtain an analysis result; Wherein, the analysis module is specifically used for: The sample robustness is calculated based on the probability that the perturbed sample falls into different category areas; Calculate the model robustness based on the distribution of recognition results of the trained network traffic analysis model on the test set under noise disturbance; When the sample robustness and the model robustness meet preset conditions, the network flow to be processed is analyzed based on the trained network traffic analysis model to obtain an analysis result; The calculating of sample robustness based on the probability of the perturbed sample falling into different category areas includes: The probability that the perturbed sample falls into the first category area is determined as the first probability, and the probability that the perturbed sample falls into the second category area is determined as the second probability, wherein the first category area is the category area with the largest confidence probability, and the second category area is the category area with the second largest confidence probability; Calculating a lower bound of the first probability and an upper bound of the second probability using a binomial proportion confidence interval; Determining the sample robustness based on a difference between a lower bound of the first probability and an upper bound of the second probability; Among them, the model robustness for: ; in, Determined based on the PA curve: ; in, is the indicator function, For the The labels of samples, For the The labels of samples are determined by the first probability, and , is the number of samples, is the coordinate value in the PA curve, For logical AND, For the The lower bound of the first probability corresponding to samples, For the The upper bound of the second probability corresponding to samples.

6. An electronic device comprising: A memory, a processor, and a program stored in the memory and executable on the processor; wherein the processor is used to read the program in the memory to implement the steps in the encrypted network traffic analysis method based on packet length sequence as described in any one of claims 1 to 4.

7. A readable storage medium for storing a program, characterized in that: When the program is executed by a processor, the steps in the encrypted network traffic analysis method based on packet length sequence as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Non-intrusive load monitoring method based on transfer learning and self-attention feature fusion

    CN116742795A

  • Encrypted traffic detection method and device based on comparative learning pre-training

    CN118555155A