Encrypted network traffic classification method based on pre-trained large language model

Through pre-training large language model, encrypted network traffic data is performed on two-byte cleaning and preprocessing, combined with bidirectional attention mechanism retraining and confrontation training, the adaptability and accuracy of the encrypted network traffic classification method in a rapidly changing environment is solved, and efficient encrypted network traffic classification is achieved.

CN120372350APending Publication Date: 2025-07-25CHONGQING UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510432176.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

After the existing encryption network traffic classification methods are widely adopted, it is difficult to adapt to rapidly changing encryption strategies and new environments. Traditional methods rely on experts to design feature generalization capabilities, and deep learning methods are highly dependent and poorly adaptable to labeled data.

Method used

The large language model is adopted to clean and pre-process the encrypted network traffic data in two bytes, and the causal language model is trained using GPT-2, combined with the two-way attention mechanism to retrain, and fine-tune it using adversarial training to generate an encrypted network traffic classification model.

Benefits of technology

It realizes learning deep semantic information in label-free data, improves the accuracy of encrypted network traffic classification and the generalization ability of the model, and adapts to the rapid classification of new environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372350A_ABST
    Figure CN120372350A_ABST
Patent Text Reader

Abstract

The invention relates to an encrypted network traffic classification method based on a pre-trained large language model, and belongs to the technical field of encrypted network traffic classification. The method mainly comprises three stages: a pre-training stage: converting original encrypted network traffic data into a double-byte hexadecimal format through preprocessing, generating and optimizing a basic vocabulary by using a byte pair coding algorithm, then constructing a large language model, and obtaining a pre-training model through distributed training; in the retraining stage, the data is subjected to head byte shuffling processing, and the pre-training model is quickly retrained to improve the generalization ability. In the fine tuning stage, to-be-classified data is preprocessed to generate hexadecimal double-byte data with labels, and classification task training is performed by using the retrained model to obtain an encrypted network traffic classification fine tuning model and classification accuracy. Through the combination of pre-training and retraining and the fine tuning of the pre-training model by means of data classification, the efficient processing and accurate classification of the complex encrypted network traffic are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of encrypted network traffic classification, and relates to an encrypted network traffic classification method based on a pre-trained large language model. Background Art

[0002] Encrypted network traffic classification aims to identify the categories of encrypted traffic from different applications or network services, and is an important technology in the fields of network management and network security. By classifying encrypted network traffic, potential threats can be effectively identified, service quality and user experience can be improved, and efficient network management can be achieved. However, with the widespread adoption of traffic encryption technologies such as the Transport Layer Security (TLS) protocol and anonymous network technologies such as Virtual Private Network (VPN) and The Onion Router (Tor), traditional traffic classification methods face great challenges. Traditional methods mainly rely on Deep Packet Inspection (DPI), which analyzes by capturing patterns and keywords from the payload. However, due to the application of encryption technologies, the packet payload becomes unreadable, and traditional methods are not applicable to encrypted traffic. In addition, classification methods for specific encrypted traffic are difficult to adapt to rapidly changing encryption policies or new environments.

[0003] Traditional machine learning methods based on statistical features: By extracting the statistical features of encrypted traffic and combining classical machine learning algorithms, these methods alleviate the challenge of the lack of plaintext information to a certain extent. However, due to the high dependence on features designed by experts, the generalization ability of these methods is limited, and it is difficult to comprehensively characterize traffic behavior.

[0004] Traffic classification methods based on deep learning: Deep learning methods can automatically learn complex patterns from raw byte-level data, significantly improving the performance of traffic classification. However, these methods are highly dependent on the quantity and distribution of labeled data, which is prone to causing model bias, and they are less adaptable when facing new encryption technologies or data distributions.

[0005] Recently, pre-trained models have shown significant advantages in fields such as Natural Language Processing (NLP) and Computer Vision (CV). The pre-training technology learns general data representations through a large amount of unlabeled data and fine-tunes through limited labeled data, achieving good results in downstream tasks.

[0006] In summary, there is a need for an encrypted network traffic classification technology based on a pre-trained model that can learn deep semantic information from unlabeled traffic, fine-tune the model through the input of labeled samples, quickly adapt to the new environment, and achieve efficient and accurate encrypted network traffic classification. Based on this, the present invention proposes an encrypted network traffic classification method based on a pre-trained large language model, which can learn the deep semantic information in the data from a large amount of unlabeled data and achieve efficient classification of encrypted network traffic data through fine-tuning. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to implement an encrypted network traffic classification method based on a pre-trained large language model, and propose a deep learning method that uses pre-training, re-training, and fine-tuning. By cleaning and preprocessing the original encrypted network traffic into double-byte format, the pre-training of the large language model is used to learn the deep semantic information in the encrypted network traffic, and fine-tuning is performed through a labeled classification data set to obtain an efficient encrypted network traffic classification model. To achieve the above purpose, the technical solution of the present invention is as follows:

[0008] An encrypted network traffic classification method based on a pre-trained large language model includes the following steps:

[0009] S1. Pre-training stage: Clean and preprocess the original encrypted network traffic data into double-byte hexadecimal format, construct an unlabeled training corpus, and use the autoregressive model GPT-2 for causal language model training to learn the deep semantic representation of encrypted traffic;

[0010] S2. Re-training stage: Shuffle the first byte of the data packet header after preprocessing, and perform rapid re-training on the pre-trained model through a bidirectional attention mechanism to improve the generalization ability of the model;

[0011] S3. Fine-tuning stage: Convert the labeled encrypted network traffic data into a token sequence containing special tokens, and perform fine-tuning training on the model by adopting an adversarial training method to generate a final model for encrypted traffic classification.

[0012] Further, the S1 specifically includes the following steps:

[0013] S1.1. Take the original encrypted network traffic data as input, clean and split the data, convert the original data into hexadecimal vocabulary, and generate a double-byte corpus containing encrypted network traffic information;

[0014] S1.2. Perform word segmentation on the double-byte corpus, initialize the word segmentation model based on the byte pair encoding algorithm, and obtain the basic vocabulary;

[0015] S1.3. Merge the high-frequency words in the basic vocabulary, optimize the word encoding in the basic vocabulary through the word-piece algorithm model, and perform serialization operations to optimize the vocabulary;

[0016] S1.4. Use the vocabulary described in S1.3 to build a large language model based on Transformer, train the deep learning model in a distributed training manner, perform forward propagation and backward propagation steps, calculate the loss value, and update the gradient based on the cumulative number of steps;

[0017] S1.5. Repeat S1.1 to S1.4 to obtain a pre-trained model for encrypted network traffic.

[0018] Furthermore, the specific steps of S2 are as follows:

[0019] S2.1. Use the corpus described in S1.1, perform word segmentation on it through the word-piece algorithm model, and obtain a final vocabulary containing special symbols;

[0020] S2.2. Shuffle the first bytes of the vocabulary described in S2.1, initialize the hyperparameters of the pre-trained model described in S1.5, and perform retraining on it to obtain a retrained model.

[0021] Furthermore, the specific steps of S3 are as follows:

[0022] S3.1. Initialize the original encrypted network traffic to be classified, and obtain encrypted network traffic data in double-byte hexadecimal with labels through data classification preprocessing;

[0023] S3.2. Perform division operations on the encrypted network traffic data described in S3.1 into training set, validation set, and test set, use the division results as input samples, and use the retrained model described in S2.2 to perform encrypted network traffic classification training on the above input samples, and finally obtain an encrypted network traffic classification fine-tuning model and the accuracy of encrypted network traffic classification.

[0024] Furthermore, the specific steps of cleaning and splitting the data described in S1.1 are as follows:

[0025] S1.1.1. After reading each data, parse the header information content of the encrypted network traffic flow-level data and divide it into encrypted network traffic packet-level data;

[0026] S1.1.2. Extract the packet-level encrypted network traffic packet-level data and save it as an encrypted network traffic data packet in hexadecimal encoding;

[0027] S1.1.3. Cut the encrypted network traffic data packets with too long lengths, and at the same time cut the strings into segments to further generate a double-byte corpus;

[0028] In S1.2, word segmentation is performed on the double-byte corpus, which specifically includes the following steps:

[0029] S1.2.1. Process the double-byte corpus in S1.1.3 through the byte pair encoding algorithm, and calculate the frequency of each byte in the corpus dataset appearing in the entire corpus;

[0030] S1.2.2. For each pair of bytes, calculate the score after merging adjacent subwords;

[0031] S1.2.3. According to the score in S1.2.2, select the byte pair with the highest frequency of occurrence for merging to generate a vocabulary;

[0032] In S1.3, the word piece algorithm model optimizes the vocabulary encoding in the basic vocabulary, which specifically includes the following steps:

[0033] S1.3.1. After initializing the basic vocabulary in S1.2, all corpus vocabulary in the vocabulary corresponds to the word piece sequence represented by the hexadecimal of the double-byte;

[0034] S1.3.2. Initialize the word piece algorithm according to the word piece sequence, set the maximum token length to control the token granularity and optimize the calculation efficiency, and finally generate a vocabulary containing special markers with more complex semantic information;

[0035] In S1.4, building a large language model based on Transformer specifically includes the following steps:

[0036] S1.4.1. Initialize the autoregressive model GPT-2 base model and train it through the causal language model to deeply understand the semantics of encrypted network traffic data;

[0037] S1.4.2. Perform positional embedding and word embedding operations on the vocabulary in S1.3.2 to construct the representation of word vectors;

[0038] S1.4.3. Train the word vectors in the way of the causal language model to obtain the probability of calculating x for the previous k - 1 tokens, expressed as the following formula: k as follows:

[0039] P L (x k |x1,…,x k-1 ) = softmax(M v h k-1 )

[0040] where P LRepresents the probability of the k-th token; k represents the total number of tokens; P L (x k |x1,…,x k-1 ) represents calculating the probability of the k-th token based on the previous k - 1 tokens; M v Represents a trainable linear transformation matrix; h k-1 Represents the representation obtained by Transformer encoding; softmax(M v h k-1 ) represents the output probability distribution;

[0041] And the objective function, expressed as the following formula:

[0042]

[0043] Among them, θ represents all the trainable parameters of the model and is optimized by the gradient descent algorithm.

[0044] Furthermore, the numeral token algorithm model in S2.1 tokenizes the corpus in S1.1, specifically including the following steps:

[0045] S2.1.1. Read the data packet length in the corpus in S1.1;

[0046] S2.1.2. Add a special character "pck" at the end of each data packet as a delimiter, and add a special character "cls" at the beginning of the data packet to identify the category of the entire data stream;

[0047] In S2.2, the vocabulary in S2.1 is shuffled by the first byte, specifically including the following steps:

[0048] S2.2.1. According to the byte size of the vocabulary, find the byte units to be shuffled in the first byte;

[0049] S2.2.2. Perform a random shuffling operation on the byte units to obtain an enhanced training data set;

[0050] In S2.2, the hyperparameters of the pre-trained model in S1.5 are initialized, specifically including the following steps:

[0051] S2.2.3. Initialize the hyperparameters of the pre-trained model in S1.5, load all parameters onto the CPU, and use the training data set in S2.2.2 to retrain the GPT-2 model in a bidirectional attention manner, and use the multi-head attention mechanism to mine the interaction relationships between features;

[0052] S2.2.4. If the hyperparameters of the pre-trained model are not obtained in S2.2.3, initialize the parameters using a normal distribution, load all the parameters onto the CPU, and use the training dataset in step S2.2.2 to retrain the GPT-2 model in a bidirectional attention manner, and use the multi-head attention mechanism to mine the interaction relationships between features.

[0053] Further, the data classification preprocessing in S3.1 specifically includes the following steps:

[0054] S3.1.1. Initialize the encrypted network traffic data packet, read and extract the payload in the traffic, replace the sensitive fields in the network data packet, convert the data packet into a hexadecimal string, and store the features in a JSON file;

[0055] S3.1.2. Split the hexadecimal string to obtain two-byte data to ensure an efficient representation of the data packet features;

[0056] S3.1.3. Perform label partitioning for different preset data representations: distribute the two-byte data and divide it into a training dataset, a validation dataset, and a test dataset;

[0057] S3.1.4. Save the dataset in TSV format;

[0058] The encrypted network traffic classification training in S3.2 specifically includes the following steps:

[0059] S3.2.1. Load the dataset in S3.1.4, read all the data packets, split the data packets with a fixed length, and if the data packet length is too long, keep the front data and cut the tail data;

[0060] S3.2.2. Initialize the training model, and the embedding layer converts the input into an embedding representation;

[0061] S3.2.3. Extract features through the embedding representation, apply an activation function and a linear layer to output the classification result, and if the target label exists, perform forward propagation to calculate the basic loss;

[0062] S3.2.4. Adopt distributed adversarial training, calculate the adversarial loss and accumulate the basic loss, and return the classification result;

[0063] S3.2.5. Use the classifier to generate the final classification accuracy.

[0064] Further, the division of the encrypted network traffic flow-level data into packet-level data in S1.1.1 specifically includes the following steps:

[0065] Read all data protocol contents according to the header information of the parsed network data traffic;

[0066] Using the protocol content as the unique identifier, use the Splitcap tool to divide the data packets belonging to the same flow, and finally obtain the encrypted network traffic packet-level data at the flow level; the protocol content includes the protocol and the port;

[0067] The generation of the vocabulary containing special marks described in S1.3.2 specifically includes the following steps:

[0068] According to the vocabulary described in S1.3.1, add special characters: "pck" between each data packet, use pad to fill the shorter sequence with a fixed length value, and randomly replace the tokens with mask for marking operations;

[0069] Repeat the above step to obtain the final vocabulary with more complex semantic information.

[0070] An encrypted network traffic classification system based on a pre-trained large language model, including:

[0071] Encrypted network traffic data pre-training module: Divide the original encrypted network traffic data into multiple data packets at the flow level, perform segmentation and padding processing on the data packet lengths to be used as the input of the model training data, and use the autoregressive model GPT-2 for pre-training of the causal language model;

[0072] Encrypted network traffic model re-training module: Based on the pre-trained model parameters, perform word-piece algorithm tokenization and embedding on the pre-processed data samples, and use the bidirectional attention method for rapid re-training of the GPT-2 model to perform efficient characterization of encrypted network traffic data for use in the fine-tuning stage of the encrypted network traffic model;

[0073] Encrypted network traffic model fine-tuning module: By converting the original encrypted network traffic into data samples with labels, perform encrypted network traffic fine-tuning classification training according to the model after re-training of the encrypted network traffic model, so as to obtain the finally fine-tuned encrypted network traffic model, and use the fine-tuned encrypted network traffic model to predict the categories of encrypted network traffic.

[0074] The beneficial effects of the present invention are as follows:

[0075] First, the present invention preprocesses data in the pre-training stage to enhance the model's ability to deeply learn the deep semantic information contained in the data. Secondly, after pre-training, the model is re-trained by byte shuffling to improve the generalization ability of the encrypted network traffic classification model. Finally, during data preprocessing, the original unlabeled encrypted network traffic can be converted into labeled double-byte data as sample inputs, and then the model is fine-tuned to obtain the final fine-tuned model for encrypted network traffic classification. Therefore, this method can learn the deep semantic information in the encrypted network traffic data, achieve efficient and accurate classification of encrypted network traffic, effectively solve the problem of low classification accuracy, and improve the accuracy of encrypted network traffic classification.

[0076] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0078] Figure 1 is a flowchart of an embodiment of the present invention;

[0079] Figure 2 is a schematic diagram of the system operation of an embodiment of the present invention;

[0080] Figure 3 is a schematic diagram of the input layer when GPT-2 of an embodiment of the present invention performs the encrypted network traffic classification task;

[0081] Figure 4 is a model structure diagram when GPT-2 of an embodiment of the present invention performs the encrypted network traffic classification task;

[0082] Figure 5 is a training schematic diagram of GPT-2 as an encrypted network traffic classification task model in an embodiment of the present invention;

[0083] Figure 6 is a schematic diagram of the attention mechanism structure during the model training process of an embodiment of the present invention;

[0084] Figure 7 is a schematic diagram of the multi-head self-attention mechanism structure during the model training process of an embodiment of the present invention;

[0085] Figure 8 is a schematic diagram of byte shuffling of an embodiment of the present invention;

[0086] Figure 9 Schematic diagram of the fine-tuning process according to an embodiment of the present invention. Detailed implementation manners

[0087] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0088] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0089] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0090] The working process of the present invention is mainly divided into a pre-training stage of the encrypted network traffic model, a re-training stage of the encrypted network traffic model, and a classification fine-tuning stage. In the pre-training stage of the encrypted network traffic model, after cleaning and preprocessing the encrypted network traffic, a pre-training is carried out using the base model of the autoregressive model (Generative Pre-trained Transformer 2, GPT-2), and the model is trained in the way of a causal language model to obtain a pre-trained model. On this basis, after shuffling the head bytes of the encrypted network traffic data, a rapid re-training of the model is carried out in a bidirectional attention manner, and finally, fine-tuning is carried out with the encrypted network traffic data with labels to obtain a final encrypted network traffic classification model for encrypted network traffic classification.

[0091] Please refer to Figure 1 , which is a flowchart of an embodiment of the present invention; the key technology of the present invention lies in the joint processing of the pre-training, re-training, and fine-tuning stages of the encrypted network traffic model. It is mainly divided into three stages. First is the pre-training stage. Since the encrypted network traffic model data includes flow-level and packet-level types, relevant preprocessing is performed on the encrypted network traffic data, including operations such as flow-level data partitioning, packet segmentation, and packet padding, so as to perform the initial pre-training of the model. Secondly, in order to improve the efficient understanding of the encrypted network traffic model for data classification, after shuffling the header bytes of the packets, rapid re-training is carried out. Finally, after labeling the original unlabeled encrypted network traffic data, it is used as the sample input for the fine-tuning model, and the encrypted network traffic classification accuracy and the model after fine-tuning of the encrypted network traffic label data are obtained.

[0092] Therefore, the encrypted network traffic model of the present invention can understand the deep semantic information in the encrypted network traffic data, including the logic between the data before and after and the category of the data. At the same time, even in the case of unlabeled encrypted network traffic data, the data preprocessing in the fine-tuning stage can convert the original encrypted network traffic data into labeled data for fine-tuning training of the encrypted network traffic classification model.

[0093] Please refer to Figure 2 , which is a schematic diagram of the system operation of an embodiment of the present invention; the specific operation process is as follows:

[0094] Create a series of encrypted network traffic preprocessing tasks. If the encrypted network traffic data is flow-level data, data partitioning is performed according to the content such as parsing the header information of the network data traffic, and each partitioned traffic comes from the same traffic category containing a complete traffic session.

[0095] Among them, the flow-level data in the original encrypted network traffic dataset is expressed as the following formula:

[0096]

[0097] where n represents the total number of flow-level data in the encrypted network traffic data; l 1 represents the first flow-level data in the encrypted network traffic data; l 2 represents the second flow-level data in the encrypted network traffic data; l n represents the nth flow-level data in the encrypted network traffic data.

[0098] The packets in the flow-level data l i are expressed as the following formula:

[0099]

[0100] Among them, i represents the i-th flow-level data in the encrypted network traffic data; m represents the total number of data packets contained in this flow-level data l i in the middle; represents the first data packet in the i-th flow-level data; represents the second data packet in the i-th flow-level data; represents the m-th data packet in the i-th flow-level data.

[0101] After dividing the data, the total packet-level encrypted network traffic data sets p1, p2,..., p n are obtained. Read the data packets, read the header bytes including IP and source / destination addresses, etc. and the payload part, and obtain the data length len(p i ). If the data packet is too long, len(p i ) > 256, then cut the data packet. If the data packet is too short, then perform padding processing on the data packet and convert it into a hexadecimal string s i . Finally, each data packet is processed into a corpus of two-byte hexadecimals.

[0102] Two-byte hexadecimal operation: First, the string s i is split into a set of single characters {c1, c2,..., c L}, and combinations g k of every two adjacent characters are generated, which is expressed as the following formula:

[0103] g k = c k + c k+1 , k ∈ {1, 2,..., L - 1}

[0104] Among them, k represents the current position in the string; c k represents the k-th single character in the original string s i ; L represents the total number of characters in the string s i .

[0105] The corpus is processed using the byte pair encoding algorithm, and the occurrence frequency f(g i ) of each byte g i in the corpus is calculated. From all the byte pairs, the byte pair (g i , g j ) with the highest occurrence frequency is selected for merging, which is expressed as the following formula:

[0106]

[0107] Among them, j represents the current byte pair (g i , g j) The second byte in it represents the j-th byte in the entire corpus; f(g i , g j ) represents the byte pair (g i , g j )'s occurrence frequency in the entire corpus.

[0108] After the merging operation, recalculate the new byte pair frequencies, repeat the merging steps to generate a high-quality vocabulary. Optimize the vocabulary encoding and serialize it using the wordpiece algorithm. Add a packet separator (PacketSeparator, pck) between each data packet, use a padding symbol (Padding, pad) to fill shorter sequences to a fixed length, and randomly replace tokens with masked tokens (Masked Token, mask) for tokenization to generate a vocabulary with special tokens containing more complex semantic information.

[0109] Please refer to Figure 3 , which is a schematic diagram of the input layer when the GPT-2 of an embodiment of the present invention performs the encrypted network traffic classification task; for the original data of encrypted network traffic, consider an encrypted network traffic packet length sequence as follows:

[0110] “04050025055003d2583c…”

[0111] The word sequence with special tokens represented after processing is:

[0112] [‘[cls]’,‘0025’,‘[mask]’,‘[03d2]’,‘[583c]’,…,‘[pck]’]

[0113] If the data packet sequence is short, the special token '[pad]' will be added before '[pck]'.

[0114] Please refer to Figure 4 , which is a model structure diagram when the GPT-2 of an embodiment of the present invention performs the encrypted network traffic classification task; that is, during the pre-training process of the autoregressive model GPT-2 base large language model using the pre-training dataset, special tokens such as pck are used for traffic data, and training is carried out through a causal language model to deeply understand the semantics of encrypted network traffic data, that is, calculate the probability of the previous k - 1 tokens x k as follows:

[0115] P L (x k |x1,…,x k-1 ) = softmax(M v h k-1 )

[0116] Among them, P L represents the probability of the k-th token; k represents the total number of tokens; P L (x k |x1,…,x k-1 ) represents calculating the probability of the k-th token based on the previous k - 1 tokens; M v represents a trainable linear transformation matrix; h k-1 represents the representation obtained by Transformer encoding; softmax(M v h k-1 ) represents the output probability distribution.

[0117] The objective function is expressed as the following formula:

[0118]

[0119] Among them, θ represents all trainable parameters of the model and is optimized by the gradient descent algorithm.

[0120] Using special marker traffic data such as pck to pre-train the autoregressive model GPT-2 base large language model, the in-depth understanding of the semantic representation of encrypted network traffic data is specifically as follows:

[0121] The GPT-2 base model mainly includes a multi-head self-attention module (Self-Attention) and a fully connected feed-forward network (FeedForward), please refer to Figure 5 , which is a training schematic diagram of GPT-2 as a model for encrypted network traffic classification tasks in an embodiment of the present invention; in the multi-head self-attention mechanism, for each input vector X, query vector Query (Q), key vector Key (K), and value vector Value (V) are generated through linear transformation. The similarity between the query vector Q and the key vector K is calculated through dot product, and it is converted into attention weights using the softmax function. In GPT-2, the multi-head attention mechanism is used to enhance the representation ability of the model. The attention results of multiple heads will be concatenated. The output of the multi-head attention mechanism will be passed to the fully connected feed-forward network for further transformation and non-linear activation. The output of the feed-forward network will be added with a residual connection and a LayerNorm operation to apply the attention weights to the value vector V, thereby generating the encoding of the input vector X. The specific calculation steps are as follows:

[0122] First, for the input sequence, considering the sequence X = [x1, x2,…, x n embedded in the d-dimensional vector space is expressed as the following formula:

[0123] E = [e1, e2,…, e n , e i= Embedding(x i )

[0124] where E represents the word embedding matrix representation of the input sequence; e1 represents the embedding vector of the first token x1 in the input sequence; e2 represents the embedding vector of the second token x2 in the input sequence; e n represents the embedding vector of the nth token x n in the input sequence; x i represents the ith token in the input sequence.

[0125] At the same time, add the positional encoding P:

[0126] H (0) = E + P

[0127] where H (0) represents the initial representation of the input sequence, and P is used to represent the information of each position in the sequence.

[0128] Project it onto three vectors through a linear transformation:

[0129] Q = H (l) W Q , K = H (l) W K , V = H (l) W V

[0130] where Q represents the query vector; l represents the index of the current layer, i.e., the layer of the Transformer; K represents the key vector; V represents the value vector; H (l) represents the input of the lth layer; W Q represents the weight matrix of the query vector; W K represents the weight matrix of the key vector; W V represents the weight matrix of the value vector.

[0131] It contains a learnable parameter matrix to calculate the self-attention.

[0132] Please refer to Figure 6 , which is a schematic diagram of the attention mechanism structure in the model training process of an embodiment of the present invention; calculate the similarity between the query vector Q and the key vector K through the dot product, and use the Softmax function to convert it into the attention weight Attention(Q, K, V) expressed as the following formula:

[0133]

[0134] k where d represents the dimension of the key vector; at the same time, to avoid the gradient problem caused by too large dot product value, set for numerical scaling.

[0135] After the operation of the attention mechanism for each part, the results will be reconnected together. The attention mechanism in the model training process is as Figure 6 shown. The similarity between the query vector Q and the key vector K is calculated through dot product, and it is converted into attention weights using the Softmax function:

[0136]

[0137] At the same time, it avoids the gradient problem caused by too large dot product value for numerical scaling.

[0138] After the operation of the attention mechanism for each part, the results will be reconnected together. Please refer to Figure 7 , which is the schematic diagram of the multi-head self-attention mechanism structure in the model training process of an embodiment of the present invention; the multi-head self-attention mechanism in the model training process combines 12 independent attention heads head i for calculation and splicing:

[0139] MultiHead(Q, K, V) = Concat(head1, head2, …, head 12 )W O

[0140] where, W O represents the output weight matrix; head i = Attention(Q i , K i , V i ), representing each attention head.

[0141] Then, residual connection and normalization are performed:

[0142] Z (l) = LayerNorm(H (l) MultiHead(Q, K, V))

[0143] where, Z (l) represents the result after residual connection and normalization of the output of the l-th layer; H (l) represents the input of the l-th layer.

[0144] In the fully connected feed-forward network, after the output of the attention mechanism, the fully connected feed-forward network (Feed-Forward Network, FFN) in the GPT-2 model is used to perform non-linear transformation on the weighted representation to enhance the representation ability of the model:

[0145] First, pass the input vector X through the first-layer linear transformation W1 and bias b1, and then apply the ReLU activation function to increase the non-linearity ability H, which is expressed as the following formula:

[0146] H = ReLU(XW1 + b1)

[0147] Generate the final feed-forward network output Y through the second-layer linear transformation W2 and bias b2, which is expressed as the following formula:

[0148] Y = HW2 + b2

[0149] The operation of the entire feed-forward network can be expressed as:

[0150] FFN(X) = ReLU(XW1 + b1)W2 + b2

[0151] The first-layer linear transformation maps the input dimension from d to a higher dimension d through W1 ff , improving the representation ability; the second-layer linear transformation maps the dimension from d ff back to the original dimension d, enabling the output to match the processing requirements of subsequent layers.

[0152] The feed-forward network is usually applied after the output of each multi-head attention module to form a complete sub-layer structure, and the training process is stabilized through residual connections and LayerNorm.

[0153] Finally, after all the processing of the input, the output will pass through the final linear layer.

[0154] The purpose of the pre-training stage is to enable the model to learn the deep semantic information contained in the network traffic from a large amount of encrypted network traffic data, while improving the generalization ability of the model, and also providing high-quality model parameters for subsequent re-training and fine-tuning stages, and applying them to the input of a small number of labeled encrypted network traffic samples for model fine-tuning.

[0155] In order to enable the model to learn the encrypted network traffic representations of different categories on the basis of understanding deep semantics; therefore, after pre-training, the model is quickly re-trained by shuffling the header bytes. Different from the traditional pre-training tasks in natural language processing, the header bytes of encrypted network traffic are relatively independent. Utilizing this unique feature can significantly improve the performance of bidirectional learning. On this basis, without affecting the corresponding traffic semantics, the header bytes are shuffled to optimize the pre-trained model and the encrypted network traffic classification task. In addition, this shuffling can also generate more data, thus alleviating the problem of small sample size. It should be noted that although changing the order of the header fields has no impact on semantics, there are still a few tasks (such as generating test traffic for private protocols) that require the correct header field order. Therefore, when constructing the encrypted network traffic classification model, this header byte shuffling adjustment is not performed during the pre-training process, but during the fine-tuning process. The specific implementation steps are as follows:

[0156] To be consistent with the vocabulary in the pre-training process, based on the vocabulary in the pre-training stage, perform header byte shuffling. Please refer to Figure 8 , which is the schematic diagram of byte shuffling in an embodiment of the present invention;

[0157] According to the protocol format of the encrypted network traffic data packet, the version and traffic class fields occupy two bytes, the flow label occupies two bytes, the payload length field occupies two bytes, and the next header length, etc. occupy two bytes. In this example, the version and the flow label are shuffled as a unit, and the payload length field and the next header bytes are shuffled as a unit. Similarly, as in the pre-training stage, consider a sequence of encrypted network traffic packet lengths:

[0158] “6a0bc00000e806f4…”

[0159] After header byte shuffling, it becomes:

[0160] “06f400e8c0006a0b…”

[0161] Similarly, add a special token cls to the vocabulary to handle the encrypted network traffic classification task. After the above processing, after initializing the hyperparameters of the large language model in the pre-training stage, perform rapid re-training of the model in a bidirectional attention manner, and also calculate the attention weights using the multi-head attention mechanism:

[0162]

[0163] And its training method is consistent with the pre-training.

[0164] Please refer to Figure 9, which is a schematic diagram of the fine-tuning process according to an embodiment of the present invention; first, encrypted network traffic data containing tags needs to be obtained and used as the training samples for fine-tuning. The specific steps are as follows:

[0165] Read the encrypted network traffic data, obtain the data content of the payload, and convert it into a hexadecimal string. Then, divide the data packet into single-byte units and generate double bytes of two adjacent units. From the original PCAP file set, generate the corresponding feature data set for each tag, divide the feature data into input X and tag Y, and save them in JSON format. To ensure the stability and reproducibility of the experiment, set the random seed and divide the data set into a training set, a validation set, and a test set in the ratio of 8:1:1 as the sample input for the fine-tuning model.

[0166] Convert the encrypted network traffic data with tags into a token sequence, add a special marker cls in front of the data packets of each category, add pck to the tails of different data packets at the same time, fill in pad for the data packets with shorter lengths, and represent them using a segmented vector seg; map the input seg, position, and token to an embedding vector Emb(x), and after being processed by the encoder and the pooling layer P, obtain the final output layer representation; it uses a two-layer fully connected network, extracts the output vector corresponding to the special marker cls as the input of the last fully connected layer for downstream tasks, and at the same time, the activation function generates a prediction:

[0167] O = Linear2(tanh(Linear1(P)))

[0168] Among them, o represents the result of the final output prediction; Linear1(P) represents the result output by the first-layer fully connected network at the pooling layer P.

[0169] Regarding the loss function, in order to better apply it to encrypted network traffic classification, prevent overfitting at the same time, improve the accuracy of classification, ensure that the training process can be more flexible and efficient, adapt to different types of problems and data distributions, and thus improve the generalization ability and performance of the model, that is, adopt the method of combining hard-label loss and soft-label loss. Therefore, the hard-label loss (cross-entropy loss) L hard and the soft-label loss (mean squared error loss) L soft are defined respectively, that is:

[0170]

[0171]

[0172] Among them, y i represents the hard label of the sample; Softmax(o) iRepresents the predicted probability belonging to the i-th class; N represents the number of encrypted traffic data samples; O i Represents the original predicted output of the model; S i Represents the soft label of the sample.

[0173] The final loss L is a weighted combination of the two:

[0174] L = αL soft +(1 - α)L hard

[0175] where α represents the weight balancing the hard label and soft label losses.

[0176] The optimizer performs an update on the model parameters θ, considering parameter grouping and using the gradient descent algorithm for model parameter update:

[0177]

[0178] To enhance the model's robustness, the anti-attack ability of the model is improved by adding adversarial samples during the training process, that is, using the FGM adversarial training method to add a perturbation δ:

[0179]

[0180] where ∈ represents the magnitude of the perturbation; Represents the gradient of the loss function θ with respect to the model parameters θ.

[0181] Finally, the relevant evaluation of the model is carried out, defining relevant evaluation indicators and parameters, and the F1 score and Ac are used to measure the overall performance, expressed as:

[0182]

[0183] where:

[0184] TP is the true positive value, that is, the number of samples that are actually positive classes and are correctly predicted as positive classes;

[0185] FP is the false positive value, that is, the number of samples that are actually negative classes and are incorrectly predicted as positive classes, also known as "false alarm";

[0186] TN is the true negative value, that is, the number of samples that are actually negative classes and are correctly predicted as negative classes;

[0187] FN is the false negative value, that is, the number of samples that are actually positive classes and are incorrectly predicted as negative classes, also known as "missed alarm".

[0188] In the verification experiment, the datasets used in the present invention are two publicly available datasets: the ISCXVPN 2016 dataset publicly released by the Canadian Institute of Cybersecurity and the USTC-TFC2016 dataset publicly released by the cybersecurity team of the University of Science and Technology of China.

[0189] The ISCXVPN 2016 dataset consists of six communication applications captured in VPN and non-VPN, mainly including encrypted communication traffic through the virtual private network (VPN) tunnel. The USTC-TFC 2016 dataset mainly includes encrypted flows of malware and benign applications, and the effectiveness of this method is verified by examples. The dataset statistical information is shown in Table 1.

[0190] Table 1

[0191] Dataset ISCXVPN2016 USTC-TFC2016 #packets 137163 97115 #flows 6023 9853

[0192] The ISCXVPN 2016 dataset contains multiple traffic types: HTTP, HTTPS, VoIP, chat, email, file transfer, etc., including both ordinary network traffic (Non-VPN) and traffic encrypted through VPN (VPN). Similarly, the USTC-TFC2016 dataset also contains typical application traffic types: including mail (Email), file transfer (FTP), web browsing (Web), peer-to-peer transfer, video streaming, chat tools, etc. Therefore, in order to effectively test the classification accuracy of services and applications, the dataset is further classified by services and applications, and the classification accuracy and adaptability of the method of the present invention for various encrypted network traffic data are fully verified, that is, multiple tasks are divided for experimental testing, as follows:

[0193] According to the ISCXVPN 2016 data, two tasks are divided: the VPN detection task (Task 1), which contains 2 categories, and the application classification task (Task 2), which contains 13 categories. According to the USTC-TFC 2016 data, two tasks are divided: the attack detection task (Task 3), which contains 2 categories, and the software identification task (Task 4), which contains 20 categories. The information statistics of the encrypted traffic classification task are shown in Table 2.

[0194] Table 2

[0195]

[0196] The hyperparameters in the model pre-training stage are consistent with GPT-2, that is, the dimension of the hidden layer and the embedding vector is set to 768, the number of heads in the multi-head attention mechanism is set to 12, etc. In the fine-tuning stage, after multiple trainings, the optimal parameters for model training are obtained. The number of model training rounds is set to 8, the number of data samples in each batch is set to 32, the learning rate is set to 0.00002, and the mean pooling method is adopted.

[0197] At the same time, it is also compared with three existing state-of-the-art methods. One is the ET-BERT model method, which is superior to the basic BERT model method and other non-pre-trained methods in the traffic classification task. The other is the GPT-2 model method based on the original foundation.

[0198] The experimental results of Task 1 and Task 2 under Packet data are shown in Table 3.

[0199] Table 3

[0200]

[0201] The experimental results of Task 3 and Task 4 under packet data are shown in Table 4.

[0202] Table 4

[0203]

[0204] The experimental results of Task 1 and Task 2 under flow data are shown in Table 5.

[0205] Table 5

[0206]

[0207] The experimental results of Task 3 and Task 4 under flow data are shown in Table 6.

[0208] Table 6

[0209]

[0210] According to Tables 3 - 6, in multiple classification tasks, compared with the ET-BERT method, the Ac index of the present invention has an accuracy improvement of 1% - 2%, and even a 39% improvement for the F1 index. Compared with the GPT-2 method, both the Ac index and the F1 index have an accuracy improvement of 1% - 7%. Through the analysis of the experimental results, it can be clearly seen that the overall classification performance of the model is better than the other two methods and can better cope with the problem of encrypted network traffic classification.

[0211] The key technical points of the present invention are as follows:

[0212] 1. A model training strategy based on pre-training, re-training, and fine-tuning is designed. First, the original encrypted network traffic data is converted into a double-byte data format, and the pre-training is carried out using the autoregressive model GPT-2 base model. The model is trained in the way of a causal language model to deeply understand the semantics of the encrypted network traffic data. Secondly, after adding special tokens to the data, the model is re-trained in the way of bidirectional attention. Finally, the data with labels is used as sample input to fine-tune the model, so as to obtain the accuracy of encrypted network traffic classification and an efficient fine-tuning model for encrypted network traffic classification.

[0213] 2. In the data preprocessing of this method, data preprocessing and cleaning are carried out on the data by means of encrypted network traffic data stream-level division, packet segmentation, and padding. At the same time, the encrypted network traffic data is divided into double-byte form for data embedding, so as to improve the learning effect in the model training process and deeply learn the semantic information contained in the data.

[0214] 3. A method for improving classification performance through re-training is designed. In order to improve the classification performance of the model, the preprocessed encrypted network traffic data is shuffled by the header byte, and the hyperparameters of the pre-trained model are initialized. The enhanced training data is used as sample input, and the model is quickly re-trained in the way of bidirectional attention, so as to improve the generalization ability of the model.

[0215] 4. The model is trained by fine-tuning. In the fine-tuning stage of this method, the data preprocessing can convert the original encrypted network traffic data into a dataset with labels and use it as sample input for model fine-tuning, and the adversarial training method is adopted to improve the robustness of the model and the performance of encrypted network traffic classification.

[0216] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for classifying encrypted network traffic based on a pre-trained large language model, characterized in that, It includes the following steps: S1. Pre-training stage: Clean and preprocess the original encrypted network traffic data into a double-byte hexadecimal format, construct an unlabeled training corpus, and use the autoregressive model GPT-2 for causal language model training to learn the deep semantic representation of encrypted traffic; S2. Re-training stage: Shuffle the first bytes of the data headers after preprocessing, and perform fast re-training on the pre-trained model through a bidirectional attention mechanism to improve the model's generalization ability; S3. Fine-tuning stage: Convert the labeled encrypted network traffic data into a token sequence with special markers, and perform fine-tuning training on the model by adopting an adversarial training method to generate a final model for encrypted traffic classification.

2. The encrypted network traffic classification method based on a pre-trained large language model according to claim 1, wherein The specific steps of S1 include the following: S1.

1. Take the original encrypted network traffic data as input, clean and split the data, convert the original data into hexadecimal vocabulary, and generate a double-byte corpus containing encrypted network traffic information; S1.

2. Perform word segmentation on the double-byte corpus, initialize the word segmentation model based on the byte pair encoding algorithm, and obtain the basic vocabulary; S1.

3. Perform a merging operation on the high-frequency vocabulary in the basic vocabulary, optimize the vocabulary encoding in the basic vocabulary through the wordpiece algorithm model, and perform serialization operations to optimize the vocabulary; S1.

4. Use the vocabulary in S1.3 to construct a large language model based on Transformer, train the deep learning model in a distributed training manner, execute the forward propagation and backward propagation steps, calculate the loss value, and update the gradient based on the cumulative number of steps; S1.

5. Repeat S1.1 to S1.4 to obtain a pre-trained model for encrypted network traffic.

3. The encryption network traffic classification method based on a pre-trained large language model according to claim 2, wherein The specific steps of S2 include the following: S2.

1. Use the corpus in S1.1, perform word segmentation on it through the wordpiece algorithm model, and obtain a final vocabulary containing special symbols; S2.

2. Shuffle the first bytes of the vocabulary in S2.1, initialize the hyperparameters of the pre-trained model in S1.5, and perform re-training on it to obtain a re-trained model.

4. The encrypted network traffic classification method based on a pre-trained large language model according to claim 3, wherein, The specific steps of S3 include the following: S3.

1. Initialize the original encrypted network traffic to be classified, and obtain encrypted network traffic data in double-byte hexadecimal format with labels through data classification preprocessing; S3.

2. Perform division operations on the encrypted network traffic data in S3.1 into training set, validation set, and test set, use the division results as input samples, and perform encrypted network traffic classification training on the above input samples using the re-trained model in S2.2 to finally obtain an encrypted network traffic classification fine-tuning model and the accuracy rate of encrypted network traffic classification.

5. The encrypted network traffic classification method based on a pre-trained large language model according to claim 2, wherein: The cleaning and splitting of the data in S1.1 specifically include the following steps: S1.1.

1. After reading each data, parse the header information content of the encrypted network traffic flow-level data and divide it into encrypted network traffic packet-level data; S1.1.

2. Extract the packet-level data of the packet-level encrypted network traffic and save it as an encrypted network traffic data packet in hexadecimal encoding; S1.1.

3. Cut the encrypted network traffic data packets with too long lengths, and at the same time split the strings into segments to further generate a double-byte corpus; In S1.2, perform word segmentation on the double-byte corpus, which specifically includes the following steps: S1.2.

1. Process the double-byte corpus in S1.1.3 through the byte pair encoding algorithm, and calculate the frequency of each byte in the corpus dataset appearing in the entire corpus; S1.2.

2. For each pair of bytes, calculate the score after merging adjacent sub-words; S1.2.

3. According to the scores in S1.2.2, select the byte pair with the highest occurrence frequency for merging to generate a vocabulary; In S1.3, the word piece algorithm model optimizes the vocabulary encoding in the basic vocabulary, which specifically includes the following steps: S1.3.

1. After initializing the basic vocabulary in S1.2, all corpus vocabulary in the vocabulary corresponds to the word piece sequence represented by the hexadecimal of the double byte; S1.3.

2. Initialize the word piece algorithm according to the word piece sequence, set the maximum token length to control the token granularity and optimize the calculation efficiency, and finally generate a vocabulary with special markers containing more complex semantic information; In S1.4, build a large language model based on Transformer, which specifically includes the following steps: S1.4.

1. Initialize the autoregressive model GPT-2 basic model, and train it through a causal language model to deeply understand the semantics of the encrypted network traffic data; S1.4.

2. Perform position embedding and word embedding operations on the vocabulary in S1.3.2 to construct the representation of word vectors; S1.4.

3. Train the word vectors using a causal language model to obtain the probability of calculating \(x\) for the previous \(k - 1\) tokens, which is expressed by the following formula: k as follows: P L (x k |x1,…,x k-1 )=softmax(M v h k-1 ) Among them, P L represents the probability of the k-th token; k represents the total number of tokens; P L (x k |x1,…,x k-1 ) represents the probability of calculating the k-th token based on the previous k - 1 tokens; M v represents a trainable linear transformation matrix; h k-1 represents the representation obtained by Transformer encoding; softmax(M v h k-1 ) represents the output probability distribution; And the objective function, which is expressed as the following formula: Among them, θ represents all trainable parameters of the model, and is optimized through the gradient descent algorithm.

6. According to the encrypted network traffic classification method based on a pre-trained large language model described in claim 3, wherein: In S2.1, the word piece algorithm model performs word segmentation on the corpus in S1.1, which specifically includes the following steps: S2.1.

1. Read the packet lengths in the corpus in S1.1; S2.1.

2. Add a special character "pck" as a separator at the end of each data packet, and at the same time add a special character "cls" at the beginning of the data packet to identify the category of the entire data stream; In S2.2, perform head byte shuffling on the vocabulary in S2.1, which specifically includes the following steps: S2.2.

1. According to the byte size of the vocabulary, find the byte units to be shuffled in the head byte; S2.2.

2. Perform a random shuffling operation on the byte units to obtain an enhanced training dataset; In S2.2, initialize the hyperparameters of the pre-trained model in S1.5, which specifically includes the following steps: S2.2.

3. Initialize the hyperparameters of the pre-trained model described in S1.5, load all parameters onto the CPU, and use the training dataset described in S2.2.2 to retrain the GPT-2 model using a bidirectional attention method, and use the multi-head attention mechanism to mine the interactive relationship between features; S2.2.

4. If S2.2.3 does not obtain the hyperparameters of the pre-trained model, use the normal distribution to initialize the parameters, load all parameters onto the CPU, use the training data set of step S2.2.2, use the bidirectional attention method to retrain the GPT-2 model, and use the multi-head attention mechanism to mine the interactive relationship between features.

7. A method for classifying encrypted network traffic based on a pre-trained large language model according to claim 4, characterized in that: The data classification preprocessing described in S3.1 specifically includes the following steps: S3.1.

1. Initialize encrypted network traffic packets, read and extract the payload in the traffic, replace sensitive fields in the network packets, convert the packets into hexadecimal strings, and store the features in a JSON file; S3.1.2, splitting the hexadecimal string to obtain double-byte data to ensure efficient characterization of data packet features; S3.1.3, performing label division according to different pre-set data representations: labeling the two-byte data and dividing it into a training data set, a verification data set and a test data set; S3.1.4, saving the data set in TSV format; The encrypted network traffic classification training described in S3.2 specifically includes the following steps: S3.2.1, load the data set described in S3.1.4, read all data packets, split the data packets using a fixed length, if the length of the data packet is too long, retain the front data, and cut the tail data; S3.2.2, initialize the training model, and the embedding layer converts the input into an embedded representation; S3.2.3, extract features through the embedding representation, apply activation function and linear layer to output classification results, and if the target label exists, perform forward propagation to calculate the basic loss; S3.2.

4. Distributed adversarial training is performed to calculate the adversarial loss and accumulate the basic loss, and the classification result is returned; S3.2.

5. Use the classifier to generate the final classification accuracy.

8. The encrypted network traffic classification method based on a pre-trained large language model according to claim 5 is characterized in that: The encrypted network traffic flow-level data is divided into packet-level data in S1.1.1, which specifically includes the following steps: Read all data protocol contents by parsing the header information of network data traffic; Using the protocol content as a unique identifier, using the Splitcap tool to divide the data packets belonging to the same flow, and finally obtaining encrypted network traffic packet-level data divided by flow level; the protocol content includes the protocol and the port; The step of generating a vocabulary containing special tags as described in S1.3.2 specifically includes the following steps: According to the vocabulary described in S1.3.1, add special characters between each data packet: "pck", use pad to fill the shorter sequence with fixed length, and randomly replace the word unit with mask for marking operation; Repeat the above step to obtain the final vocabulary with more complex semantic information.

9. An encrypted network traffic classification system based on a pre-trained large language model, characterized in that, Including: Encrypted network traffic data pre-training module: The original encrypted network traffic data is divided into multiple data packets at the flow level, the data packet lengths are segmented and padded as the input of the model training data, and the autoregressive model GPT-2 is used for the pre-training of the causal language model; Encrypted network traffic model retraining module: Based on the pre-trained model parameters, the pre-processed data samples are tokenized and embedded using the WordPiece algorithm, and the GPT-2 model is quickly retrained in a bidirectional attention manner for efficient representation of encrypted network traffic data classification, for use in the fine-tuning stage of the encrypted network traffic model; Encrypted network traffic model fine-tuning module: By converting the original encrypted network traffic into data samples with labels, the encrypted network traffic is fine-tuned and classified based on the model after the encrypted network traffic model is retrained, so as to obtain the final fine-tuned encrypted network traffic model, and the fine-tuned encrypted network traffic model is used to predict the encrypted network traffic categories.

Citation Information

Cited By

  • Industrial internet network intrusion detection method based on pre-trained large language model

    CN121644231A

  • Pre-training encrypted traffic classification method based on self-distillation dynamic reasoning acceleration

    CN121644468A

  • DoH tunnel detection method based on feature fusion and large language model

    CN121864426A