Encrypted traffic classification method based on two-stage adaptive architecture

Through the encrypted traffic classification method of the two-stage adaptive architecture, the hybrid OOD detection mechanism and Transformer classification strategy are used to solve the problems of insufficient processing capabilities of out-of-distribution data and difficulty in fine-grained identification of unknown types in the existing technology, and the encrypted traffic classification with high precision and strong generalization capabilities is achieved.

CN119996039AActive Publication Date: 2025-05-13KUNMING UNIV OF SCI & TECH

Patent Information

Application Number
CN202510269495.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-13
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The existing encrypted traffic classification methods have limited processing capabilities when processing Out-of-Distribution (OOD) data, and cannot identify emerging encryption protocols or variant attack traffic in a timely manner, and cannot achieve fine-grained identification of multiple unknown types.

Method used

The two-stage adaptive architecture is adopted. The first stage detects the traffic type through a hybrid OOD detection mechanism, and the second stage uses Transformer-based precise classification strategy and semantic enhancement prompt strategy to perform fine-grained classification of unknown traffic based on the detection results.

Benefits of technology

It realizes accurate identification of known and unknown encrypted traffic types, improves the robustness and applicability of the system, significantly reduces the misjudgment rate, and maintains good detection and classification performance in different encryption scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996039A_ABST
    Figure CN119996039A_ABST
Patent Text Reader

Abstract

The invention relates to an encrypted traffic classification method based on a two-stage adaptive architecture, and belongs to the technical field of computer network security. The method comprises the following steps: firstly, distinguishing data in-distribution (ID) and out-of-distribution (OOD) traffic through a mixed OOD detection mechanism, and processing a detection result by adopting a self-adaptive classification strategy: for the ID traffic, accurately classifying known categories by adopting a transformer-based encoder; for OOD traffic, a classification task is converted into a generation task by combining a large language model (LLM) and a newly proposed semantic enhancement prompt strategy (SPS), so that flexible fine-grained recognition of unknown traffic types is realized. The SPS strategy comprises three levels of a strict mode, a complete mode and an expansion mode, and a flexible generation space is provided while the classification precision is ensured. Through the innovative two-stage design, the problem that the emerging network application cannot be accurately identified by the existing method is effectively solved while the high-precision classification of the ID traffic is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an encrypted traffic classification method based on a two-stage adaptive architecture, which belongs to the field of computer network security technology. Specifically, it relates to a classification method that can simultaneously process encrypted traffic in-data distribution (ID) and out-of-distribution (OOD), which realizes accurate identification of known and unknown encrypted traffic types through an innovative hybrid OOD detection mechanism and a semantic enhancement prompt strategy. Background Art

[0002] In recent years, Internet technology has experienced explosive growth, and waves of innovation in network applications have emerged one after another, which has led to an increasing complexity and diversity of network traffic. On the one hand, existing network applications continue to deeply optimize and update communication protocols and encryption methods; on the other hand, emerging applications continue to emerge like mushrooms after rain, each adopting a unique and significantly differentiated traffic model. In such a rapidly changing network ecological environment, traffic classification technology faces unprecedented severe challenges. As the cornerstone of network security, accurate traffic classification is not only the core support for intrusion detection and abnormal traffic analysis, but also a key link in improving the level of service quality management. It has irreplaceable and important significance for ensuring the security and stability of cyberspace.

[0003] At present, in the field of encrypted traffic classification, the mainstream methods can be roughly summarized into two categories: traditional methods based on feature engineering and cutting-edge methods based on deep learning. Early research work mainly focused on manually designed statistical features. Researchers achieved traffic classification through detailed analysis of basic parameters such as packet size and arrival time interval. With the deepening of research, classification methods based on session behavior have gradually emerged. This method achieves more accurate traffic classification by deeply mining the behavioral characteristics in traffic sessions. Although these traditional methods have good interpretability to a certain extent and can provide researchers with clear classification logic, they generally have cumbersome feature engineering processes, and the limitations of generalization ability are becoming increasingly prominent when facing complex and changing network environments.

[0004] The vigorous development of deep learning technology has brought revolutionary changes and opportunities to encrypted traffic classification. Deep learning models represented by convolutional neural networks (CNN) achieve automatic feature extraction by cleverly converting raw traffic data into images or other structured expressions, greatly improving classification efficiency and accuracy. At the same time, the hybrid model combining recurrent neural networks (RNNs) and their derivative long short-term memory networks (LSTMs) with convolutional neural networks further improves classification performance by accurately modeling the temporal characteristics of traffic data. In recent years, the introduction of the attention mechanism has injected new vitality into deep learning models, which significantly enhances the model's ability to capture long-term dependencies in traffic data and provides a new technical path for achieving higher-precision classification.

[0005] However, most of the existing encrypted traffic classification methods focus on the classification task of known categories of traffic, and are relatively weak in processing unknown traffic (i.e., OOD data) that has never appeared in the training data. Although some studies have tried to alleviate this problem by introducing an "unknown" category, this simple processing method makes it difficult to achieve fine distinction and accurate identification of unknown categories. In actual network application scenarios, network traffic is highly dynamic and diverse, and new applications and protocols continue to emerge like a tide, making the network environment increasingly complex. When encountering traffic patterns outside the scope of training data, existing classification methods often experience significant performance degradation. In addition, lumping all unknown traffic into a single "other" category is far from meeting the strict requirements of network security monitoring for accurate identification.

[0006] In summary, how to achieve fine-grained identification of unknown traffic while ensuring high-precision classification of known traffic has become a key technical problem that needs to be overcome in the current field of encrypted traffic classification. In order to effectively respond to the emerging new encrypted traffic in the network environment and meet the growing high standards of modern network security, it is urgent to innovate and design a new technical solution to adapt to the complex and changing network ecological environment. Summary of the invention

[0007] The technical problem to be solved by the present invention is to provide an encrypted traffic classification method based on a two-stage adaptive architecture to solve the following problems existing in the prior art:

[0008] (1) Existing encrypted traffic classification methods mainly focus on in-distribution (ID) data and have limited processing capabilities for out-of-distribution (OOD) data, resulting in the inability to timely identify newly emerging encryption protocols or variant attack traffic scenarios;

[0009] (2) Most existing methods simply classify all unknown traffic into the “other” category, which is not able to achieve fine-grained identification of multiple unknown types in encrypted traffic;

[0010] (3) With the continuous emergence of new applications and protocols, network traffic patterns are diverse and evolving. There is an urgent need for a classification method that has both high generalization and adaptability and can handle unknown traffic types in a timely manner.

[0011] In response to the above problems, the present invention proposes a two-stage encrypted traffic classification architecture that combines "hybrid detection + adaptive classification". Compared with the traditional model that is only trained on known data, the present invention adds a hybrid OOD detection mechanism in the first stage, and in the second stage, the Transformer-based precise classification strategy and semantic enhancement prompt strategy (SPS) are used to perform fine-grained classification of unknown traffic according to different judgment results. This method makes full use of the advantages of deep learning models in sequence feature extraction, distribution difference measurement, and cross-layer feature smoothness analysis. It can not only achieve high-precision identification of known traffic types, but also dynamically adapt to unknown traffic introduced by new protocols and new applications, greatly improving the robustness and applicability of the overall system.

[0012] The technical solution adopted by the present invention is as follows: an encrypted traffic classification method based on a two-stage adaptive architecture, the specific steps are:

[0013] Step 1: For the input encrypted traffic data sequence, the final feature vector is calculated through the gate control units and state updates of the LSTM network;

[0014] Step 2: Using the final eigenvector obtained, decompose the eigenspace based on the principal component analysis method, calculate the eigencovariance matrix, perform eigenvalue decomposition to obtain eigenvalues ​​and eigenvectors, and construct the principal component space and residual space;

[0015] Step 3: Use the residual space to project the final feature vector and calculate the residual space projection score;

[0016] Step 4: Calculate the smoothness score of the transformation between deep network layers. For the feature representation of adjacent layers in the network, calculate its L 2 The sum of the norm differences is used as the smoothness score;

[0017] Step 5: Set the balance factor, perform a weighted combination of the residual space projection score and the smoothness score to obtain the final hybrid OOD score;

[0018] Step 6: Determine the mixed OOD score based on the preset threshold. When the mixed OOD score is greater than the preset threshold, it is determined as OOD traffic, otherwise it is determined as ID traffic.

[0019] Step 7: Based on the determined encrypted traffic data type, an adaptive classification strategy is used to classify ID and OOD traffic.

[0020] The specific update of each gating unit and state through the LSTM network is as follows:

[0021] For the input encrypted traffic data sequence {x 1 , x 2 , …, x n}, using LSTM network for feature extraction, we get:

[0022] i t =σ(W i [h t-1 , x t ]+b i )

[0023] f t =σ(W f [h t-1 , x t ]+b f )

[0024] o t =σ(W o [h t-1 , x t ]+b o )

[0025]

[0026] h t =o t ⊙tanh(c t )

[0027]

[0028] Among them, i t is the input gate, f t For the forget gate, t is the output gate, is the candidate cell state, c t is the cell state, h t is the hidden state, x t is the encrypted traffic data sequence, W i , W f , W o , Wc is the weight matrix, b i , b f , b o , b c is the bias term, σ is the sigmoid activation function, tanh is the hyperbolic tangent activation function, ⊙ represents element-by-element multiplication, and h n represents the last hidden state of the sequence, n represents the sequence length, h t-1is the hidden state at time t-1, c t-1 is the cell state at time t-1, Representing the final feature vector, by introducing LSTM, longer-distance dependencies can be captured in encrypted traffic sequence modeling, providing a robust and rich representation of timing information for subsequent judgments.

[0029] The Step 2 is specifically as follows:

[0030] After obtaining the sequence features of the LSTM output, in order to further distinguish normal (known distribution) traffic from potential unknown (abnormal) traffic, the present invention combines principal component analysis (PCA) to perform orthogonal decomposition of the feature space, specifically including:

[0031] First calculate the covariance matrix C:

[0032]

[0033] Where N is the number of ID training samples, is the feature vector extracted by LSTM, μ is the mean vector, and T represents the transposed matrix operation;

[0034] Perform eigenvalue decomposition on matrix C:

[0035] C=VΛV T

[0036] Where Λ=diag(λ 1 , …, λ m ) is the eigenvalue, V=[v 1 , …, v m ] is the feature vector, m is the dimension of the feature vector;

[0037] Determine the number of principal components k according to the cumulative variance ratio γ:

[0038]

[0039] Among them, m' represents the number of principal components selected, and γ is the preset variance ratio threshold;

[0040] Construct the principal component space P and residual space R:

[0041] P = span {v 1 , …, v k}

[0042] R = span {v k+1 , …, v m}

[0043] Among them, span{v 1 ,…,v k} represents the vector v1 to v k The linear space of Zhang Cheng can be expressed as α 1 v 1 +α 2 v 2 +…+α k v k A set of vectors of the form α 1 to α k is an arbitrary real coefficient. In principal component analysis (PCA), P represents the principal component space spanned by the first k principal component vectors, capturing the main direction of change in the data; R represents the residual space spanned by the remaining principal component vectors, corresponding to the secondary direction of change and noise in the data.

[0044] The residual space projection score is specifically:

[0045]

[0046] Among them, P R is the projection operator of the residual space R, s 1 is the residual space projection score.

[0047] The Step 4 is specifically as follows:

[0048] In view of the drastic fluctuations in inter-layer feature mapping that may occur when deep networks process encrypted traffic, the present invention further calculates the feature differences between adjacent layers of the transformer and accumulates them to obtain an overall smoothness score. The specific formula is as follows:

[0049] First, obtain the feature representation of each layer:

[0050] F l (X i )=TransformerLayer l (X i )

[0051] Among them, l represents the network layer index, TransformerLayer l represents the l-th layer transformer encoder;

[0052] Calculate the difference of features of adjacent layers:

[0053] d l =||F l (X i )-F l-1 (X i )|| 2

[0054] Among them, d lrepresents the difference between the feature representation of the lth layer and the feature representation of the l-1th layer, F l (X i ) represents the input X of the deep network layer l i The feature representation of F l-1 (X i ) represents the input X of the deep network layer l-1 i The feature representation of ||·|| 2 Indicates L 2 norm;

[0055] Cumulative smoothness scores for all layers:

[0056]

[0057] Among them, L is the total number of network layers, s 2 Score for smoothness;

[0058] If the input traffic is OOD, it often causes large differences in multiple adjacent network layers, making s 2 This inter-layer smoothness analysis combined with principal component decomposition can cross-validate the input in terms of both spatial distribution and network representation dynamics.

[0059] The Step 5 is specifically as follows:

[0060] In order to comprehensively utilize the different focuses provided by residual space projection and smoothness analysis, the present invention combines them into a hybrid OOD score S(X i ), thereby achieving a more robust judgment, the hybrid OOD score S(X i ) is calculated as:

[0061] S(X i )=α·s 1 +(1-α)·s 2

[0062] Specifically:

[0063]

[0064] Among them, S(X i ) is the input X i The hybrid OOD score is α, which is a balancing factor ranging from [0,1], and Σ represents the sum of the differences of all network layers.

[0065] The adaptive classification strategy is specifically:

[0066] If it is determined to be traffic data of ID, a multi-layer transformer encoder is used to process the input feature sequence to obtain the hidden layer representation H i, the feature weights are calculated based on the attention mechanism to obtain the weighted feature representation, the probability distribution P(y|X) of the known categories is calculated through the fully connected layer and the softmax function, and the category with the highest probability is selected as the final classification result; its core idea is to use the attention mechanism to deeply mine the contextual dependencies in the traffic sequence, and combine the multi-head attention structure to model the segment features in parallel. The specific process includes:

[0067] Feature encoding:

[0068]

[0069] Among them, H i is the hidden state between layers, v i It is a classification token representation. represents the transformer encoder with a classification head, and θ represents the model parameters;

[0070] Attention calculation:

[0071]

[0072] Among them, Q, K, V are query, key value and value matrices, d k is the key-value vector dimension, T represents the transposition operation, and softmax represents the normalization function;

[0073] Classification probability calculation:

[0074] P(y|X i )=softmax(W·A+b)

[0075] Among them, W and b are learnable parameters, P(y|X i ) indicates that given input X i The conditional probability distribution of category y at this time. Through this step, the predicted distribution of each known traffic type can be obtained. Since a large amount of known label data is used in the training phase, the present invention can maintain extremely high accuracy and stability when processing conventional encrypted traffic.

[0076] If the traffic data is determined to be OOD, the traffic features are converted into a standardized format for processing by the large language model, and a three-layer prompt template is constructed based on the semantic enhancement prompt strategy. The processed features and prompt template are input into the large language model, and specific application labels are obtained by generating tasks, which can further explore the differences within the unknown traffic and achieve multi-category segmentation. The process includes:

[0077] Feature Normalization:

[0078] Q std =StandardizeFeatures(X i )

[0079] Among them, Q std Represents the standardized feature representation, StandardizeFeatures represents the feature standardization function;

[0080] Select the prompt mode according to the actual scenario:

[0081] T k ∈{T 1 , T 2 , T 3}

[0082] The first level is strict mode, the expression is T 1 ={(X i ,y i )|y i ∈Y OOD}, where y i Indicates the corresponding category label, Y OOD Represents a predefined set of OOD categories;

[0083] The second layer is the complete mode, expressed as T 2 ={(X i ,y i )|y i ∈(Y ID ∪Y OOD )}, where Y ID Represents a known set of ID categories;

[0084] The third layer is the extended mode, the expression is Among them, j represents the index of different data sets, represents the ID category set in the jth data set, Represents the set of OOD categories in the jth dataset.

[0085] Build Tips:

[0086] p={T k ,Q std}

[0087] Where p represents the complete prompt of the build.

[0088] Generate classification labels:

[0089]

[0090] in, is the predicted label sequence, L seq Indicates the length of the sequence of generated labels, y t is the tth token of the generated tag, y 1:t-1represents the subsequence from the 1st to the t-1th token, and Π represents the multiplication operation.

[0091] Through this semantic enhancement prompt mechanism, the present invention can make more detailed category determinations on the detected OOD traffic based on the full mining of implicit information, thereby overcoming the limitation of traditional methods that can only be generalized and marked as "other". In addition, when dealing with emerging new protocols or variant traffic, the reasoning and generation characteristics of the large language model can further improve the interpretability and scalability of unknown categories.

[0092] The present invention achieves the following beneficial effects through the above technical solution:

[0093] (1) Hybrid OOD detection mechanism has excellent performance

[0094] The detection accuracy rate reached 96.81% on the CHNAPP dataset. Due to the integration of residual space projection and inter-layer transformation smoothness analysis, the present invention can maintain stable unknown detection performance under multiple types of encryption protocols and significantly reduce the misjudgment rate.

[0095] (2) Maintain high-precision classification of ID traffic

[0096] A macro-average accuracy of 96.90% was achieved in the ISCXVPN dataset, indicating that the present invention can balance high accuracy and real-time performance when processing common VPN type traffic; the adaptive modeling of network layer features greatly reduces the misidentification of traffic.

[0097] (3) Innovatively implement fine-grained classification of OOD traffic

[0098] A classification accuracy of 97.70% was achieved on the ISCXTor dataset, which not only avoids the rough practice of simply classifying it as "others", but also enables further differentiation of different unknown traffic types, providing a refined basis for network security monitoring and threat warning.

[0099] (4) Strong generalization ability

[0100] Through testing in different encryption scenarios, it is found that whether it is traditional TLS (Transport Layer Security Protocol) traffic, or the emerging QUIC (Quick UDP Internet Connection), SSH (Secure Shell Protocol) or custom encryption protocol, the present invention can maintain good detection and classification performance, and show cross-protocol and cross-distribution robustness.

[0101] (5) Significantly improved computing efficiency

[0102] Compared with existing methods, the average processing time is reduced by 35%, which has significant advantages in online analysis of real-time massive traffic. In addition, the phased processing architecture can be flexibly deployed on the server side or the edge side to meet the needs of various application scenarios.

[0103] It can be seen that the present invention further proposes a two-stage adaptive architecture with high accuracy and strong generalization ability based on the paper. By using key technical means such as LSTM sequence feature extraction, PCA principal component decomposition, inter-layer transformation smoothness measurement, and semantic enhancement prompt strategy (SPS) to guide large language models, it successfully solves the common problems of "unknown traffic cannot be subdivided" and "weak out-of-distribution detection" in existing encrypted traffic classification, and has broad application value and promotion prospects. This technology can be applied to network intrusion detection, real-time traffic supervision, transport layer protocol optimization, and higher-level network security audits. It also has good adaptability and scalability to the network ecology with the continuous emergence of new applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0104] Figure 1 is a flow chart of the steps of the present invention;

[0105] Figure 2 It is a two-stage adaptive architecture framework diagram of the present invention;

[0106] Figure 3 It is a schematic diagram of the hybrid OOD detection mechanism of the present invention;

[0107] Figure 4 This is a performance comparison chart of the present invention on different data sets;

[0108] Figure 5 It is a confusion matrix analysis diagram of the present invention. DETAILED DESCRIPTION

[0109] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific implementation methods.

[0110] Example 1: Figure 1 , Figure 2As shown, the present invention provides an encrypted traffic classification method based on a two-stage adaptive architecture, which includes two stages: OOD detection and adaptive classification. The first stage detects the traffic type through a hybrid mechanism, and the second stage adopts a corresponding classification strategy based on the detection result to achieve accurate classification of ID and OOD traffic. Specifically, the first stage adopts a hybrid OOD detection mechanism to distinguish ID and OOD traffic by combining the smoothness of transformer inter-layer transformation and feature analysis. The second stage adopts different processing strategies based on the detection results: the transformer encoder is used for classification of ID traffic, and the large language model combined with the semantic enhancement prompt strategy is used for processing OOD traffic.

[0111] Figure 3 The hybrid OOD detection mechanism of the present invention is presented in detail, specifically showing the complete process of LSTM-based feature extraction, feature space decomposition, inter-layer smoothness analysis, and score fusion. The calculation formula of the hybrid OOD score S(X i )=α·s 1 +(1-α)·s 2 , where s 1 is the residual space projection score, s 2 Score for inter-layer smoothness.

[0112] Specifically, the mechanism first extracts the traffic feature vector through LSTM. For the input encrypted traffic sequence {x 1 ,x 2 ,…,x n}, LSTM extracts features through gating mechanism and state update. Input gate i t 、Forget gate t and output gate o t The update of follows the following calculation process:

[0113] i t =σ(W i [h t-1 ,x t ]+b i )

[0114] f t =σ(W f [h t-1 ,x t ]+b f )

[0115] o t =σ(W o [h t-1 ,x t ]+b o )

[0116] Among them, Wi , W f , W o is the corresponding weight matrix, h t-1 is the hidden state at time t-1, x t is the current input, b i 、b f 、b o is the bias term. Cell state c t and the hidden state h t The update formula is:

[0117] c t =f t ⊙c t-1 +i t ⊙tanh(Wc[h t-1 ,x t ]+b c )

[0118] h t =o t ⊙tanh(c t )

[0119] Among them, ⊙ represents element-by-element multiplication operation, c t-1 is the cell state at time t-1, Represents the final feature vector, which is composed of the hidden state h at the last moment of the sequence n Given, that is

[0120] After obtaining the feature vector, such as Figure 3 As shown, the present invention uses principal component analysis to decompose the feature space. First, the covariance matrix is ​​calculated:

[0121]

[0122] Where N is the number of ID training samples, is the feature vector extracted by LSTM, μ is the mean vector, and T represents the transposed matrix operation;

[0123] The eigenvalues ​​{λ i} and the eigenvector {v i}. Determine the number of principal components k according to the cumulative variance ratio γ, and construct the principal component space P and residual space R:

[0124] P = span {v 1 ,…,v k}

[0125] R = span {v k+1 ,…,v m}

[0126] Among them, span{v 1 ,…,v k} represents the vector v 1 to v k The linear space of Zhang Cheng can be expressed as α 1 v 1 +α 2 v 2 +…+α k v k A set of vectors of the form α 1 to α k is an arbitrary real coefficient. In principal component analysis (PCA), P represents the principal component space spanned by the first k principal component vectors, capturing the main direction of change in the data; R represents the residual space spanned by the remaining principal component vectors, corresponding to the secondary direction of change and noise in the data.

[0127] At the same time, the present invention analyzes the transformation smoothness of feature representations of adjacent layers in a deep network. Figure 3 As shown on the right, for the L-layer feature representation in the network {F 1 ,…,F L}, calculate L between adjacent layers 2 Norm difference, finally get the smoothness score s of all layers 2 :

[0128]

[0129] Among them, F l (X i ) represents the input X of the deep network layer l i The feature representation of F l-1 (X i ) represents the input X of the deep network layer l-1 i The feature representation of ||·|| 2 Indicates L 2 norm,;

[0130] The final hybrid OOD score is calculated as Figure 3 As shown, the residual space projection score s 1 and smoothness score s 2 Perform weighted fusion:

[0131] S(X i )=α·s 1 +(1-α)·s 2

[0132] Among them, s 1 Score the residual space projection: P Ris the projection operator of the residual space R, α is the balance factor, and its value range is [0,1]. i ) exceeds the preset threshold δ, it is determined as OOD traffic, otherwise it is determined as ID traffic. Experiments show that when α=0.6, δ=0.75, the detection effect is optimal.

[0133] The present invention designs a three-layer semantic enhancement prompt strategy (SPS). This strategy guides the large language model to generate accurate classification labels through prompt templates at different levels. Specifically, it includes:

[0134] Strict mode (T 1 ) contains only specific application categories. Its prompt template is as follows:

[0135] “Classify this encrypted network traffic packet into one of these known application categories:QQMail,QQMusic,Youku,TaoBao.Consider the trafficpacket characteristics including:1.Protocol behavior(TCP flags,window size)2.Packet structure(length,fragmentation)3.Encryptedpayloadpatterns.”

[0136] Full Mode (T 2 ) is extended to all application categories in the dataset. Its prompt template is:

[0137] "Given this encrypted network traffic packet, classify it into one of these possible applications: QQMail, QQMusic, Youku, TaoBao, WeChat, Weibo. Analyze the characteristics including TCP protocol behaviors, packet structure patterns, and encrypted payload features."

[0138] Extended Mode (T 3 ) integrates cross-dataset knowledge. Its prompt template is:

[0139] "Analyze this encrypted network traffic packet and classify it based on its characteristics. Consider the following applications across different platforms: QQMail, QQMusic, Youku, TaoBao, WeChat, Weibo, Gmail, Facebook, Skype, YouTube. "

[0140] The generation space corresponding to each pattern is gradually expanded, achieving a trade-off between classification accuracy and generalization ability.

[0141] It is worth noting that the present invention uses English prompt templates to more accurately describe and match protocol characteristics. This is because hexadecimal data is represented by ASCII: network data packets are usually displayed in hexadecimal form in packet capture tools (such as Wireshark), and most of the displayable ASCII characters are in English. Although the payload of encrypted traffic is encrypted, the protocol handshake process, header structure and other features still retain this representation. The use of English in the prompt template is more consistent with these original data representations. The RFC documents and technical specifications of encryption protocols such as TLS, QUIC, and SSH are all written in English, and the keywords, status codes, and error messages in the protocols are also in English.

[0142] like Figure 4 As shown in the figure, the horizontal axis represents different evaluation indicators (macro-average precision, macro-average F1 value, etc.), and the vertical axis represents the performance score. By comparing with existing methods (PacRep, BART-Large, GPT-4o, etc.), the performance advantages of the present invention are intuitively demonstrated. Specifically, the present invention is comprehensively evaluated on three datasets: CHNAPP, ISCXVPN, and ISCXTor. The evaluation indicators include macro-average precision (M-Prec), macro-average F1 value, and performance score. 1 Value(MacroF 1 ), micro average F 1 Value (MicroF 1 ) and recall rate (Recall). Experimental results show that the proposed method is significantly better than the existing technology. On the CHNAPP dataset, a macro-average precision of 96.81% is achieved, which is 38.68 percentage points higher than 58.13% of PacRep. On the ISCXVPN dataset, a macro-average precision of 96.90% is achieved, which is 14.97 percentage points higher than 81.93% of GPT-4o.

[0143] Figure 5The confusion matrix analysis results of the present invention are shown. The figure shows the classification results on three data sets: CHNAPP, ISCXVPN and ISCXTor. The distribution of diagonal and off-diagonal elements reflects the classification effect of the present invention on different types of traffic. Specifically, it can be clearly seen from the confusion matrix that the method of the present invention exhibits a strong classification ability when processing ID traffic. Taking the CHNAPP data set as an example, the number of accurate recognitions of QQMusic reached 1085 and that of Youku reached 1207. High accuracy was also maintained when processing OOD traffic, and the accurate number of recognitions of WeCHat and Weibo reached 970 and 1014 respectively. This excellent performance is due to the design of a two-stage adaptive architecture.

[0144] Further analysis shows that the advantages of the present invention are particularly prominent when processing VPN encrypted traffic. Figure 5 As shown in the confusion matrix of the ISCXVPN dataset, the present invention can still maintain clear classification boundaries even in the face of complex traffic patterns encrypted by VPN. This high-precision classification effect mainly comes from the effectiveness of the hybrid OOD detection mechanism and the precise guidance of the semantic enhancement prompt strategy.

[0145] In practical applications, the specific deployment process of the present invention is as follows:

[0146] First, data preprocessing and partitioning are performed. Taking the CHNAPP dataset as an example, 485,782 ID category data (QQMail, QQMusic, Youku, TaoBao) are selected from a total of 614,575 traffic data for training, and the remaining 128,793 data include WeChat and Weibo as OOD category data. Among these 128,793 non-training data, a validation set and a test set are constructed, and each set maintains a 7:3 ratio of ID to OOD data. Specifically, the validation set contains 64,391 data (45,074 of which are ID data and 19,317 are OOD data), and the test set contains 64,402 data (45,081 of which are ID data and 19,321 are OOD data). This division ensures that the model can fully access different proportions of known and unknown types of traffic during the validation and testing stages.

[0147] In the model training phase, the OOD detection part uses the Adam optimizer, the learning rate is set to 2e-5, and the training is performed for 20 rounds. The calculation of the hybrid OOD score uses the optimal parameter configuration verified by experiments: balance factor α = 0.6, detection threshold δ = 0.75. These parameter settings also show good adaptability on the ISCXVPN and ISCXTor datasets.

[0148] The ID classification branch uses a 12-layer transformer encoder, with 8 attention heads in each layer. The training uses the AdamW optimizer, with a learning rate of 2e-5 and a total of 30 rounds of training. In the experiment, it was found that this configuration can well balance the model's expressiveness and computational overhead.

[0149] The OOD classification branch uses GPT-4o as the basic language model, with the temperature parameter set to 0.7 and the top-p sampling parameter set to 0.95. During the generation process, the strict mode prompt strategy showed the best classification effect. Experiments have shown that this parameter configuration can maintain moderate output diversity while ensuring the accuracy of generated labels.

[0150] When processing Tor network traffic, such as Figure 5 The results of the ISCXTor dataset show that the present invention exhibits strong robustness. Even in the face of the multi-layer encryption and anonymization processing of the Tor network, the classification accuracy remains at a high level. This proves the practical value of the two-stage adaptive architecture proposed in the present invention in a complex network environment.

[0151] The experiment also found a phenomenon: as the prompt mode changes from strict to extended, although the classification accuracy decreases slightly, the model shows stronger generalization ability. This provides an important reference for how to balance accuracy and generalization in practical applications.

[0152] In terms of computational efficiency, thanks to the reasonable architecture design, the present invention shows significant performance advantages under standard hardware configuration (NVIDIA RTX4090 GPU). Compared with the existing methods, the average processing time is reduced by 35%. This is of great significance for practical scenarios that require real-time processing of large amounts of network traffic.

[0153] It should be noted that the above embodiments are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. A person skilled in the art may make various modifications and improvements without departing from the essence of the technical solution of the present invention, and these modifications and improvements also fall within the scope of protection of the present invention.

Claims

1. A method for classifying encrypted traffic based on a two-stage adaptive architecture, characterized in that: The steps include: Step 1: For the input encrypted traffic data sequence, the final feature vector is calculated through the gate control units and state updates of the LSTM network; Step 2: Using the final eigenvector obtained, decompose the eigenspace based on the principal component analysis method, calculate the eigencovariance matrix, perform eigenvalue decomposition to obtain eigenvalues ​​and eigenvectors, and construct the principal component space and residual space; Step 3: Use the residual space to project the final feature vector and calculate the residual space projection score; Step 4: Calculate the smoothness score of the deep network layer transformation. For the feature representations of adjacent layers in the network, calculate the sum of their L2 norm differences as the smoothness score. Step 5: Set the balance factor, perform a weighted combination of the residual space projection score and the smoothness score to obtain the final hybrid OOD score; Step 6: Determine the mixed OOD score based on the preset threshold. When the mixed OOD score is greater than the preset threshold, it is determined as OOD traffic, otherwise it is determined as ID traffic. Step 7: Based on the determined encrypted traffic data type, an adaptive classification strategy is used to classify ID and OOD traffic.

2. According to claim 1, a method for classifying encrypted traffic based on a two-stage adaptive architecture is characterized in that: The specific update of each gating unit and state through the LSTM network is as follows: I t =σW i h t-1 ,x t +b i f t =σW f h t-1 ,x t +b f o t =σW o h t-1 ,x t +b o h t =o t ⊙tanhc t Among them, i t is the input gate, f t For the forget gate, t is the output gate, is the candidate cell state, c t is the cell state, h t is the hidden state, x t is the encrypted traffic data sequence, W i , W f , W o , W c is the weight matrix, b i , b f , b o , b c is the bias term, σ is the sigmoid activation function, tanh is the hyperbolic tangent activation function, ⊙ represents element-by-element multiplication, and h n represents the last hidden state of the sequence, n represents the sequence length, h t-1 is the hidden state at time t-1, c t-1 is the cell state at time t-1, represents the final feature vector.

3. According to claim 1, a method for classifying encrypted traffic based on a two-stage adaptive architecture is characterized in that: The Step 2 is specifically as follows: Calculate the covariance matrix C: Where N is the number of ID training samples, is the feature vector extracted by LSTM, μ is the mean vector, and T represents the transposed matrix operation; Perform eigenvalue decomposition on matrix C: C=VΛV T Among them, Λ=diagλ1,...,λ m is the eigenvalue, V = v1,…,v m is the feature vector, m is the dimension of the feature vector; Determine the number of principal components k according to the cumulative variance ratio γ: Among them, m' represents the number of principal components selected, and γ is the preset variance ratio threshold; Construct the principal component space P and residual space R: P=span{v1,…,v k } R=span{v k+1 ,…,v m } Where span{v1,…,v k } means from vector v1 to v k In principal component analysis, P represents the principal component space spanned by the first k principal component vectors; R represents the residual space spanned by the remaining principal component vectors.

4. According to claim 1, a method for classifying encrypted traffic based on a two-stage adaptive architecture is characterized in that: The residual space projection score is specifically: Among them, P R is the projection operator of the residual space R, and s1 is the residual space projection score.

5. According to claim 1, a method for classifying encrypted traffic based on a two-stage adaptive architecture is characterized in that: The Step 4 is specifically as follows: Get the feature representation of each layer: F l X i =TransformerLayer l X i Among them, l represents the network layer index, TransformerLayer l represents the l-th layer transformer encoder; Calculate the difference of features of adjacent layers: d l =|F l X i -F l-1 X i |2 Among them, d l represents the difference between the feature representation of the lth layer and the feature representation of the l-1th layer, F l X i Represents the input X of the deep network layer l i The feature representation of F l-1 X i Represents the input X of the deep network layer l-1 i The feature representation of , ||·||2 represents the L2 norm; Cumulative smoothness scores for all layers: Among them, L is the total number of network layers and s2 is the smoothness score.

6. According to claim 1, a method for classifying encrypted traffic based on a two-stage adaptive architecture is characterized in that: The hybrid OOD score is specifically: Among them, SX i For input X i The hybrid OOD score is α, which is a balancing factor ranging from 0 to 1. Σ represents the sum of the differences of all network layers.

7. The encrypted traffic classification method based on a two-stage adaptive architecture according to claim 1 is characterized in that: The adaptive classification strategy is specifically: If it is determined to be traffic data of ID, a multi-layer transformer encoder is used to process the input feature sequence to obtain the hidden layer representation H i , the feature weights are calculated based on the attention mechanism to obtain the weighted feature representation, the probability distribution Py|X of the known categories is calculated through the fully connected layer and the softmax function, and the category with the highest probability is selected as the final classification result; If the traffic data is determined to be OOD, the traffic features are converted into a standardized format for processing by the large language model. A three-layer prompt template is constructed according to the semantic enhancement prompt strategy. The processed features and prompt template are input into the large language model, and specific application labels are obtained by generating tasks.

8. The encrypted traffic classification method based on a two-stage adaptive architecture according to claim 7 is characterized in that: The three-layer prompt template constructed according to the semantic enhancement prompt strategy is specifically: The first level is strict mode, the expression is T1={(X i ,y i )|y i ∈Y OOD }, where y i Indicates the corresponding category label, Y OOD Represents a predefined set of OOD categories; The second layer is the complete mode, and the expression is T2={(X i ,y i )|y i ∈(Y ID ∪Y OOD )}, where Y ID Represents a known set of ID categories; The third layer is the extended mode, the expression is Among them, j represents the index of different data sets, represents the ID category set in the jth data set, Represents the set of OOD categories in the jth dataset.

Citation Information

Patent Citations

  • Computer-implemented method, computer program product and system for data analysis

    CN112639834A

  • Encrypted traffic classification method and device and electronic equipment

    CN113822331A

  • Out-of-distribution network flow data detection method based on calculated likelihood ratio

    CN114844840A

  • System and method for network traffic classification using snippets and on the fly built classifiers

    US11165675B1

Cited By

  • Unknown encrypted traffic identification method and system based on small sample incremental learning, and storage medium

    CN120602237A

  • An unknown encrypted traffic identification method and system based on small sample incremental learning and a storage medium

    CN120602237B

  • Unknown encrypted traffic fine granularity identification method and system based on semi-supervised learning, and storage medium

    CN122112779A

  • A semi-supervised learning-based unknown encrypted traffic fine-grained identification method and system and storage medium

    CN122112779B