QUIC encrypted traffic classification method based on multi-model fusion

By adopting multi-model fusion technology in QUIC traffic classification, combining LGBM and IP-based classifiers, the problem of low adaptability and accuracy of QUIC traffic classification methods in actual network environments is solved, and the classification effect of high accuracy and high recall is achieved.

CN120217301APending Publication Date: 2025-06-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510367362.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing QUIC traffic classification method is difficult to adapt to the actual network environment, has low accuracy, and IP-based classifiers are not suitable for co-hosted service traffic identification classification.

Method used

The QUIC encrypted traffic classification method based on multi-model fusion is adopted, combined with LGBM and IP-based classifiers, and through dynamic weighting mechanism, probability calibration and voting threshold optimization, QUIC traffic classification with high accuracy and high recall is achieved.

Benefits of technology

Improves the accuracy and stability of QUIC traffic classification, and is suitable for different network environments, especially in co-hosted service traffic identification classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217301A_ABST
    Figure CN120217301A_ABST
Patent Text Reader

Abstract

A QUIC encrypted traffic classification method based on multi-model fusion belongs to the field of communication, and comprises the following steps: a model training stage: dividing a data set, training each model in a high-precision model group and a high-recall-rate model group by using a training set, evaluating model performance by using a verification set, if a dynamic weight mechanism is started, calculating the weight of the model according to a model evaluation result, and if the dynamic weight mechanism is started, calculating the weight of the model; carrying out normalization processing on the weight, and searching an optimal decision threshold by using a plurality of candidate thresholds; in the model prediction stage, prediction data are input, each model generates a respective prediction probability, and if a dynamic weight mechanism is started, the prediction probability of each model is subjected to weighted averaging according to the weight obtained in the training stage so as to perform probability calibration; and combining the calibration probability of each model, judging the calibration probability through an optimal decision threshold, and generating a final classification prediction result. According to the method, the risks of overfitting and poor generalization ability of a single model are reduced, and QUIC encrypted traffic classification with high accuracy and stability is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technologies, and particularly relates to a QUIC encrypted traffic classification method based on multi-model fusion. Background Art

[0002] With the rapid development of network technologies, people's awareness of privacy security has gradually increased, and the use of encrypted traffic has become increasingly frequent, which also increases the difficulty of network security supervision. Accurately classifying network traffic can collect the usage habits and needs of network users, so as to provide them with high-quality services and enhance network management. Traditional traffic detection technologies cannot directly analyze the content of encrypted traffic, and classifying and analyzing encrypted traffic will become an important research direction for network security monitoring and management.

[0003] Traffic identification and classification technologies can effectively manage and optimize network traffic, thereby improving the quality and response speed of network services and maintaining the security and stability of the network environment. Specifically, in terms of network management, this technology can effectively identify and classify network traffic, thereby improving the level of network management; in terms of network services, this technology can optimize network traffic, remove redundant and garbage traffic in the network, improve the quality and response speed of network services, and improve the user experience; in terms of network security, this technology can help network administrators monitor and analyze network traffic in real time, understand the traffic characteristics and usage in the network, discover and solve abnormal traffic and security hazards in the network, thereby maintaining the security and stability of the network environment. Therefore, traffic identification and classification technologies are widely used in many fields such as network traffic analysis, service quality management, and intrusion detection systems, and are key technologies for effectively managing Internet traffic.

[0004] The following introduces several different encrypted traffic classification methods.

[0005] 1. Research on Encrypted Traffic Classification Method Based on Multi-Model Fusion;

[0006] The idea of model fusion is similar to Ensemble Learning, both of which construct and combine multiple learners to learn tasks. However, in Ensemble Learning, the learners are homogeneous, while in model fusion, the learners are heterogeneous. The following introduces several widely used fusion methods:

[0007] (1) Voting method

[0008] The voting method obtains the final prediction result by voting on the prediction results of multiple learners, with the minority obeying the majority. The voting method is divided into the ordinary voting method and the weighted voting method. The weighted weights can be set subjectively by humans or according to the model evaluation scores. The voting method requires 3 or more models. Using the voting method among homogeneous models cannot achieve good performance because there is a strong correlation between the results obtained by homogeneous models.

[0009] (2) Averaging method

[0010] It is applicable to regression and classification tasks and averages the results of the learners. The advantage of the averaging method is that it can reduce overfitting. Common averaging methods include: arithmetic averaging method, geometric averaging method, and weighted averaging method.

[0011] (3) Stacking method

[0012] The idea of the Stacking stacking method is based on the original data. Multiple base learners are trained, and then the prediction results of the base learners are combined into a new training set to train a new learner. That is, in the first layer, various machine learning algorithms are used, and the predicted values obtained are used as the input features of the meta-model in the second layer. Through the learning of the meta-model in the second layer, the final predicted value is output. This structure is beneficial for the second-layer model to correct the errors of the first-layer model.

[0013] (4) Blending method

[0014] The idea of the Blending mixing method is to divide the original data set into a smaller hold-out set. For example, 10% of the training set is left for training the original learner, and 90% of the data is used for training the base learners. In this way, the base learners and the meta-learner are trained with different data sets, thus avoiding information leakage and causing overfitting.

[0015] (5) Bagging method

[0016] Bagging is based on bootstrap (self-sampling), that is, sampling with replacement. The size of the training subset is the same as the size of the original data set. The training of the base learners can be carried out in parallel. For a training set of m samples, m random sampling operations with replacement are performed to obtain m sampling sets of samples. In this way, approximately 36.8% of the samples in the training set are not sampled. Repeating in the above manner, t data sets containing m samples can be collected, and thus t base learners can be trained. Finally, the outputs of these t base learners are combined.

[0017] (6) Boosting method

[0018] The Boosting method is a serial mechanism, that is, there is a dependency relationship between the training of individual learners, and the subsequent model will correct the prediction results of the previous model. Its basic idea is to increase the weight of the samples with prediction errors during the training of a base learner, so that subsequent base learners can pay more attention to these samples with large errors, correct these errors as much as possible, and continue serially downward until the required t base learners are generated. Finally, these t learners are combined with weights.

[0019] 2. Research on Encrypted Traffic Classification Method Based on Machine Learning;

[0020] Machine learning uses the knowledge obtained through inhibition as experience, trains a large amount of data to achieve the effort of flexibly processing various data, and applies the internal logic of the learned data to new data to achieve high-accuracy and high-precision prediction. Machine learning takes the form of an instance dataset as input, where an instance refers to an independent instance in the dataset, and each instance is characterized by its feature values, which measure different aspects of the instance. The dataset finally presents as a matrix of instances and features. For example, if the input data is labeled to establish a key-value pair mapping relationship between the input variables X and Y, then (X, Y) belongs to a supervised machine learning model, which is mainly used for classification and regression. Common algorithms of this type include decision trees, random forests, support vector machines, etc. If the input data has no prior processing, group the instances with similar features into clusters, and in association learning, look for any associations between features. This mode is unsupervised learning, and the predicted result is not a discrete class but a numerical quantity, which is mostly used for clustering. Common algorithms include K-Nearest Neighbor, PCA, K-Means algorithm, etc. The output of machine learning is a description of the learned knowledge, and how the specific results of the learning process are represented depends largely on the specific machine learning algorithm used.

[0021] Regarding the categories of encrypted traffic, there are encrypted protocols, abnormal encrypted traffic, application categories, encrypted services, etc. To perform refined identification, it is necessary to rely on means of machine learning or even deep learning for refined identification. Generally, the encryption object of encryption technology is payload information rather than traffic data characteristics, so that algorithms relying on statistical thinking and machine learning are less affected by encryption technology. Therefore, the mainstream idea of encrypted traffic identification technology is to train the model algorithm of machine learning, and the machine learning classification and identification method based on traffic statistical characteristics is widely used.

[0022] The technical means adopted are different for different objects of encrypted traffic identification. The key to identifying communication traffic lies in feature selection and extraction of the data packets, flow characteristics, and behavior characteristics of the identification object, which is also the key to optimizing the identification algorithm. However, during the encryption process, the characteristics of the data content are interfered, which greatly limits the optimization of the traffic identification algorithm.

[0023] The quality of the feature set is crucial for the performance of machine learning algorithms. Using irrelevant or redundant features is not conducive to the accuracy of most machine learning algorithms and may make the system computationally more expensive. Therefore, an ideal feature subset should be small enough but retain the key and necessary useful information.

[0024] 3. Traffic classification technology based on LGBM;

[0025] The Light Gradient Boosting Model (LGBM) was initially proposed by Microsoft and has many advantages of XGBT (Extreme Grandient Boosting Tree), such as good training effect and not being prone to overfitting. Its main idea is to use weak classifiers (decision trees) for iterative training to obtain the optimal model. The main difference between it and XGBT lies in the tree generation strategy. XGBT trees grow in a level-wise manner, while LGBM uses a leaf-wise algorithm with a depth limit. The Gradient-based One-Side Sampling (GOSS) algorithm and Exclusive Feature Bundling are the main reasons for LGBM to execute faster and with higher accuracy.

[0026] The GOSS algorithm is an important sampling strategy used in LightGBM to handle large-scale data. Its core idea is to reduce the number of training samples while maintaining the data distribution characteristics, thereby improving the training efficiency. Based on the observation that instances with larger gradients contribute more to model training, GOSS believes that the larger the gradient, the greater the prediction error of the current model for that instance and the more attention is needed. Retaining instances with large gradients can ensure that the model learns the key patterns.

[0027] EFB is used to handle high-dimensional sparse features. Its core idea is to bundle mutually exclusive features (features that rarely take non-zero values simultaneously) together, thereby reducing the number of features. In specific implementation, first, a feature conflict graph is constructed, the conflict degree (the frequency of being non-zero simultaneously) between any two features is calculated, and it is judged whether they are mutually exclusive according to the set threshold. Then, feature bundling is performed, converting the problem into a graph coloring problem, and the greedy algorithm is used to group the mutually exclusive features. Each group of features is bundled into a new feature. Finally, the bundled features are encoded to ensure that the value ranges of different features do not overlap and support feature restoration.

[0028] The decision tree algorithm for selecting histograms has the following basic idea: bin the feature values, discretize continuous floating-point feature values into k integers to form bins, and simultaneously construct a histogram with a width of k. Then traverse the data, accumulate statistical information in the histogram using the discrete values as indices, and then find the optimal splitting point by traversing the discrete values obtained from the histogram. Since the histogram algorithm does not require additional storage resources to save the pre-sorted results, only the discretized values are needed, so LightGBM can effectively reduce memory usage.

[0029] The QUIC protocol was proposed and developed by Google in 2013 to address issues such as long connection establishment time and head-of-line blocking in HTTP / 2.0. It is a low-latency transport layer protocol based on UDP. QUIC provides reliable transmission and can establish a connection within one RTT. QUIC has many functional designs superior to TCP-based transport protocols. It has functions such as congestion control, flow control, and packet loss recovery, and can manage states such as the establishment, maintenance, migration, and termination of network connections. QUIC also has TLS 1.3 built-in, uses the QUIC record layer instead of TLS 1.2 to encrypt and decrypt packets, and has higher security. QUIC uses a connection identifier CID (Connection ID) to represent a unique network flow, which gives QUIC the feature of connection migration. Therefore, designing an efficient QUIC traffic classification method to improve the accuracy of QUIC traffic classification will optimize network security, monitoring, and quality of service. However, due to features such as full verification, full confidentiality, 0-RTT connection establishment, connection migration, forward error correction, and multiplexing in QUIC, the feature dimensions extracted from QUIC traffic are fewer than those extracted from traditional protocols.

[0030] J. Luxembur et al. evaluated the classification effect of a QUIC encrypted traffic classifier based on LGBM and selected three models: 1) multi-modal based on convolutional neural networks; 2) LGBM; 3) IP-based classifier, and tested and evaluated the characteristics and accuracy of these three models. The experimental results show that the LGBM classifier has a better accuracy than mm-CNN within three weeks of training and performs better on more popular services (such as Google and Facebook), but has a poor recall rate (Luxemburk J, Hynek K, T. Encrypted Traffic Classification: The QUIC case [C]. 2023 7th Network Traffic Measurement and Analysis Conference (TMA), IEEE, 2023: 1 - 10.). S. Almuhammadi et al. studied and tested five different ensemble learning techniques to solve the QUIC network traffic classification problem (Almuhammadi S, Alnajim A, Ayub M. QUIC network traffic classification using ensemble machine learning techniques [J]. Applied Sciences, 2023, 13(8): 4725.). They trained the models with different numbers of features in different scenarios and conducted performance evaluations. The results showed that XGBT and LGBM outperformed other models, and LGBM outperformed other methods in terms of accuracy, precision, recall, and F1 - score, up to over 99%, and LGBM and XGBT still achieved a performance score of 92% using a small number of features (such as 15 packets).

[0031] 4. The QUIC encrypted traffic classification algorithm implemented based on the IP algorithm;

[0032] TCP and UDP provide multiplexing of multiple flows between public IP ports by using port numbers. In practical applications, many applications also use the "well - known" ports on the local host as the rendezvous points where other hosts can initiate communication. Based on the network - layer classifier, it only needs to look for TCP SYN packets (the first step of the TCP three - way handshake during session establishment) to know the server side of the new client - server TCP connection. Then, the application is inferred by looking up the destination port number of the TCP SYN packet in the registered port list of the Internet Assigned Number Authority (IANA). UDP is similar (although UDP does not establish or maintain connection status).

[0033] However, this method also has limitations. First, some applications may not register their ports with IANA (e.g., peer-to-peer applications such as Napster and Kazaa). Applications can use ports other than their well-known ports to avoid operating system access control restrictions (e.g., unprivileged users on Unix-like systems may be forced to run an HTTP server on a port other than port 80). Additionally, in some cases, server ports are dynamically allocated as needed. For example, RealVidel streaming allows for dynamic negotiation of the server port used for data transfer, which is negotiated on the initial TCP connection established using the well-known RealVideo control port.

[0034] Moore and Papagiannaki combined port-based and payload-based techniques to identify network applications (Moree A W. Toward the Accurate Identification of Network Applications[J]. Pam, 2005. DOI: doi:10.1007 / 978-3-540-31966-5_4.). The classification process starts by examining the port numbers of the flows. If well-known ports are not used, the flow is passed to the next stage. In the second stage, the first packet is examined to see if it contains a known signature. If none is found, the packet is examined to see if it contains a known protocol. If these tests fail, the protocol signature in the first Kbyte of the flow is studied. After this stage, unclassified traffic has its entire traffic payload examined. Their results showed that port information alone was able to correctly classify 69% of the total bytes. Including the information observed in the first Kbyte of each flow increased the accuracy to nearly 79%. Higher accuracy can only be achieved by examining the entire remaining unclassified payload. Although payload-based inspection avoids dependence on fixed port numbers, it brings greater complexity and processing load to traffic identification devices. The accuracy of the algorithm needs to keep pace with extensive knowledge of application protocol semantics and may require concurrent analysis of large volumes of traffic, and this method will face greater challenges when faced with encrypted traffic.

[0035] In the research of Nguyen Phong Hoang et al., it was found that even with encryption enabled, users would disclose information about the domains they accessed through DNS queries and the TLS Server Name Indication (SNI) extension (Hoang N P, AkhavanNiaki A, Borisov N, et al. Assessing the privacy benefits of domain name encryption[C]. Proceedings of the 15th ACM Asia Conference on Computer and Communications Security, 2020: 290-304.). They quantified the privacy gains provided by ESNI for different hosting and CDN through different metrics, namely the k-anonymity brought by co-hosting and the dynamics of IP address changes, and found that 20% of the domains studied and tested did not gain any privacy benefits because there would be a one-to-one mapping between their hostnames and IP addresses, and only 7.7% of the domains would change their hosted IP addresses daily.

[0036] Jan Luxemburk et al. designed an IP-based classifier (Luxemburk J, Hynek K, T. Encrypted Traffic Classification: The QUIC case[C]. 2023 7th Network Traffic Measurement and Analysis Conference (TMA), IEEE, 2023: 1-10.). During the training process, for each IP address and its / P IP prefix, the IP-based classifier algorithm stores the hosting service and the number of occurrences into a dictionary. For classification, an exact matching test will be conducted. When there is an unknown IP address or a complete matching failure due to multiple co-hosting services of a given IP (the score of the service with the most occurrences is less than exact_T), a subnetwork matching will be performed, and the service with the largest number of occurrences in the subnetwork will be selected. When the subnetwork does not exist in the training set, the classifier also does not make a prediction.

[0037] In summary, within the three weeks of training, the accuracy rate of the LGBM classifier is better than that of the mm-CNN, but the recall rate is poor, which means that LGBM performs better on more popular services such as Google and Facebook. However, for deep learning-based classifiers, the mm-CNN and LGBM classifiers often make mistakes in classification among the same service providers. For example, when classifying the traffic of Google Pay and Google Analytics, the performance of both classifiers will drop significantly. The IP-based classifier, on the other hand, has stable performance throughout the test. However, as long as the server changes all its IP addresses, the IP-based classifier will cause the recall rate to drop to 0 and is not applicable to the identification and classification of co-hosted service traffic. Moreover, compared with LGBM, the machine learning-based model is more vulnerable to data drift. Summary of the Invention

[0038] To solve the problems of difficulty in adapting to the actual network environment and low accuracy rate existing in the current QUIC traffic classification research, the present invention provides a QUIC encrypted traffic classification method based on multi-model fusion, aiming to fuse the QUIC traffic classification model based on LGBM and the QUIC traffic classification model based on IP through the mainstream model fusion method at the present stage, give full play to the advantages of the two technologies, complement each other's disadvantages, and realize a QUIC traffic classifier with high accuracy rate of the QUIC traffic classification model based on LGBM, high recall rate of the QUIC traffic classification model based on IP and stability.

[0039] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0040] A QUIC encrypted traffic classification method based on multi-model fusion provided by the present invention includes the following steps:

[0041] Step 1: Data preprocessing;

[0042] Step 2: Model training stage;

[0043] S201: Divide the training set and the validation set;

[0044] S202: Train each basic model in the high-precision model group and the high-recall model group with the training set;

[0045] S203: Evaluate the performance of each basic model with the validation set;

[0046] S204: If the dynamic weight mechanism is enabled, calculate the corresponding dynamic weight according to the performance of each basic model on the validation set;

[0047] S205: Normalize the obtained dynamic weights;

[0048] S206: Use multiple candidate thresholds to find the optimal decision threshold;

[0049] Step 3: Model prediction stage;

[0050] S301: Input the prediction data, and each base model generates its respective prediction probability;

[0051] S302: If the dynamic weight mechanism is enabled, the prediction probabilities of each base model will be weighted and averaged according to the dynamic weights obtained in the training stage for probability calibration;

[0052] S303: Combine the calibrated probabilities of each base model;

[0053] S304: Judge the combined calibrated probabilities through the optimal decision threshold to generate the final classification prediction result.

[0054] Furthermore, in Step 1, the data is sourced from the CESNET - QUIC22 dataset.

[0055] Furthermore, in Step 1, the method of data preprocessing is: standardize the data, normalize the packet histogram, and perform robust scaling on FLOWSTATS.

[0056] Furthermore, the base model in the high - precision model group is a QUIC traffic classification model based on LGBM; the base model in the high - recall model group is a QUIC traffic classification model based on IP.

[0057] Furthermore, in Step S203, by sequentially traversing all the base models in the high - precision model group and the high - recall model group, each base model is independently trained.

[0058] Furthermore, in Step S204, for the high - precision model group, the weights are calculated using the training accuracy values; for the high - recall model group, the weights are calculated using the recall rate.

[0059] Furthermore, in Step S206, using the voting threshold optimization method, optimize the decision threshold through voting and search for the optimal decision threshold on the validation set to improve the model performance.

[0060] Furthermore, in Step S206, select the threshold that can obtain the best F1 - score as the optimal decision threshold.

[0061] Furthermore, in Step 3, design a classification probability prediction method based on the IP address, and perform classification prediction through the network affiliation relationship of the IP address.

[0062] Furthermore, during the classification prediction process, through a two-level search strategy, the first level searches through the website's network, and the second level searches through network prefixes. After finding, count the total number of tags in this category, calculate the proportion of each tag. If it is not found in both the IP dictionary and the network prefix dictionary, the probability is set to 0.

[0063] The beneficial effects of the present invention are:

[0064] Based on the multi-model fusion mechanism of voting, the present invention proposes a QUIC encrypted traffic classification algorithm based on LGBM-IP, combines the advantages of different models, complements each other's learned domain knowledge, and averages their respective noise differences, thereby reducing the risk of overfitting and poor generalization ability of a single model, so as to achieve a QUIC encrypted traffic classification with higher accuracy and stability. Description of the Drawings

[0065] Figure 1 It is a schematic diagram of a QUIC encrypted traffic classification method based on multi-model fusion provided by the present invention.

[0066] Figure 2 It is a flowchart of a QUIC encrypted traffic classification method based on multi-model fusion provided by the present invention. Detailed Embodiments

[0067] The following further elaborates on the present invention in conjunction with the drawings.

[0068] The main work of the present invention is to study the use of multi-model fusion technology, and fuse the LGBM algorithm in machine learning and the IP-based classifier to achieve high-accuracy and stable service-level classification of encrypted traffic based on the QUIC protocol. Since under the LGBM model, the existing QUIC traffic classification methods have high accuracy but poor recall rate and poor stability; while under the IP model, the QUIC traffic classification method has high stability and good accuracy, but is not applicable to the identification and classification of co-hosted service traffic. To integrate the advantages of the two QUIC encrypted traffic classifications, the present invention introduces a multi-model fusion mechanism based on soft voting, selects the QUIC traffic classification model based on LGBM and the QUIC traffic classification model based on IP as sub-models to obtain the best results. The present invention uses the lightweight gradient boosting model LGBM, the IP address-based classification algorithm, and multi-model fusion for the scheme design, meeting the two key requirements of high efficiency and stability for QUIC encrypted traffic classification, and improving the classification performance through four key technologies: hierarchical prediction mechanism, dynamic weight mechanism, probability calibration, and voting threshold optimization.

[0069] A QUIC encrypted traffic classification method based on multi-model fusion provided by the present invention specifically uses a multi-model fusion classification algorithm to achieve fine-grained classification of the application type level of QUIC encrypted traffic.

[0070] The multi - model fusion classification algorithm adopted by the present invention mainly consists of four parts, namely, a hierarchical prediction mechanism, a dynamic weight mechanism, probability calibration, and voting threshold optimization.

[0071] First, use the hierarchical prediction mechanism based on soft voting. By setting a confidence threshold to distinguish different prediction scenarios, different decision strategies are adopted for different confidence intervals.

[0072] Then, through the dynamic weight mechanism, evaluate the performance of each base model in the high - precision model group and the high - recall model group on the validation set, and decide whether to use dynamic weights based on the evaluation results of the validation set (if so, calculate the dynamic weights; if not, use equal weights). By dynamically adjusting the weights of each base model during the fusion process, ensure that better models can play a greater role.

[0073] Secondly, use the probability calibration method to calibrate the prediction probabilities of each base model in the high - precision model group and the high - recall model group by weighted average to improve the possibility of the final prediction.

[0074] Finally, use the voting threshold optimization method to optimize the decision threshold through voting, that is, search for the optimal decision threshold on the validation set to further improve the performance of the model.

[0075] The core idea of the multi - model fusion classification algorithm is how to more efficiently combine the advantages of the base models in the high - precision model group and the high - recall model group to improve the overall classification and prediction performance of the algorithm. This multi - model fusion classification algorithm adopts the idea of "precision - recall" balance, allowing the algorithm to trade - off between precision and comprehensiveness in different scenarios, so that the fusion result can take into account two important evaluation dimensions. In addition, the algorithm introduces a dynamic weight mechanism to dynamically adjust the weights of each base model during the fusion process based on the actual performance of each base model on the validation set, ensuring that better - performing models can play a greater role while suppressing the negative impact of poorly - performing models. This multi - model fusion classification algorithm uses a probability calibration mechanism to perform weighted average on the prediction accuracies of each base model to ensure the reliability of the final prediction. This multi - model fusion classification algorithm also uses the voting threshold optimization method to find the optimal decision threshold on the validation set instead of simply fixing it at the average threshold, which can adapt to more multi - model fusion scenarios.

[0076] See Figure 1 and Figure 2 For illustration, a QUIC encrypted traffic classification method based on multi - model fusion provided by the present invention has the following specific implementation process:

[0077] Step S1: Data pre - processing;

[0078] In the present invention, the QUIC traffic dataset is specifically selected as the CESNET-QUIC22 dataset. The data is standardized through the data processing layer, the packet histogram is normalized, and FLOWSTATS is robustly scaled to suppress the negative impact of outliers.

[0079] Among them, when standardizing the data, the packet size and time can be specifically standardized using the mean and variance (z-score normalization). The packet direction is encoded as ±1 and does not need to be standardized.

[0080] Among them, after normalizing the packet histogram, the sum of the bins of each packet histogram is 1.

[0081] In the present invention, for FLOWSTATS features without a fixed range, such as the number of transmitted packets, robust scaling is selected, and the median and interquartile range are used instead of the mean and variance to limit the negative impact of outliers. The FLOWSTATS features are standardized to the 0.99 quantile to reduce extreme values.

[0082] After data preprocessing, the packet time is compressed to a maximum of 15 seconds; the packet size is trimmed to a maximum of 1460 bytes, which is the maximum size of an unfragmented UDP packet when using the typical MTU of 1500.

[0083] In the present invention, the original QUIC traffic dataset is first preprocessed, using the metadata sequence (PSTATS) and traffic statistics (FLOWSTATS) features in the dataset without using the traffic payload. The metadata sequence is the sequence of the packet size, direction, and inter-packet time of the first 30 packets in each flow, and the traffic statistics contain the summary information of the bidirectional flow, such as the number of bytes transmitted, the number of packets, the duration of the flow, and the reason for flow export in both directions. In the data preprocessing stage, the packet size and time are standardized using the mean and variance (z-score normalization), the packet histogram is normalized so that the sum of the bins of each histogram is 1, and for the FLOWSTATS feature data without a fixed range, such as the number of transmitted packets, the file is scaled, and the median and interquartile range are used instead of the mean and variance to limit the negative impact of outliers.

[0084] Step S2: Model training stage;

[0085] In the model training stage, the preprocessed complete dataset is divided into a training set and a validation set according to the division ratio parameter of the validation set. The training set is used to train the high-precision model group and the high-recall model group respectively, ensuring that the models can be fully verified during the training process and avoiding overfitting problems. Then, the validation set is used to evaluate the performance of each basic model in the high-precision model group and the high-recall model group (the basic model in the high-precision model group is the QUIC traffic classification model based on LGBM, and the basic model in the high-recall model group is the QUIC traffic classification model based on IP), that is, by traversing all the basic models in the high-precision model group and the high-recall model group in turn, and independently training each basic model. If the dynamic weight mechanism is enabled, the corresponding dynamic weight of each basic model is calculated according to its performance on the validation set. Among them, for the high-precision model group, the weight is calculated using the training precision value (precision_score), and for the high-recall model group, the recall rate (recall_score) is used to calculate the weight, and the obtained weights are normalized. After completing the training of the basic models and the calculation of the dynamic weights, multiple candidate thresholds obtained through multiple experiments are used to find the optimal decision threshold to balance the prediction precision and recall rate, and finally the threshold that can obtain the best F1 score is selected as the final decision threshold.

[0086] Step S3: Model prediction stage;

[0087] In the model prediction stage, the processing flow of the algorithm is more intuitive. First, by inputting the prediction data, the prediction probabilities are generated by each basic model in the high-precision model group and the high-recall model group respectively. If the dynamic weight mechanism is enabled, the prediction probability of each generated basic model will be weighted and averaged according to the normalized weight obtained in the training stage for probability calibration. Then, by combining the calibration probabilities of each basic model in the high-precision model group and the high-recall model group, the reliability of the prediction result is ensured. Finally, the combined calibration probability is judged by the optimal decision threshold to generate the final classification prediction result.

[0088] The idea of multi-model fusion is similar to Ensemble Learning, both of which are to construct and combine multiple learners to learn tasks. However, in Ensemble Learning, the learners are homogeneous, while in model fusion, the learners are heterogeneous. The present invention uses the weighted voting method. The idea of the weighted voting method is to assign different weights to the votes of different models. The weight setting can be based on the performance score of the model on the validation set, expert experience setting, and cross-validation results. The final prediction result is determined by the highest score of the weighted vote.

[0089] Based on the IP-based classifier algorithm, the present invention designs a classification probability prediction method based on IP addresses. The core technical route is to perform classification prediction through the network attribution relationship of IP addresses. During the prediction process, through a two-level lookup strategy, the first level is to search through the network of the website, and the second level is to search through the network prefix. After finding it, count the number of all labels under this category, calculate the proportion of each label. If it is not found in both the IP dictionary (IPDict) and the network prefix dictionary (PrefixDict), the probability is set to 0.

[0090] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A QUIC encrypted traffic classification method based on multi-model fusion, characterized in that: The following steps are involved: Step 1: Data preprocessing; Step 2: Model training phase; S201: Divide the training set and the validation set; S202: training each basic model in the high-precision model group and the high-recall model group using the training set; S203: Evaluate the performance of each base model using the validation set; S204: If the dynamic weight mechanism is enabled, the corresponding dynamic weight of each basic model is calculated according to its performance on the validation set; S205: normalizing the obtained dynamic weight; S206: using multiple candidate thresholds to find an optimal decision threshold; Step 3: Model prediction stage; S301: input prediction data, and each basic model generates its own prediction probability; S302: If the dynamic weight mechanism is enabled, the predicted probability of each basic model will be weighted averaged according to the dynamic weight obtained in the training phase to perform probability calibration; S303: combining the calibration probabilities of each basic model; S304: The combined calibration probability is judged by an optimal decision threshold to generate a final classification prediction result.

2. The QUIC encrypted traffic classification method based on multi-model fusion according to claim 1 is characterized in that: In step 1, the data comes from the CESNET-QUIC22 dataset.

3. The QUIC encrypted traffic classification method based on multi-model fusion according to claim 1 is characterized in that: In step 1, the data preprocessing method is: standardizing the data, normalizing the data packet histogram and robustly scaling FLOWSTATS.

4. The QUIC encrypted traffic classification method based on multi-model fusion according to claim 1 is characterized in that: The basic model in the high-precision model group is a QUIC traffic classification model based on LGBM; the basic model in the high-recall model group is a QUIC traffic classification model based on IP.

5. The QUIC encrypted traffic classification method based on multi-model fusion according to claim 1 is characterized in that: In step S203, all basic models in the high-precision model group and the high-recall model group are traversed in sequence, and each basic model is trained independently.

6. The QUIC encrypted traffic classification method based on multi-model fusion according to claim 1 is characterized in that: In step S204, the weights are calculated using the trained precision values ​​for the high-precision model group, and the weights are calculated using the recall rates for the high-recall model group.

7. The QUIC encrypted traffic classification method based on multi-model fusion according to claim 1 is characterized in that: In step S206, the voting threshold optimization method is used to optimize the decision threshold by voting, and the optimal decision threshold is searched on the validation set to improve the model performance.

8. The QUIC encrypted traffic classification method based on multi-model fusion according to claim 1 is characterized in that: In step S206, a threshold that can obtain the best F1 score is selected as the optimal decision threshold.

9. The QUIC encrypted traffic classification method based on multi-model fusion according to claim 1 is characterized in that: In step three, a classification probability prediction method based on IP addresses is designed to perform classification prediction based on the network affiliation of IP addresses.

10. The QUIC encrypted traffic classification method based on multi-model fusion according to claim 9 is characterized in that: In the classification prediction process, a two-level search strategy is used. The first level is to search through the website's network, and the second level is to search through the network prefix. After finding the label, the number of all labels in the category is counted, and the proportion of each label is calculated. If it is not found in both the IP dictionary and the network prefix dictionary, the probability is set to 0.

Citation Information

Cited By

  • Intelligent decision-making method and device based on multi-modal data fusion

    CN121256670A

  • Generated power prediction method and device, electronic equipment and storage medium

    CN121352147A

  • Roof illegal building identification method and device for urban existing building

    CN121904595A

  • A method and device for identifying illegal buildings on roofs of existing buildings in cities

    CN121904595B

  • Event recognition model training method and device and event recognition method and device

    CN122244599A