An Uncertainty Quantification Method for Encrypted Network Traffic Classification Prediction

By using kernel density estimation and KL divergence to calculate OOD scores in the encrypted network traffic classification model, the problem of quantitative analysis of uncertainty in the encrypted network traffic in the prior art is solved, the credibility and generalization performance of the model are improved, and the reliability of network traffic classification is enhanced.

CN118353842BActive Publication Date: 2025-06-13CHINESE PEOPLES LIBERATION ARMY UNIT 61660
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410483175.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-22
Publication Date
2025-06-13
Estimated Expiration
2044-04-22

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively conduct uncertainty quantitative analysis of encrypted network traffic, resulting in poor generalization performance of the model, and the existing methods have problems such as high computational complexity, balance of interpretability and performance, low generalization, and insufficient credibility test.

Method used

An uncertainty quantization method is adopted, including model training prediction and uncertainty quantization steps. The specific steps include using the trained machine learning model to predict the test sample, calculating the density distribution of the test sample and the training data set sample through the kernel density estimation method, measuring the similarity between the distributions using KL divergence, and converting the KL divergence into OOD scores to judge the confidence of the prediction results.

Benefits of technology

This method can effectively quantify the uncertainty of the prediction results of the encrypted network traffic classification model, improve the credibility and interpretability of the model, enhance the generalization performance and robustness of the model, and improve the reliability of network traffic classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118353842B_ABST
    Figure CN118353842B_ABST
Patent Text Reader

Abstract

The present invention relates to an uncertainty quantification method for encrypted network traffic classification prediction, belonging to the field of network security. The method of the present invention includes steps such as model training and prediction, Gaussian kernel density estimation, KL divergence calculation, OOD score calculation, hypothesis testing, etc. The present invention can perform uncertainty quantification analysis on the prediction results of existing encrypted network traffic classification models, detect the credibility of their prediction results, and evaluate the performance of existing models. During the prediction process, the method can quantify the uncertainty of the prediction results, calculate the OOD score through Gaussian kernel density estimation and KL divergence, measure the similarity and difference between in-distribution and out-of-distribution samples, help the system identify and adapt to new data, thereby improving the scalability and robustness of the model. The present invention improves the reliability and interpretability of network traffic classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network security, and particularly relates to an uncertainty quantification method for encrypted network traffic classification prediction. Background Art

[0002] With the rapid development of Internet technology, the public's network security awareness has been continuously improved, and the security and privacy protection of network communication have become increasingly important. Network traffic encryption technology has become one of the key means to maintain communication security. Technologies for identifying and classifying encrypted traffic have emerged as the times require. Machine learning and deep learning models have also been introduced into the field of network security for network traffic research. With the continuous in-depth research on traffic classification, mainstream traffic classification methods can achieve the classification and recognition of different types of traffic and can obtain good results. However, since machine learning models usually need to predict new data during the traffic classification process, traffic data often has uncertainty and diversity, so the prediction results of the models will also have uncertainty. At this time, the generalization performance of machine learning models will be relatively poor. In order to enable the model to make accurate predictions and adapt to new traffic data, it is necessary to perform uncertainty quantification analysis on the prediction results of the model.

[0003] Current research uses Explainable Artificial Intelligence (XAI) technology for network traffic uncertainty quantification, which can study interpretability and credibility through XAI technology and provide sample-based and global explanations; there is uncertainty quantification for the feature description of network traffic data; there is the use of Out-of-Distribution (OOD) detection methods for the quantification analysis of prediction results. However, these technologies all have different limitations, and problems such as high computational complexity, trade-off between interpretability and performance, low generalization, and insufficient credibility testing will occur. These problems are still a huge challenge for the uncertainty quantification analysis of network traffic. And at the current stage, research rarely involves the uncertainty quantification analysis of traffic classification results, far less in-depth than in the field of computer vision. Therefore, the research on uncertainty quantification in the field of encrypted network traffic classification needs to be further deepened, and the existing technologies need to be improved and enriched.

[0004] The present invention proposes an uncertainty quantification method for encrypted network traffic classification prediction, aiming to perform uncertainty quantification analysis on the prediction results of existing encrypted network traffic classification models and study the interpretability and credibility of the models. Summary of the Invention

[0005] (1) Technical Problems to be Solved

[0006] The technical problem to be solved by the present invention is how to provide an uncertainty quantification method for encrypted network traffic classification prediction to solve the problem of uncertainty quantification analysis of network traffic.

[0007] (2) Technical Solution

[0008] To solve the above technical problems, the present invention proposes an uncertainty quantification method for encrypted network traffic classification prediction, and the method includes the following steps:

[0009] S1. Model Training and Prediction

[0010] Use a machine learning model for encrypted network traffic classification to perform training and prediction;

[0011] S2. Uncertainty Quantification

[0012] S21. Use the trained machine learning model to predict the test samples, and determine the category to which the test samples belong according to the prediction results;

[0013] S22. Then use the kernel density estimation method to calculate the density distributions of the test samples and the samples in the training data set;

[0014] S23. After obtaining the density distributions of the above test samples and the samples in the training data set, use the KL divergence to measure the similarity between the two distributions;

[0015] S24. Use a conversion function to convert the KL divergence of the test sample and the training data set sample distributions into an OOD score; judge whether to accept or reject the null hypothesis according to the OOD score.

[0016] (3) Beneficial Effects

[0017] The present invention proposes an uncertainty quantification method for encrypted network traffic classification prediction, which can perform uncertainty quantification analysis on the prediction results of existing encrypted network traffic classification models, detect the credibility of their prediction results, and evaluate the performance of existing models. During the prediction process, the method can quantify the uncertainty of the prediction results, calculate the OOD score through Gaussian kernel density estimation and KL divergence, measure the similarity and difference between in-distribution and out-of-distribution samples, help the system identify and adapt to new data, thereby improving the scalability and robustness of the model. The present invention proposes an uncertainty quantification method for existing encrypted network traffic classification machine learning models, which can provide uncertainty quantification analysis of the classification prediction results given by the model, and improve the reliability and interpretability of network traffic classification. Description of the Drawings

[0018] Figure 1 is the overall architecture of the present invention;

[0019] Figure 2 is the uncertainty quantification flowchart of the present invention. Detailed Embodiments

[0020] To make the objectives, content, and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention with reference to the accompanying drawings and embodiments.

[0021] The present invention adopts an uncertainty quantification method for encrypted network traffic classification prediction, which can perform uncertainty quantification analysis on the prediction results of existing encrypted network traffic classification models, detect the credibility of their prediction results, and evaluate the performance of existing models. By using Gaussian kernel density estimation and KL divergence to calculate the OOD score, the similarity and difference between in-distribution and out-of-distribution samples are measured, helping the system identify and adapt to new data, thereby improving the generalization and robustness of the model, and enhancing the reliability and interpretability of network traffic classification.

[0022] The present invention proposes an uncertainty quantification method for encrypted network traffic classification prediction, and this method includes the following steps:

[0023] S1. Model training and prediction

[0024] As an embodiment of the present invention, the main innovation of the present invention lies in the uncertainty quantification analysis method for encrypted network traffic classification prediction. The design and training of the classification model are not the focus of the present invention. Therefore, no innovation will be made to it in the present invention. Only existing machine learning models for encrypted network traffic classification will be used for training and prediction, and then uncertainty quantification analysis will be performed on the prediction results of the model. Similarly, no requirements are made for the acquisition and processing of encrypted network traffic data.

[0025] S2. Uncertainty quantification

[0026] Training a large amount of traffic data on existing machine learning models can obtain a trained traffic classification model. However, the quality of the model performance needs to be evaluated based on the visualization results. Therefore, the problem of uncertainty quantification needs to be considered in the above classification model. As an embodiment of the present invention, the OOD score is used to perform uncertainty quantification analysis on the prediction results of the model.

[0027] As an embodiment of the present invention, the null hypothesis and the alternative hypothesis should be determined before performing uncertainty quantification analysis on the prediction results of the test samples. Here, the null hypothesis and the alternative hypothesis are respectively set as:

[0028] H 0 : There is no significant difference between the test sample and the samples in the training dataset;

[0029] H 1 : There is a significant difference between the test sample and the samples in the training dataset;

[0030] S21. As an embodiment of the present invention, first, a trained machine learning model is used to predict a test sample, and the category to which the test sample belongs is determined according to the prediction result.

[0031] Next, the kernel density estimation method will be used to calculate the density distributions of the test sample and the training dataset samples, and the uncertainty of the model prediction result will be measured.

[0032] S22. Use the kernel density estimation method to calculate the density distributions of the test sample and the training dataset samples.

[0033] Kernel density estimation is a non-parametric method for estimating the probability density function of sample points. In the present invention, the Gaussian kernel function is used to calculate the kernel density distribution estimation. The Gaussian kernel function can be expressed as:

[0034]

[0035] where t is the sample point of the input data. Substituting into the formula, we can get:

[0036]

[0037] where x m represents any sample point in the training dataset, x j represents the j-th sample point in the training dataset, and h represents the bandwidth of the Gaussian kernel function.

[0038] x m The Gaussian kernel density distribution estimation of can be expressed as:

[0039]

[0040] where x m ≠x j , and N represents the number of samples in the training dataset.

[0041] Similarly, the Gaussian kernel density distribution estimation of the test sample point x t can be calculated as:

[0042]

[0043] Formula (3) obtains the Gaussian kernel density distribution estimation of x m by using the distance between the sample point x m in the training dataset and the remaining sample points in the training dataset, while formula (4) obtains the Gaussian kernel density distribution estimation of x t by using the distance between the test sample point x t and the sample points in the training dataset;

[0044] S23. After obtaining the density distributions of the above-mentioned test samples and training dataset samples, use the KL divergence to measure the similarity between the two distributions;

[0045] After obtaining the density distributions of the above-mentioned test samples and training dataset samples, use the KL divergence to measure the similarity between the two distributions:

[0046]

[0047] For the convenience of measurement, use the normalization technique to map the KL divergence to the interval [0, 1]:

[0048]

[0049] By calculating the KL divergence, the distribution conditions of the test sample points and the training data sample points can be evaluated. If the KL divergence is large, the difference between the two distributions is large, and the test sample is more likely to belong to an unknown category; on the contrary, if the KL divergence is small, it means that the similarity between the two distributions is high, and the test sample is more likely to belong to the predicted category.

[0050] S24. Use a conversion function to convert the KL divergence between the test sample and the training dataset sample distribution into an OOD score, and judge whether to accept or reject the null hypothesis according to the OOD score.

[0051] As an embodiment of the present invention, in order to make the test process of the prediction result easier to understand, use a conversion function to convert the KL divergence between the test sample points and the training dataset sample points distribution into an OOD score:

[0052]

[0053] KL divergence D * KL The smaller it is, the smaller the OOD score is, indicating that the test sample is more likely to belong to the predicted category, and the confidence level of the model prediction result is high, and the null hypothesis should be accepted; on the contrary, KL divergence D * KL The larger it is, the larger the OOD score is, indicating that the test sample is more likely to belong to an unknown category, and the confidence level of the model prediction result is low, and the null hypothesis should be rejected.

[0054] As an embodiment of the present invention, Gaussian kernel density estimation and KL divergence are two key components used in the network to calculate the OOD score. Gaussian kernel density estimation is used to obtain the probability density values of the training dataset samples and test samples relative to the entire traffic dataset, and the KL divergence is used to measure the similarity between the training dataset samples and the test sample distribution and convert it into an interpretable outlier score. The combination of these components helps the model identify and handle the differences between in-distribution and out-of-distribution samples, improving the performance of the model. The specific process of uncertainty quantification is asFigure 2 as shown

[0055] By adding the above-mentioned uncertainty quantification process to the existing encrypted network traffic classification machine learning model and performing uncertainty quantification analysis on the classification prediction results of the model, the reliability and stability of the model can be evaluated, and the confidence level of the model prediction results in different situations can be determined; the unknown traffic data can be classified and predicted, its distribution can be determined, and the model can be helped to adapt to new data, improving the expansion ability and adaptability of the classification model.

[0056] The present invention proposes an uncertainty quantification method for encrypted network traffic classification prediction, which can perform uncertainty quantification analysis on the prediction results of the existing encrypted network traffic classification model, detect the credibility of its prediction results, and evaluate the performance of the existing model. During the prediction process, this method can quantify the uncertainty of the prediction results, calculate the OOD score through Gaussian kernel density estimation and KL divergence, measure the similarity and difference between in-distribution and out-of-distribution samples, help the system identify and adapt to new data, thereby improving the scalability and robustness of the model. The present invention proposes an uncertainty quantification method for the existing encrypted network traffic classification machine learning model, which can provide uncertainty quantification analysis of the prediction results for the classification prediction results given by the model, improving the reliability and interpretability of network traffic classification.

[0057] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for quantifying uncertainty in classification prediction of encrypted network traffic, characterized in that: The method comprises the following steps: S1. Model training prediction Use machine learning models for encrypted network traffic classification for training predictions; S2. Uncertainty Quantification S21. Use the trained machine learning model to predict the test sample, and determine the category to which the test sample belongs based on the prediction result; S22, then use the kernel density estimation method to calculate the density distribution of the test sample and the training data set sample; S23, after obtaining the density distribution of the test sample and the training data set sample, use KL divergence to measure the similarity between the density distribution of the test sample and the training data set sample; S24, using a conversion function to convert the KL divergence of the density distribution of the test sample and the training data set sample into an OOD score; judging whether to accept or reject the null hypothesis according to the OOD score; in, In S22, a Gaussian kernel function is used to calculate a kernel density distribution estimate; The Gaussian kernel function is expressed as: Where t is the sample point of the input data; Substituting into formula (1), we get: Among them, x m Represents any sample point in the training data set, x j represents the jth sample point in the training data set, and h represents the bandwidth of the Gaussian kernel function; In S22, x m The Gaussian kernel density distribution estimate is expressed as: Among them, x m ≠x j , N represents the number of samples in the training data set, and formula (3) is obtained by using the sample points x in the training data set m The distance between the rest of the training data set sample points is x m Gaussian kernel density distribution estimation of ; In S22, the test sample point x t The Gaussian kernel density distribution of is estimated as: Formula (4) is obtained by testing the sample point x t The distance between the training data set sample points is obtained by t Gaussian kernel density distribution estimation of ; The S23 specifically includes: after obtaining the density distribution of the test sample and the training data set sample, using KL divergence to measure the similarity between the density distribution of the test sample and the training data set sample:

2. The uncertainty quantification method for classification prediction of encrypted network traffic according to claim 1, characterized in that: The step before S21 also includes: before performing uncertainty quantitative analysis on the prediction results of the test sample, the null hypothesis and the alternative hypothesis should be determined, and the null hypothesis and the alternative hypothesis are respectively set as: H0: There is no significant difference between the test sample and the training dataset sample; H1: There are significant differences between the test samples and the training dataset samples.

3. The uncertainty quantification method for classification prediction of encrypted network traffic according to claim 1, characterized in that: The S23 further includes: using a normalization technique to map the KL divergence to the [0,1] interval: By calculating the KL divergence, the distribution of the test sample points and the training data sample points is evaluated; if the KL divergence is large, the difference between the two distributions is large, and the test sample is more likely to belong to the unknown category; on the contrary, if the KL divergence is small, it means that the similarity between the two distributions is high, and the test sample is more likely to belong to the predicted category.

4. The uncertainty quantification method for classification prediction of encrypted network traffic according to claim 3, characterized in that: The S24 specifically includes: using a conversion function to convert the KL divergence of the density distribution of the test sample point and the training data set sample point into an OOD score: KL divergence D * KL The smaller the KL divergence D, the smaller the OOD score, which means that the test sample is more likely to belong to the predicted category, the confidence of the model prediction result is higher, and the null hypothesis should be accepted; on the contrary, the KL divergence D * KL The larger the value, the larger the OOD score, which means that the test sample is more likely to belong to the unknown category, the confidence in the model prediction result is lower, and the null hypothesis should be rejected.

5. The uncertainty quantification method for classification prediction of encrypted network traffic according to claim 4, characterized in that: Gaussian kernel density estimation is used to obtain the probability density values ​​of the training dataset samples and the test samples relative to the entire traffic dataset. KL divergence is used to measure the similarity between the distributions of the training dataset samples and the test samples and convert them into interpretable outlier scores.

Citation Information

Patent Citations

  • Encrypted traffic classification method based on classifier and network structure

    CN113746707A

  • Wind turbine generator regulation rate prediction method and system based on variational inference

    CN115130743A