Efficient open set encryption traffic identification method for super-large scale traffic
Through the improved MobileNetV3 model and incremental learning technology, combined with PCA dimensionality reduction and confidence-based open set recognition algorithm, the challenge of encrypted traffic classification in resource-constrained and dynamic network environments is solved, and efficient encrypted traffic recognition and highly adaptable classification performance are achieved.
Patent Information
- Application Number
- CN202510280551.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing encrypted traffic classification methods have high demand for computing resources in resource-constrained scenarios and are difficult to adapt to category drift and new categories in dynamic network environments.
The improved MobileNetV3 model is used to combine PCA dimensionality reduction, confidence-based open set recognition algorithm and incremental learning technology of knowledge distillation to extract features and classify, identify drift samples and unknown categories, update model parameters and retain old knowledge.
While maintaining low resource consumption, the model's adaptability in an open-set environment is significantly improved, the adaptability and classification performance to complex feature distributions is improved, and it is suitable for encrypted traffic classification in lightweight network environments.
Smart Images

Figure CN120151017A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security, and particularly to an efficient open-set encrypted traffic recognition method for ultra-large-scale traffic. Background Art
[0002] With the popularization of network encryption technology, encrypted traffic has become the mainstream of network communication. According to the Google Transparency Report, as of 2023, encrypted HTTPS traffic has accounted for more than 95% of the total network traffic. Encrypted traffic classification technology analyzes the statistical characteristics of traffic (such as packet length distribution, flow duration, arrival time interval, etc.) to identify traffic behavior types, which is the key to ensuring network security. However, traditional encrypted traffic classification methods face two major challenges:
[0003] High computational resource requirements: Complex deep learning models (such as ResNet, VGG) perform well in feature extraction and classification, but their high computational resource requirements (for example, ResNet-50 requires 2.35 billion floating-point operations) limit their application in resource-constrained scenarios (such as edge computing, Internet of Things).
[0004] Insufficient adaptability to dynamic network environments: In the actual network environment, traffic characteristics often change dynamically (class drift), and new types of attack traffic continuously emerge, which poses a major challenge to static classification models.
[0005] To solve the problem of high computational cost of deep learning models, lightweight models (such as MobileNet, ShuffleNet) have become a research hotspot. These models have achieved a reduction in computational complexity through the following techniques: Depthwise separable convolution decomposes the standard convolution into depthwise convolution and pointwise convolution, significantly reducing the number of parameters and computational overhead. Network architecture search (NAS) searches for the best model structure to achieve performance similar to that of complex models with fewer resource requirements. The attention mechanism SE module (Squeeze-and-Excitation) improves the model's attention to key information by reallocating the weights of feature channels.
[0006] Regarding the problems of class drift and new classes in dynamic network environments, incremental learning and open-set recognition techniques have gradually been applied to encrypted traffic classification tasks. Incremental learning aims to learn new classes without forgetting old classes, and open-set recognition identifies unknown classes through confidence analysis. These two techniques provide new directions for dynamic encrypted traffic classification, but existing methods still have limitations in terms of efficiency and classification performance. Summary of the Invention
[0007] To solve the technical problems of the existing technologies in terms of resource efficiency, feature expression, and open-set recognition, the present invention provides an efficient open-set encrypted traffic recognition method for ultra-large-scale traffic, including:
[0008] S1. Input a network traffic data set, and generate reduced-dimension data with the optimal dimension through the PCA dimensionality reduction algorithm;
[0009] S2. Input the reduced-dimension data with the optimal dimension into an improved MobileNetV3 model, extract features and output classification results;
[0010] S3. Use an open-set recognition algorithm based on confidence to determine whether the classification result belongs to a known category, a drifting sample, or an unknown category;
[0011] S4. If the classification result is a drifting sample or an unknown category, initiate incremental learning based on knowledge distillation, update the parameters of the improved MobileNetV3 model, and retain the old knowledge.
[0012] Further, S1 specifically includes the following steps:
[0013] S11. Initialize the network traffic data set;
[0014] S12. Check whether the initial dimension is greater than 10. If the initial dimension is greater than 10, set the score to 0, set the current dimensionality reduction target dimension i to 10, and then execute the next step. Otherwise, end the process and perform performance verification on the reduced-dimension data with the optimal dimension;
[0015] S13. Select i as the dimensionality reduction target for analysis;
[0016] S14. Detect whether the selected i dimension is sufficient to retain the information of the original data by calculating the variance explanation rate. If the variance explanation rate is less than 0.99 and does not meet the requirements, this dimension is not suitable for dimensionality reduction, and directly execute S18; if the conditions are met, continue with the next operation;
[0017] S15. Ensure that the original data is not modified through the deep copy method, and then apply dimensionality reduction operations to it to obtain a reduced-dimension data set;
[0018] S16. Divide the reduced-dimension data set into a training set and a test set;
[0019] S17. Use the LMCA network to extract features and classify the reduced-dimension data set;
[0020] S18. Calculate and record the accuracy, F1 score, model size, and inference time of the LMCA network under this dimensionality reduction configuration, and then return to step S13 to re-select a new i value;
[0021] S19. Among all the different i-value configurations tried, compare the scores in each case and select the optimal configuration plan from them;
[0022] S110. Conduct final performance verification on the data of the optimal configuration plan.
[0023] Furthermore, in S2, the improved MobileNetV3 model includes:
[0024] An initial feature extraction layer that uses a 3×3 convolutional kernel to expand the input channels from 1 to 16, combined with the BatchNorm2d and HSwish activation functions to maintain the spatial resolution of the original features;
[0025] A feature learning backbone composed of 8 MobileBottleneck modules connected in series, gradually adjusting the number of channels to 16, 24, 40, and 80, and selectively using the SE attention mechanism to enhance the feature extraction ability;
[0026] A feature enhancement layer that expands the number of channels from 80 to 160 through 1×1 convolution to further fuse and enhance the features;
[0027] A global feature modeling layer that uses adaptive average pooling to compress the 3×3 feature map into 1×1 to extract global context information;
[0028] A classification head that adopts a three-layer fully connected network structure, with a dimension reduction design of 160→128→64→11, and outputs the probability distributions of 11 categories.
[0029] Furthermore, in S3, the confidence-based open-set recognition algorithm includes:
[0030] S31. Define confidence analysis to calculate the credibility of the model's classification of samples;
[0031] S32. Define a confidence threshold for discriminating the boundary values of known categories, drifted samples, or unknown categories;
[0032] S33. Define confidence drift detection to identify drifted samples and unknown categories through dynamic threshold selection;
[0033] S34. Define a composite detection metric that combines the drift ratio and the original sample loss rate to evaluate the recognition performance of the model.
[0034] Furthermore, in S4, the knowledge distillation-based incremental learning includes:
[0035] S41. Define the total incremental learning loss function, which combines the knowledge distillation loss and the weighted cross-entropy loss to balance the learning of old knowledge and new categories;
[0036] S42. Define a feature attention mechanism to enhance the attention to key features by weighting the input features;
[0037] S43. Use the class drift incremental learning algorithm to solve the problem of degraded classification performance caused by class feature drift through knowledge distillation and the feature attention mechanism;
[0038] S44. Use the new class incremental learning algorithm to improve the classification performance of new classes by independently modeling the features of new classes and fusing them with the features of the old model.
[0039] The beneficial effects of the present invention are as follows:
[0040] The present invention proposes an efficient open-set encrypted traffic recognition method for ultra-large-scale traffic. First, a differential feature drift generation strategy is designed. For DoS attacks, a feature scaling method is adopted, and Gaussian noise perturbation is used for noise-sensitive classes, effectively simulating the feature changes of different types of traffic. Second, an open-set recognition mechanism is established based on confidence analysis, and precise recognition of drift samples and unknown classes is achieved through fine-grained threshold selection. Finally, an incremental learning mechanism combining knowledge distillation and feature fusion is used. Through the design of the feature extraction module and the weighted knowledge distillation strategy, efficient learning of drift classes and new classes is achieved. While maintaining low resource consumption, the adaptability of the model in the open-set environment is significantly improved, providing an effective solution for encrypted traffic classification in lightweight network environments. Description of the Drawings
[0041] Figure 1 It is a flowchart of an efficient open-set encrypted traffic recognition method for ultra-large-scale traffic;
[0042] Figure 2 It is a flowchart of the PCA-based dimensionality reduction algorithm;
[0043] Figure 3 It is a flowchart of encrypted traffic recognition;
[0044] Figure 4 It is a channel diagram of the extended MobileNet architecture. Detailed Embodiments
[0045] In order to make the technical solutions and advantages in the embodiments of the present invention clearer and more understandable, the following further describes the exemplary embodiments of the present invention in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0046] Embodiment 1: Refer to Figures 1-4This embodiment describes an efficient open-set encrypted traffic recognition method for ultra-large-scale traffic, including:
[0047] S1. Input the network traffic dataset and generate the dimensionality-reduced data with the optimal dimension through the PCA dimensionality reduction algorithm;
[0048] S2. Input the dimensionality-reduced data with the optimal dimension into the improved MobileNetV3 model, combine the composite structure of the improved SE attention mechanism, Dense connection, and residual connection to extract features and output the classification result;
[0049] S3. Use the open-set recognition algorithm based on confidence to determine whether the classification result belongs to a known category, a drifting sample, or an unknown category;
[0050] S4. If the classification result is a drifting sample or an unknown category, initiate incremental learning based on knowledge distillation, update the parameters of the improved MobileNetV3 model, and retain the old knowledge.
[0051] Combined with Figure 1 It can be seen that the present invention proposes an efficient and highly adaptable innovative method for network traffic classification and incremental learning scenarios. Based on the improved MobileNetV3 network, through the combination of Dense connection and residual connection structures, the propagation efficiency of deep features is enhanced, and combined with the improved SE attention mechanism, the model's perception ability of key features is improved. While ensuring high accuracy, a lightweight design is achieved, which is suitable for large-scale encrypted traffic detection. For the problems of class drift and new classes in incremental learning, the present invention balances the memory of old class knowledge and the learning of new classes by introducing knowledge distillation and class weight mechanisms. At the same time, a feature attention mechanism and a feature fusion strategy are designed to further improve the model's adaptability and classification performance to complex feature distributions. In addition, through the open-set recognition method based on confidence analysis, the efficient detection of class drift and new class samples is realized, enhancing the model's robustness in unknown scenarios.
[0052] S1 specifically includes the following steps:
[0053] S11. Initialize the network traffic dataset;
[0054] S12. Check whether the initial dimension is greater than 10. If the initial dimension is greater than 10, set the score to 0, set the current dimensionality reduction target dimension i to 10, and then execute the next step. Otherwise, end the process and perform performance verification on the dimensionality-reduced data with the optimal dimension;
[0055] S13. Select i as the dimensionality reduction target for analysis;
[0056] S14. Detect whether the selected i - dimension is sufficient to retain the information of the original data by calculating the variance explanation rate. If the variance explanation rate is less than 0.99 and does not meet the requirements, this dimension is not suitable for dimensionality reduction, and directly execute S18; if the conditions are met, continue with the next step;
[0057] S15. Ensure that the original data is not modified through the deep - copy method, and then apply the dimensionality reduction operation to it to obtain the dimensionality - reduced dataset;
[0058] S16. Divide the dimensionality - reduced dataset into a training set and a test set;
[0059] S17. Use the LMCA network to perform feature extraction and classification on the dimensionality - reduced dataset;
[0060] S18. Calculate and record the accuracy, F1 - score, model size, and inference time of the LMCA network under this dimensionality - reduction configuration, and then return to step S13 to re - select a new i value;
[0061] S19. Among all the tried different i - value configurations, compare the scores score in each case and select the optimal configuration scheme;
[0062] S110. Perform the final performance verification on the data of the optimal configuration scheme.
[0063] Specifically, combined with Figure 1 As can be seen, S1 processes the network traffic dataset through the PCA dimensionality - reduction algorithm to generate the dimensionality - reduced data of the optimal dimension, providing a more efficient and informative input for the subsequent classification model and reducing the computational overhead.
[0064] In S2, the improved MobileNetV3 model includes:
[0065] The initial feature extraction layer uses a 3×3 convolutional kernel to expand the input channels from 1 to 16, combined with the BatchNorm2d and HSwish activation functions to maintain the spatial resolution of the original features;
[0066] The feature learning backbone is composed of 8 MobileBottleneck modules connected in series, gradually adjusting the number of channels to 16, 24, 40, and 80, and selectively using the SE attention mechanism to enhance the feature extraction ability;
[0067] The feature enhancement layer expands the number of channels from 80 to 160 through 1×1 convolution to further fuse and enhance the features;
[0068] The global feature modeling layer uses adaptive average pooling to compress the 3×3 feature map into 1×1 to extract global context information;
[0069] The classification head adopts a three-layer fully connected network structure. Through the dimension reduction design of 160→128→64→11, it outputs the probability distributions of 11 categories.
[0070] Specifically, the process of the improved MobileNetV3 model for identifying encrypted traffic is as Figure 2 shown. This model first proposes an improved SE attention mechanism, enhancing the model's perception ability of key features; secondly, it designs a composite structure of Dense connection and residual connection, effectively solving the problem of feature propagation in deep networks; finally, through the modification and replacement strategy of the network architecture, the lightweight of the model is achieved. The architecture channels of the improved MobileNetV3 model are as Figure 3 shown. This model realizes the efficient abstraction and reasoning of the original data through the selection of the convolution kernel size, the adjustment of the number of channels, and the enhancement of the attention mechanism in the feature extraction layer, the composition of the feature learning backbone and the enhancement of the fusion in the feature representation layer, the feature compression in the global feature modeling layer, and the step-by-step mapping design of the classification head.
[0071] In S3, the confidence-based open-set recognition algorithm includes:
[0072] S31. Define confidence analysis to calculate the credibility of the model's classification of samples;
[0073] S32. Define the confidence threshold, which is used to discriminate the boundary values of known categories, drifting samples, or unknown categories;
[0074] S33. Define confidence drift detection to identify drifting samples and unknown categories through dynamic threshold selection;
[0075] S34. Define a composite detection index, which combines the drift ratio and the original sample loss rate to evaluate the recognition performance of the model.
[0076] Specifically, confidence analysis is a calculation method based on the confidence distribution characteristics in the event coordinate space, used to determine the credibility of the model's classification of samples, denoted as C(x). For a given sample x, its confidence is defined as the maximum value in the Softmax output. Confidence analysis is through:
[0077] C(x) = max(Softmax(f(x)))
[0078] to calculate, where f(x) is the Logits output of the model for sample x.
[0079] The confidence threshold is an important parameter in confidence analysis, used to discriminate the boundary values of samples belonging to known categories, category drift, or new categories, denoted as θ. The percentile threshold is through
[0080] θ p= Percentile(C(T c ),p)
[0081] Calculate, where T c represents the set of correctly classified samples, and p is the percentile value.
[0082] For the input sample x, if its confidence C(x) satisfies:
[0083] C(x) < θ
[0084] Then determine that the sample is a drifting sample. Based on this detection method, the category drift ratio DriftRatec and the overall drift rate DriftRate can be statistically calculated.
[0085] The composite detection index M is a comprehensive evaluation index that combines the drift ratio and the original sample loss rate, and is defined as:
[0086] M = ω 1 ·UnknownRate + ω 2 ·(1 - LossRate)
[0087] where w 1 , w 2 are weight parameters, UnknownRate represents the unknown rate of the detected new category, and LossRate represents the loss rate of the original sample.
[0088] The process of selecting the confidence threshold is as follows: First, initialize an empty set to store candidate thresholds and calculate the correctly classified samples in the training sample set T. Then, calculate the confidence score for each sample and traverse it within a given percentile range with a step size of s. During the traversal, calculate the percentile threshold and add it to the candidate threshold set Θ. Then, evaluate each candidate threshold, calculate the composite detection index M and the drift rate. Finally, select the threshold with the optimal performance as the final confidence threshold θ and return it as the output. The overall process aims to find the most suitable confidence threshold for the current dataset and task requirements through multiple evaluations and comparisons to improve the performance and robustness of the model.
[0089] In S4, the incremental learning based on knowledge distillation includes:
[0090] S41. Define the total loss function of incremental learning, combine the knowledge distillation loss and the weighted cross-entropy loss, and balance the learning of old knowledge and new categories;
[0091] S42. Define the feature attention mechanism, and enhance the attention to key features by weighting the input features;
[0092] S43. Use the category drift incremental learning algorithm to solve the problem of decreased classification performance caused by category feature drift through knowledge distillation and feature attention mechanism;
[0093] S44. Use the new category incremental learning algorithm to improve the classification performance of new categories by independently modeling new category features and fusing them with old model features.
[0094] Specifically, in order to improve the adaptability of the model to drifted categories and new categories, while maintaining the recognition ability for the original categories, a total loss function is constructed:
[0095] L_total = α·LdistilI+(1-α)·LCE
[0096] Where, Ldistill represents the knowledge distillation loss, LCE is the weighted cross-entropy loss with category weights introduced, and α is the weight parameter.
[0097] The feature attention mechanism enhances the attention to key features and suppresses irrelevant features by weighting the input features. It is achieved through:
[0098] x enhanced = σ(W2·ReLU(W1·x))
[0099] where, W 1 and W 2 are the dimension reduction and dimension increase weight matrices; x is the input feature; σ is the Sigmoid function.
[0100] The category drift incremental learning algorithm aims to solve the problem of decreased classification performance caused by category feature drift in the incremental learning scenario. Introducing the feature attention mechanism further enhances the model's attention to key features by automatically adjusting the importance of input features, thereby reducing the interference of feature drift on the model performance.
[0101] The category drift incremental learning process is shown in Table 1:
[0102] Table 1 Category Drift Incremental Learning Process Table
[0103]
[0104] The new category incremental learning algorithm focuses on solving the problems of forgetting old category knowledge and insufficient new category features during the incremental learning process. Therefore, by independently modeling new category features and fusing them with old model features, it fully captures the unique features of new categories while retaining the memory of old categories. This method uses a two-layer convolutional network to extract new category features and adopts a channel splicing and fusion strategy to improve the classification performance of new categories while avoiding interfering with old category knowledge. The new category incremental learning algorithm process is shown in Table 2:
[0105] Table 2 Flow Chart of the New Category Incremental Learning Algorithm
[0106]
[0107] Although the present invention has been described in terms of a limited number of embodiments, those skilled in the art will appreciate that other embodiments can be contemplated within the scope of the invention as thus described. Additionally, it should be noted that the language used in this specification has been principally selected for readability and instructional purposes rather than to limit or define the subject matter of the invention. Accordingly, many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure herein is illustrative and not restrictive, and the scope of the invention is defined by the appended claims.
Claims
1. An efficient open set encrypted traffic identification method for ultra-large-scale traffic, characterized in that: include: S1. Input the network traffic data set and generate the optimal dimension of reduced dimension data through PCA dimensionality reduction algorithm; S2. Input the optimal dimension-reduced data into the improved MobileNetV3 model, combine the improved SE attention mechanism, the composite structure of Dense connection and residual connection, extract features and output the classification results; S3. Use the confidence-based open set recognition algorithm to determine whether the classification result belongs to a known category, a drift sample, or an unknown category; S4. If the classification result is a drift sample or an unknown category, start incremental learning based on knowledge distillation, update the parameters of the improved MobileNetV3 model and retain the old knowledge.
2. According to claim 1, a highly efficient open set encrypted traffic identification method for ultra-large-scale traffic, characterized in that: S1 specifically includes the following steps: S11. Initialize the network traffic data set; S12. Check whether the initial dimension is greater than 10. If the initial dimension is greater than 10, set the score to 0, set the current dimension reduction target dimension i to 10, and then execute the next step. Otherwise, end the process and perform performance verification on the dimension reduction data of the optimal dimension. S13. Select i as the dimension reduction target for analysis; S14. Calculate the variance explanation rate to detect whether the selected i dimension is sufficient to retain the information of the original data. If the variance explanation rate is less than 0.99 and does not meet the requirements, the dimension is not suitable for dimensionality reduction, and S18 is directly executed; if the conditions are met, proceed to the next step; S15. Ensure that the original data is not modified by a deep copy method, and then apply a dimensionality reduction operation to it to obtain a reduced dimensionality data set; S16. Divide the dimension-reduced dataset into a training set and a test set; S17. Use LMCA network to extract features and classify the reduced dimension dataset; S18. Calculate and record the accuracy, F1 score, model size and inference time of the LMCA network under the dimensionality reduction configuration, and then return to step S13 to reselect a new i value; S19. Compare the scores of all the different i value configurations tried and select the best configuration solution; S110. Perform final performance verification on the data of the optimal configuration solution.
3. According to claim 1, a highly efficient open set encrypted traffic identification method for ultra-large-scale traffic, characterized in that: In S2, the improved MobileNetV3 model includes: In the initial feature extraction layer, a 3×3 convolution kernel is used to expand the input channels from 1 to 16, and the BatchNorm2d and HSwish activation functions are used to maintain the spatial resolution of the original features. The feature learning backbone consists of 8 MobileBottleneck modules connected in series, with the number of channels gradually adjusted to 16, 24, 40, and 80, and the SE attention mechanism is selectively used to enhance the feature extraction capability; Feature enhancement layer, which expands the number of channels from 80 to 160 through 1×1 convolution to further fuse and enhance features; The global feature modeling layer uses adaptive average pooling to compress the 3×3 feature map into 1×1 to extract global context information; The classification head adopts a three-layer fully connected network structure and outputs the probability distribution of 11 categories through a dimensionality reduction design from 160→128→64→11.
4. According to claim 1, a highly efficient open set encrypted traffic identification method for ultra-large-scale traffic, characterized in that: In S3, the confidence-based open set recognition algorithms include: S31. Define confidence analysis to calculate the degree of confidence of the model in classifying samples; S32. Define a confidence threshold for distinguishing the boundary value of a known category, a drift sample, or an unknown category; S33. Define confidence drift detection to identify drift samples and unknown categories through dynamic threshold selection; S34. Define a composite detection index that combines the drift ratio and the original sample loss rate to evaluate the recognition performance of the model.
5. According to claim 1, a highly efficient open set encrypted traffic identification method for ultra-large-scale traffic, characterized in that: In S4, incremental learning based on knowledge distillation includes: S41. Define the total loss function for incremental learning, combining knowledge distillation loss and weighted cross entropy loss to balance old knowledge and new category learning; S42. Define the feature attention mechanism to enhance the focus on key features by weighting the input features; S43. Use the category drift incremental learning algorithm to solve the problem of classification performance degradation caused by category feature drift through knowledge distillation and feature attention mechanism; S44. Use the new category incremental learning algorithm to improve the new category classification performance by independently modeling new category features and fusing them with old model features.
Citation Information
Patent Citations
Lightweight open-set landmark identification method for mobile terminal
CN112818893A
Lightweight Tor flow classification method and system based on quaternary feature fusion graph
CN114187485A
Lightweight Internet of Things malicious traffic identification method based on knowledge distillation space-time neural network
CN116260642A
Feature-enhanced lightweight network FGNet facial expression recognition method
CN116311414A
Strawberry fruit identification method based on improved YOLOv5s
CN118397427A
Cited By
Unknown encrypted traffic identification method and system based on small sample incremental learning, and storage medium
CN120602237A
An unknown encrypted traffic identification method and system based on small sample incremental learning and a storage medium
CN120602237B