Encryption proxy traffic classification method based on non-proxy traffic
By performing feature alignment of non-proxy traffic and training of Seq2Seq model, a simulated encrypted proxy traffic feature sequence is generated, which solves the problems of degradation of classification accuracy and data dependence in the existing methods, and realizes efficient encrypted proxy traffic classification.
Patent Information
- Application Number
- CN202510754139.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-01
AI Technical Summary
The existing machine learning-based encrypted proxy traffic classification method has significantly reduced classification accuracy when the training set and the test set feature sequence distribution is inconsistent, and it is costly to build multiple proxy protocol models, making it difficult to adapt to rapidly changing proxy protocols.
By collecting non-proxy traffic and encrypted proxy traffic data pairs, extracting the load length sequence as features, and then aligning the feature sequence of simulated encrypted proxy traffic is used to generate a simulated encrypted proxy traffic feature sequence, training the encrypted traffic classification model, and reducing dependence on real encrypted proxy traffic.
It significantly improves the adaptability and robustness of encrypted proxy traffic classification, reduces data acquisition costs, and realizes efficient classification of different proxy protocols.
Smart Images

Figure CN120415876A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of network communication and network security, and particularly relates to a method for classifying encrypted proxy traffic based on non-proxy traffic. Background Art
[0002] With the popularization of the Internet, encrypted proxy traffic is widely used to bypass network restrictions and protect privacy. Proxy software encapsulates and encrypts non-proxy traffic according to the proxy protocol specification, and uses the proxy server as a springboard to enable users to access restricted content. However, there are significant differences in the distribution of feature sequences between non-proxy traffic and encrypted proxy traffic. Machine learning-based encrypted proxy traffic classification methods usually rely on the feature sequence distributions of the training set and the test set to tend to be consistent. When the feature sequence distribution changes, the classification accuracy of the model will drop significantly. To improve the classification effect, existing methods often need to separately construct and train classification models for the encrypted traffic of each proxy protocol (such as Shadowsocks / VMess / Trojan / VLESS), which not only greatly increases the cost of data collection and annotation, but also is difficult to adapt to the challenges brought by the rapid evolution of proxy protocols.
[0003] In the actual scenario of encrypted proxy traffic identification, the cost of collecting and annotating the encrypted proxy traffic dataset is relatively high. If non-proxy traffic is used as the training set (or for generating simulated training data), the collection cost can be greatly reduced. This scenario where the training and test data distributions are inconsistent is called an "asymmetric identification scenario". In an asymmetric scenario, since the training set does not directly cover the traffic characteristics of the proxy protocol, the classification accuracy usually drops significantly, making it difficult to meet the actual deployment requirements.
[0004] To balance transmission efficiency, encrypted proxy protocols generally do not perform compression and padding when encapsulating and encrypting non-proxy traffic into encrypted proxy traffic. Therefore, there is a correlation between the distribution of feature sequences of non-proxy traffic and encrypted proxy traffic. To solve the problems existing in the existing methods, the present invention classifies encrypted proxy traffic based on non-proxy traffic, which can reduce the dependence on encrypted proxy traffic training data, achieve efficient classification, and improve the adaptability and robustness of the classification model. Summary of the Invention
[0005] The object of the present invention is to propose a classification method for encrypted proxy traffic based on non-proxy traffic. The load length sequence is used as the feature sequence of non-proxy traffic and encrypted proxy traffic. After feature alignment of the feature sequences of non-proxy traffic and encrypted proxy traffic, the feature-aligned feature sequence is used to train a Seq2Seq (sequence-to-sequence) model; a simulated encrypted proxy traffic feature sequence is generated by the trained Seq2Seq model; the existing encrypted traffic classification model is trained using the simulated encrypted proxy traffic feature sequence, thereby realizing the classification of encrypted traffic.
[0006] The present invention specifically includes the following steps: Step 1: Collect data pairs of non-proxy traffic and its corresponding encrypted proxy traffic, and construct a training set and a test set.
[0007] Step 2: Perform feature extraction and feature alignment on the data pairs collected in Step 1. The load length sequence is used as the feature sequence of non-proxy traffic and encrypted proxy traffic. According to different proxy protocols, the load length sequences of non-proxy traffic and encrypted proxy traffic are extracted from the original packet sequences of non-proxy traffic and its corresponding encrypted proxy traffic.
[0008] Furthermore, to retain the direction information of the traffic, the load length is encoded: when the transmission direction of the packet is from the client to the server C2S, its load length is recorded as a positive value; when the transmission direction of the packet is from the server to the client S2C, its load length is recorded as a negative value.
[0009] Feature alignment is performed on the load length sequence according to different proxy protocols, so that the non-proxy traffic feature sequence and the encrypted proxy traffic feature sequence tend to be consistent in length distribution, transmission direction, and statistical characteristics, reducing the differences introduced by the network transmission layer between the two.
[0010] Furthermore, the feature alignment includes load length extraction, recombination, and segmentation processing, normalizing the original features of non-proxy traffic and encrypted proxy traffic, reducing the differences in length distribution, transmission direction, and statistical characteristics between the non-proxy traffic feature sequence and the encrypted proxy traffic feature sequence, and effectively avoiding feature deviation problems caused by protocol handshakes, MTU fragmentation, and network fluctuations.
[0011] Step 3: Use the non-proxy traffic feature sequence and the encrypted proxy traffic feature sequence after feature alignment in Step 2 to train the Seq2Seq model. After any non-proxy traffic feature sequence is feature-aligned as described in Step 2, it is input into the trained Seq2Seq model, and a simulated encrypted proxy traffic feature sequence with corresponding encrypted proxy traffic statistical characteristics and semantic features is output.
[0012] Furthermore, the Seq2Seq model includes an encoder and a decoder. The input is the feature sequence of non-proxy traffic after feature alignment, and the output is the generated simulated encrypted proxy traffic feature sequence.
[0013] The encoder transforms the input feature sequence into a hidden vector by extracting its context information. The hidden vector retains the temporal structure and global dependencies of the original non-proxy traffic feature sequence and provides a unified feature representation for the decoder.
[0014] The decoder generates the target simulated feature sequence based on the hidden vector. The generation of each target simulated feature sequence depends on the current hidden vector and the sequence already generated by the decoder.
[0015] Using the data pairs collected in Step 1, the model learns the mapping relationship between non-proxy traffic and encrypted proxy traffic in terms of time series and feature distribution. After the model training is completed, any new non-proxy traffic feature sequence (after feature alignment processing) can be accepted as input to generate a simulated feature sequence with the characteristics of the corresponding encrypted proxy traffic.
[0016] Step 4: Combine the simulated encrypted proxy traffic feature sequence obtained in Step 3 with its corresponding category to form a training data set, and input it into any encrypted traffic classification model to train the encrypted traffic classification model.
[0017] Furthermore, the encrypted traffic classification model is a deep learning-based classifier (such as a Transformer model) or a traditional machine learning model (such as a random forest, etc.).
[0018] Step 5: Input any real encrypted proxy traffic feature sequence after feature alignment into the encrypted traffic classification model trained in Step 4, and predict the category of the real encrypted proxy traffic by analyzing the feature sequence.
[0019] The present invention uses the payload length sequence as the feature sequence of non-proxy traffic and encrypted proxy traffic, performs feature alignment on the feature sequences of non-proxy traffic and encrypted proxy traffic according to different proxy protocols, effectively bridges the differences in the distribution of feature sequences between non-proxy traffic and encrypted proxy traffic, and significantly reduces the differences in the distribution of feature sequences between non-proxy traffic and encrypted proxy traffic. It can also adapt to the characteristic distributions of different proxy protocols, and successfully solves the problem of poor adaptability of the classification model caused by the diversification of proxy protocols.
[0020] Train the Seq2Seq model using the feature sequence after feature alignment. Generate a simulated encrypted proxy traffic feature sequence using non-proxy traffic. The generated simulated feature sequence is highly consistent with the real encrypted proxy traffic under the target proxy protocol in terms of payload length distribution, direction information, and transmission semantics. Provide high-quality simulated data input for subsequent classification tasks, and achieve model training without the participation of a large amount of real encrypted proxy traffic.
[0021] Existing encrypted traffic classification models can directly use the generated simulated feature sequence for training to achieve effective classification of real encrypted proxy traffic. By using the simulated features as the training data input, the classification model can learn the statistical features and sequence distribution characteristics of encrypted proxy traffic under different proxy protocols, so as to accurately identify the traffic types generated by different proxy protocols. Significantly improve the adaptability and robustness of encrypted proxy traffic classification, and at the same time avoid relying heavily on real traffic data under multiple proxy protocols.
[0022] Different from traditional methods that require a large amount of real encrypted proxy traffic data for classification model training, the present invention does not need to collect a large amount of real encrypted proxy traffic data, eliminates the dependence on a large amount of real encrypted proxy traffic data in the model training process, and significantly reduces the data collection cost. Brief Description of the Drawings
[0023] Figure 1 is a flowchart of the method of the present invention; Figure 2 is a schematic diagram of the data collection platform in the embodiment. Detailed Embodiments
[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.
[0025] As Figure 1 shown, an encrypted proxy traffic classification method based on non-proxy traffic specifically includes the following steps: Step S1, construct a paired data set: Build a data collection platform to synchronously capture non-proxy traffic and its corresponding encrypted proxy traffic generated after being encapsulated and encrypted by a specific proxy protocol. These paired non-proxy traffic and encrypted proxy traffic form data pairs and are stored in the paired data set for subsequent training of the Seq2Seq feature conversion model.
[0026] As Figure 2 shown, the data collection platform includes a client device, a proxy client, and a proxy server. To ensure synchronous capture of traffic data and feature consistency, the client device, the proxy client, and the proxy server are deployed in the same local area network.
[0027] The client device simulates the network access behavior of real users (e.g., initiating access requests to a set of target websites via a preset script using the HTTP or SOCKS proxy protocol), generating non-proxy traffic. This non-proxy traffic is sent to the proxy client through the HTTP or SOCKS proxy protocol interface. To ensure data diversity and coverage, multiple different types of target websites are selected for multiple rounds of access.
[0028] After receiving the non-proxy traffic generated by the client device, the proxy client encapsulates and encrypts the non-proxy traffic according to the set proxy protocol, generating encrypted proxy traffic and forwarding it to the proxy server. The process of encapsulating and encrypting proxy traffic includes extracting the data payload of the non-proxy traffic, encrypting it, and adding headers and other information according to the protocol specifications, which changes the performance of the non-proxy traffic in terms of packet length, payload distribution, etc., and may introduce protocol-specific header structures or padding strategies, resulting in a detectable mapping relationship in the distribution of the feature sequence, thus achieving pairing.
[0029] The proxy client is also equipped with a traffic capturer to record the non-proxy traffic sent from the client device and the encrypted proxy traffic after being encapsulated and encrypted by the proxy client.
[0030] The proxy server is responsible for decrypting and de-encapsulating the encrypted proxy traffic sent by the proxy client, then forwarding the original request to the target server to complete the access request, and sending the data returned by the target server back to the client device along the original path.
[0031] Step S2, Feature Extraction and Alignment: Feature extraction, feature recombination, and feature segmentation are performed on the original packet sequences of the non-proxy traffic and encrypted proxy traffic collected in step S1 to align the feature sequences of the non-proxy traffic and encrypted proxy traffic. This step aims to eliminate the differences introduced by network transmission mechanisms (such as TCP fragmentation, MTU limitations) and protocol handshakes, etc., so that the subsequent model can focus more on learning the transformation brought by the proxy protocol itself.
[0032] The feature extraction process is as follows: The payload length sequences are respectively extracted from the original packet sequences of the non-proxy traffic and encrypted proxy traffic, and the payload length sequences are used as the core feature sequences of the traffic, so that the payload features are standardized into a unified payload length sequence.
[0033] The original packet sequence is represented as , where represents the th packet, usually an Ethernet frame. The payload length of packet is .
[0034] To retain the direction information of the traffic, the payload length is encoded: when the transmission direction of the data packet is from the client to the server C2S, its payload length is recorded as a positive value; when the transmission direction of the data packet is from the server to the client S2C, its payload length is recorded as a negative value, that is ; After encoding, the payload length feature sequence of non-proxy traffic is represented as: where is the number of data packets of non-proxy traffic.
[0035] Similarly, for encrypted proxy traffic, the encrypted payload length sequence is extracted according to its bearer protocol type: where is the number of data packets of encrypted proxy traffic.
[0036] For proxy protocols based on TCP bearer (such as some Shadowsocks configurations), the payload length of the TCP stream is extracted as the feature sequence for feature recombination and feature segmentation. For proxy protocols based on TLS bearer (such as the common forms of VMess, Trojan, and VLESS), the length of the application data (Application Data) recorded in the TLS stream is extracted as the feature sequence for feature segmentation.
[0037] In this embodiment, taking the feature recombination and feature segmentation of the TCP stream as an example, the feature recombination and feature segmentation are specifically described. The feature recombination process for the TCP stream after feature extraction is as follows: The TCP protocol will fragment the payload data larger than the maximum transmission unit (MTU) into multiple data packets for transmission. This fragmentation process will result in incomplete payload information within a single data packet. To restore the length information of the complete application layer data unit (PDU) or its fragments, feature recombination is performed for the fragmentation mechanism of the TCP protocol.
[0038] By identifying the Push bit flag in the TCP header, the payload lengths of consecutive TCP segments in the same direction and belonging to the same application message can be accumulated until a data packet with the PUSH flag bit set is encountered, forming a recombined payload length.
[0039] For example, a payload sequence in the C2S direction (length, PUSH flag) is {(100,0), (200,1), (150,0), (250,1)}. By detecting the Push bit flag, the recombined result is {(100 + 200),(150 + 250)} = {300,400}.
[0040] The specific recombination method is as follows: For data packets with the same transmission direction, accumulate their payload lengths , until a data packet with the Push bit marked as 1 is encountered, output the accumulated payload length and insert it into the new feature sequence. The aggregated payload length is defined as: ; Through feature recombination, the original payload fragment sequence is recombined into a feature sequence containing more complete application layer information unit lengths , effectively eliminating the deviation introduced by TCP fragmented transmission.
[0041] The algorithm expression of feature recombination is as follows: Input: Packet sequence Output: Recombined sequence features 1. Initialize , , 2. For each data packet in : 2.1 If the payload length : 2.1.1 If the direction is C2S: If the Push bit is marked as 1: .append( ) 2.1.2 If the direction is S2C: If the Push bit is marked as 1: .append( ) 3. Return The process for TCP flowing through the feature segments after feature recombination is as follows: During the feature recombination process, if the TCP Push flag is occasionally lost or abnormal, it may lead to over-aggregation, causing some of the recombined payload lengths to be abnormally large and deviate from the normal distribution range. To address this issue, a dynamic segmentation mechanism based on statistical analysis is adopted.
[0042] By calculating the segmentation threshold split the overly long recombinant payload fragments. The segmentation threshold can be calculated according to the distribution of the lengths of the recombinant payloads in the dataset . For example, select the mode of those values that exceed (MTU - TCP header length) as , ensuring is greater than a minimum reasonable threshold (such as MTU - TCP.head). The specific formula is: where MTU is the maximum transmission unit in the network configuration.
[0043] For recombinant payload fragments with absolute values exceeding the segmentation threshold , split them at a fixed length until the absolute value of the length of the remaining fragment is less than or equal to . The segmentation rules are as follows: ; This step is applied to , and through the segmentation operation, the overly long payload fragments are reasonably split to ensure that the length distribution of the generated feature sequence is more stable and more in line with the characteristics of the target proxy protocol..
[0044] The algorithm for feature segmentation is expressed as follows: Input: Aggregated sequence features Output: Segmented sequence features 1. Initialize 2. Calculate the segmentation threshold 3. for each in : 3.1 If : Split into multiple sub - fragments Add the sub - fragments to 3.2 Otherwise: Directly add to 4. Return After the above-mentioned feature extraction, recombination, and segmentation steps, the non-proxy traffic feature sequence and the encrypted proxy traffic feature sequence are more standardized and consistent in terms of length distribution, directionality, and statistical characteristics, significantly reducing the differences introduced at the network transmission level between the two and achieving feature alignment. This lays a foundation for the Seq2Seq model in the subsequent step S3 to learn the essential mapping relationship introduced by the proxy protocol encapsulation and encryption between the two.
[0045] Step S3: Feature transformation based on the Seq2Seq model: The Seq2Seq model includes an encoder and a decoder, and the Seq2Seq model can adopt advanced sequence model architectures such as Transformer.
[0046] Input stage: The Seq2Seq model receives the non-proxy traffic feature sequence after feature alignment in step S2, denoted as , where is a value containing the payload length and C2S / S2C direction information, and represents the length of the input feature sequence.
[0047] Encoder: Convert the input sequence into a set of high-dimensional hidden vector representations . These hidden vectors encode the global context information and temporal dependencies of the input sequence. The encoder can be implemented using a recurrent neural network (RNN), a long short-term memory network (LSTM), or a Transformer layer based on the attention mechanism.
[0048] Decoder: Based on the hidden vectors generated by the encoder (and possibly the attention mechanism), gradually generate the target simulated encrypted proxy traffic feature sequence . At each time step of decoding, generate the th element of the target simulated encrypted traffic feature sequence depending on the current hidden vector and the sequence of outputs that the decoder has already generated. The generated target sequence should be highly similar to the real encrypted proxy traffic (corresponding to a specific proxy protocol) in terms of payload length distribution, direction information, and statistical characteristics.
[0049] Model training: For the training of the Seq2Seq model, use the non-proxy traffic feature sequence collected in step S1 and feature-aligned in step S2 as the input, and its corresponding real encrypted proxy traffic feature sequence As the target output. The training process optimizes the Masked Softmax Cross-Entropy Loss function , which is used to process variable-length sequences and ignore padding bits, so that the simulated sequence Y generated by the model is as consistent as possible with the true target sequence Y′ in various features.
[0050] , where is the mask weight, which is used to ignore the padded part; is the non-proxy traffic feature sequence collected in step S1 and feature-aligned in step S2.
[0051] Through multiple rounds of iterative optimization, the Seq2Seq model can learn and capture the complex mapping relationship from the non-proxy traffic feature sequence aligned in step S2 to the encrypted proxy traffic feature sequence aligned in step S2. Learn non-proxy traffic and specific encrypted proxy traffic.
[0052] The trained Seq2Seq model can receive new non-proxy traffic feature sequences aligned in step S2 and generate simulated encrypted proxy traffic feature sequences with corresponding encrypted proxy traffic statistical characteristics and semantic features, providing data for the classifier training in step S4.
[0053] Step S4, Training of the encrypted proxy traffic classification model: Use the simulated encrypted proxy traffic feature sequences generated in step S3 and their corresponding labels to train the encrypted proxy traffic classification model.
[0054] Training data preparation: First, align the features of a batch of non-proxy traffic (for example, traffic accessing different known websites or applications, with clear category labels L) through step S2, and then input them into the Seq2Seq feature conversion model trained in step S3 to generate a corresponding set of simulated encrypted proxy traffic feature sequences . These simulated feature sequences and their corresponding labels L together constitute the training data set of the encrypted proxy traffic classification model.
[0055] Classification model training: Select a suitable encrypted traffic classification model, such as a deep learning-based model, a traditional machine learning model, or other existing encrypted traffic classification methods. Use the prepared simulated encrypted proxy traffic feature sequences Y and their labels L to train the classification model. During the training process, the classification model learns to identify patterns from the simulated feature sequences to distinguish different categories.
[0056] Step S5, Classify and identify real encrypted proxy traffic: The trained classification model receives the real, newly captured encrypted proxy traffic feature sequence (which also needs to go through the feature alignment process in step S2 first to match the data format during training) as input, and outputs the predicted class label by analyzing the features of the input sequence.
[0057] In the above way, the present invention realizes training a model that can effectively classify real encrypted proxy traffic only relying on non-proxy traffic data (used to generate simulated proxy traffic), significantly reducing the data acquisition cost and improving the adaptability to emerging proxy protocols.
[0058] To demonstrate the effectiveness of the method described in the present invention, the following verification experiments are carried out: First, proxy traffic and non-proxy traffic sample pairs are collected for 60 websites. A total of 6,374,760 non-proxy traffic and proxy traffic sample pairs constitute the experimental data set. The specific traffic data set information is shown in Table 1.
[0059] Table 1: Statistical information of the traffic data set used in experimental verification To evaluate its classification performance, appropriate classification evaluation metrics need to be defined. For the specific traffic data set being analyzed, the Macro-F1 value metric is defined to evaluate the classification performance of the classifier: ; where N is the total number of classes. represents , where represents the precision rate of class i , represents the recall rate of class i ; ; ; The higher the Macro-F1 value, the better the comprehensive performance of the classification model in each class.
[0060] In the actual training and testing process, there are the following several ways to select the training set and the testing set: (1) using non-proxy traffic as the training set and using proxy traffic as the testing set; (2) using proxy traffic as the training set and using proxy traffic as the testing set; (3) using non-proxy traffic as the training set and using non-proxy traffic as the testing set.
[0061] In this experiment, non-proxy traffic was selected as the training set and proxy traffic was selected as the test set, and the Transformer model was selected as the Seq2Seq model. The existing FS-Net, ETC-PS, random forest, XGBoost, and Transformer classification models were tested using the method described in the present invention and without using the method described in the present invention. The test results are shown in Table 2. It can be seen from Table 2 that using the method described in the present invention improves the Macro-F1 value of the existing model.
[0062] Table 2: Experimental results of existing models on the test dataset before and after using the present invention It should be noted that only one preferred embodiment of the present invention is disclosed above. It should be stated that the scope of the rights of the present invention cannot be limited by this. Therefore, all equivalent changes made according to the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A method for classifying encrypted proxy traffic based on non-proxy traffic, characterized in that: Specifically, it includes the following steps: Step 1: Collect pairs of non-proxy traffic and its corresponding encrypted proxy traffic data, and construct a training set and a test set; Step 2: Extract features and align features for the data pairs collected in Step 1. Take the payload length sequence as the feature sequence of non-proxy traffic and encrypted proxy traffic, and extract the payload length sequences of non-proxy traffic and encrypted proxy traffic from the original packet sequences of non-proxy traffic and its corresponding encrypted proxy traffic according to different proxy protocols; Align the payload length sequences according to different proxy protocols, so that the non-proxy traffic feature sequence and the encrypted proxy traffic feature sequence tend to be consistent in length distribution, transmission direction and statistical characteristics, reducing the differences introduced by the network transmission layer between the two; Step 3: Use the non-proxy traffic feature sequence and the encrypted proxy traffic feature sequence after feature alignment in Step 2 to train the Seq2Seq model; After aligning the features of any non-proxy traffic feature sequence according to Step 2, input it into the trained Seq2Seq model to output a simulated encrypted proxy traffic feature sequence with corresponding encrypted proxy traffic statistical characteristics and semantic features; Step 4: Combine the simulated encrypted proxy traffic feature sequence obtained in Step 3 with its corresponding category to form a training data set, and input it into any encrypted traffic classification model to train the encrypted traffic classification model; Step 5: Input any truly encrypted proxy traffic feature sequence after feature alignment into the encrypted traffic classification model trained in Step 4, and predict the category of the truly encrypted proxy traffic by analyzing the feature sequence.
2. The method for classifying encrypted proxy traffic based on non-proxy traffic according to claim 1, characterized in that: In the second step, to retain the direction information of the traffic, the payload length is encoded: when the transmission direction of the data packet is from the client to the server C2S, its payload length is recorded as a positive value; when the transmission direction of the data packet is from the server to the client S2C, its payload length is recorded as a negative value.
3. The method for classifying encrypted proxy traffic based on non-proxy traffic according to claim 1, wherein: The feature alignment described in Step 2 includes payload length extraction, recombination and segmentation processing.
4. The method for classifying encrypted proxy traffic based on non-proxy traffic according to claim 1, characterized in that: The Seq2Seq model described in Step 3 includes an encoder and a decoder. The input is the feature sequence of non-proxy traffic after feature alignment, and the output is the generated simulated encrypted proxy traffic feature sequence; The encoder converts the input feature sequence into a hidden vector by extracting the context information of which preserves the temporal structure and global dependencies of the original non-proxy traffic feature sequence and provides a unified feature representation for the decoder; The decoder generates the target simulated feature sequence based on the hidden vector. The generation of each target simulated feature sequence depends on the current hidden vector and the sequence already generated by the decoder.
5. The method for classifying encrypted proxy traffic based on non-proxy traffic according to claim 1, wherein: The encrypted traffic classification model described in Step 4 is a classifier based on deep learning or a traditional machine learning model.