Proxy encrypted traffic identification method based on context multi-stream association
Through the proxy encryption traffic recognition method based on context multi-stream association, multivariate feature analysis and random forest algorithms are used to generate fingerprint vectors, which solves the problem of identifying encrypted proxy traffic in a real network environment, and achieves high-precision and robust traffic recognition.
Patent Information
- Application Number
- CN202510615785.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-05
AI Technical Summary
The existing technology cannot effectively identify encrypted proxy traffic in a real network environment, resulting in limited practicality of network behavior recognition technology, especially when training and testing are inconsistent, traditional methods cannot work properly.
A proxy encrypted traffic recognition method based on context multi-stream association is adopted, and fingerprint vectors are generated through multivariate feature analysis, dual-label calibration, and random forest algorithms, website matching indicators and serialized fingerprint importance scores are calculated, and a random forest classification model is constructed for traffic recognition.
It improves fingerprint recognition accuracy and model robustness in complex network environments, solves the training-test asymmetry problem, and enhances the generalization ability and recognition accuracy in actual testing.
Smart Images

Figure CN120433997A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic identification, and in particular to a proxy encrypted traffic identification method based on context multi-flow association. Background Art
[0002] The present invention relates to the high-intensity and covert development trend of network confrontation. The detection of network encrypted traffic plays an important role in promoting the in-depth development of traffic detection capabilities. In encrypted communication scenarios, encrypted proxies are widely used to transmit information between two communicating parties. Typical encrypted proxies include the Tor anonymous network for multi-hop communication and Shadowsocks(R), V2Ray, etc. for single-hop communication. Although encrypted proxies provide anonymity and confidentiality for users' website access behavior, they bring technical challenges to the supervision and profiling of sensitive website access behavior. Detection of proxy stream encryption volume has become an important demand in the field of network security.
[0003] Even before HTTPS encrypted transmission, network administrators could use DPI technology to analyze plaintext traffic content and obtain user network behavior characteristics. With the adoption of HTTPS encrypted transmission, relying on SSL or TLS encryption technology to ensure connection security, DPI technology is no longer applicable. However, network administrators can still infer the websites users visit through basic characteristics such as DNS requests and network layer IP addresses. However, after the widespread use of encrypted proxies such as Shadowsocks(R) and V2Ray, basic characteristics such as DNS requests and network layer IP addresses have also been encrypted and hidden, making these methods unable to identify the websites users visit. When network proxy encrypted traffic becomes the sole data source, website access behavior can only be detected by identifying the website fingerprint characteristics of encrypted proxy traffic in terms of statistics, timing, and other aspects.
[0004] The traditional training model is as follows: in a controllable network environment (i.e., without any interference from background traffic), a script is used to repeatedly simulate browser behavior to access a specific website, and the start and end time of the access are recorded. All data packets collected during this time interval are considered to be the traffic generated by accessing the website. After collecting a fixed number of samples from multiple websites of interest, a large number of traffic samples from other websites are collected, which serve as the raw data for training the classifier. In such a model, traffic statistical features and classification algorithms designed to distinguish between access behaviors to different websites are generally considered to be the key to network behavior recognition technology. However, when the trained classifier is deployed in a real network environment, the trained classifier cannot work properly. This is because the traffic collected over a period of time in a real network environment is not a pure and complete traffic sample generated by a user visiting a single website. In fact, it is usually a mixed traffic generated by a user visiting multiple websites. The present invention defines this inconsistency between training and test samples as a training-test asymmetry problem, which fundamentally limits the practicality of existing network behavior recognition technology for encrypted traffic. Summary of the Invention
[0005] In response to the problems existing in the existing technology, a proxy encrypted traffic identification method based on contextual multi-flow association is provided, which can implement contextual multi-flow association traffic identification based on machine learning for encrypted proxies without directly cracking the encrypted traffic.
[0006] The technical solution adopted by the present invention is as follows: a proxy encrypted traffic identification method based on contextual multi-flow association, comprising:
[0007] Multivariate feature analysis: Collect mixed traffic samples generated by each round of website visits and perform data flow extraction; perform dual-label calibration and statistical feature extraction on each data flow;
[0008] Characterizing spatial relationships of data streams based on contextual semantic recognition: Based on the results of multivariate feature analysis, a random forest algorithm is used to generate fingerprint vectors for each data stream and calculate the target website matching index;
[0009] Characterizing the temporal relationship of data streams based on contextual semantic recognition: This involves calculating the website indicator coefficient of data streams extracted through multivariate feature analysis. Traffic samples accessing the same website are split into data stream instance sequences, and the LCS is calculated to generate a serialized fingerprint. The importance score of the serialized fingerprint is calculated using the website indicator coefficient of the data stream.
[0010] Deep identification of website fingerprints associated with multiple streams: Build a random forest classification model and perform iterative training of the random forest classification model based on the fingerprint vector of the data stream, the target website matching index, the serialized fingerprint of the data stream, and the importance score. Traffic identification can be completed through the trained classification model.
[0011] As a preferred solution, the double-label marking includes: the first label is the URL of the visited website, and the second label is the identifier of each data flow in all data flows of the website to which it belongs.
[0012] As a preferred solution, the statistical feature extraction includes: overall statistical features, data packet sequence features, data packet time features, data packet size features, head packet and tail packet features; wherein, the overall statistical features are a comprehensive description of each flow; the data packet sequence features involve the interaction sequence between requests and responses and the density distribution of data packets in the time dimension; the data packet time features involve the time characteristics of the arrival of data packets in the data flow; the data packet size features involve the size distribution of each data packet in the data flow.
[0013] As a preferred solution, the fingerprint vector generation process includes: for each data stream sample, using the output results of all decision trees in the random forest algorithm to construct a t-dimensional vector as the fingerprint vector of the data stream sample, where t is the number of decision trees in the random forest; the decision tree is trained based on different training data stream samples and statistical features.
[0014] As a preferred solution, the website matching index is equal to the number of target websites output by the decision tree divided by the total number of decision trees in the random forest.
[0015] As a preferred solution, the calculation process of the website indication coefficient of the data stream includes:
[0016] For the jth data flow flow(i,j) of sample i under website I, its website indication coefficient is defined as W2I ij ;
[0017] The calculation method is as follows: the data flow sample labeled bi-label(i,j) is called an instance of flow(i,j). For the data flow sample labeled bi-label(i,j), calculate the proportion of samples labeled bi-label(i,j) in its K nearest neighbor samples. Then, the average of the proportion values calculated for all data flow samples labeled bi-label(i,j) is taken as the W2I of flow(i,j). ij .
[0018] As a preferred solution, the serialized fingerprint generation process includes:
[0019] The traffic samples generated by visiting the same website are split into data stream instance sequences, and a serialized fingerprint is obtained by calculating the LCS between each sequence.
[0020] As a preferred solution, the importance score calculation process of the serialized fingerprint includes:
[0021]
[0022] Among them, L is the currently scored serialized fingerprint, #occur(L) represents the number of times the serialized fingerprint appears in all data stream sequences, W2I ij Indicates the coefficient of the website.
[0023] As a preferred solution, the iterative training of the random forest classification model includes:
[0024] The fingerprint vector of the data stream and the target website matching index, the serialized fingerprint of the data stream and the importance score are input into the random forest classification model for iterative training. During the iterative training process, the training results are adjusted according to the target website matching index and the importance score.
[0025] As a preferred solution, the traffic identification can be completed by using a trained classification model, including: inputting the actual mixed traffic into a trained random forest classification model, and the random forest classification model can determine whether the target website has been visited.
[0026] Compared with the existing technology, the beneficial effects of adopting the above technical solution are:
[0027] (1) Multi-flow fusion association recognition is proposed to improve fingerprint recognition accuracy: Traditional methods usually perform website fingerprint recognition based on a single network flow, which is easily affected by network traffic congestion and insufficient feature similarity. This method establishes a fine-grained mapping relationship between website resources and network flows through multi-flow fusion association analysis. It can mine the similarity of flow features caused by shared resources between different websites, thereby improving fingerprint recognition accuracy in complex network environments.
[0028] (2) Solve the training-test asymmetry problem and enhance model robustness: Traditional fingerprint recognition methods often suffer from performance degradation during training and testing due to inconsistent traffic feature distribution, especially in complex network environments. This method proposes a new multi-stream fusion method through spatiotemporal fine-grained analysis, which can effectively solve the training-test asymmetry problem caused by traffic congestion and significantly improve the generalization ability and robustness of the model in actual testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a flow chart of the proxy encrypted traffic identification method based on contextual multi-flow association proposed by the present invention.
[0030] Figure 2 Schematic diagram of data stream statistical feature extraction in one embodiment of the present invention. DETAILED DESCRIPTION
[0031] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar modules or modules with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application. On the contrary, the embodiments of the present application include all changes, modifications, and equivalents that fall within the spirit and scope of the appended claims.
[0032] In order to solve the problems existing in the prior art, the embodiment of the present application proposes a proxy encrypted traffic identification method based on context multi-flow association, please refer to Figure 1 ,This method mainly includes three aspects : multi-dimensional feature analysis, ,contextual semantic recognition and deep website fingerprint recognition based on ,multi-stream association. These three aspects are explained one by one ,below.
[0033] (1) Multivariate feature analysis
[0034] In this embodiment, the multivariate feature analysis is mainly performed by collecting mixed traffic samples generated by each round of website visits, extracting data streams, and then performing dual-label calibration and statistical feature extraction on the data streams.
[0035] Data stream extraction is a crucial first step in implementing this solution. In this example, the packets in the pcap file are divided into multiple data streams. This division of data streams relies not only on network protocols and ports but also on the access log files of the encryption proxy. The access log records the source of each request, including information such as the request time, the target domain name, the HTTP method, and the response status.
[0036] Then, each data stream is double-labeled, and the first label in the double label of the data stream is the URL of the website to which the data stream belongs. This information is obtained directly from the data source file to ensure that each stream can clearly identify its source website. The second label is the identifier of each data stream among all data streams of the website to which it belongs. When a user visits a website, multiple TCP connections are usually established to download and render the various resources required for HTML pages. In this case, there may be multiple data streams requesting different resources under the same domain name. For example, a web page may need to load multiple resources such as text, images, style sheets, and scripts. Therefore, in the data source file, there will be multiple data stream sequences under the same domain name, and these sequences correspond to requests for different resources.
[0037] To address the flow alignment problem, this embodiment uses a density-based, noise-robust spatial clustering algorithm (DBSCAN). This algorithm uses regions of sufficient density as distance centers and continuously grows these regions. A cluster can be uniquely identified by any core object within it. The algorithm utilizes the concept of density-based clustering, which requires that the number of objects within a certain region of the cluster space must be no less than a given threshold. Data flows from different accesses that are clustered into the same category have the same flow identifier. Each network flow is converted into a feature vector. This vector contains three global features used for flow alignment: the number of packets, the total size of sent packets, and the total size of received packets. Before all data flows from a particular website are input into DBSCAN, each data flow is represented using these three global features. Among all resulting clusters, flows that appear in more than 90% of access rounds are retained, while other flows are removed to prevent unnecessary interference from dynamic resources in the construction of the network behavior recognition model. The retained clusters are randomly sorted and sequentially assigned an integer starting from 1. This assigned integer is inherited by all data flows in the cluster as their common flow identifier.
[0038] Finally, in this embodiment, a set of statistical features is designed for each data flow to comprehensively characterize the network behavior represented by the data flow. Figure 2 , the statistical characteristics are as follows:
[0039] (a) Overall statistical characteristics:
[0040] In the process of analyzing data flows, overall statistical characteristics are used to comprehensively describe each flow, including:
[0041] Total Bidirectional Packets: The total number of all packets in the flow.
[0042] Total number of unidirectional packets: Counts the number of packets in each direction (i.e. sending and receiving directions).
[0043] Direction Ratio: Calculates the ratio of the number of packets in each direction to the total number of packets in both directions.
[0044] (b) Packet sequence characteristics:
[0045] The sequential characteristics of data packets focus on the interaction sequence between requests and responses and the density distribution of data packets in the time dimension, including:
[0046] Interaction order: Count the number of bidirectional data packets before its arrival, and then calculate the mean and standard deviation of these numbers as features reflecting the interaction order between requests and responses.
[0047] Time density: Counts the number of packets in each consecutive one-second time window and calculates the mean and standard deviation of these numbers.
[0048] Temporal density of packet size: The cumulative distribution of packet sizes is sampled and 120 interpolation points are extracted. These interpolation points serve as features that reflect the temporal distribution of packet sizes.
[0049] (c) Packet time characteristics:
[0050] Packet timing features focus on the time characteristics of packet arrival within a flow. They reveal the flow's temporal behavior and analyze packet arrival patterns and interrelationships, including:
[0051] Total number of packets before arrival: Extract the first 250 features of each flow. These features are derived from the total number of packets before each packet arrives, which helps understand the delay characteristics of the flow and the interaction order of packets.
[0052] The number of incoming packets between adjacent packets: Another 250 features are calculated by counting the number of incoming packets between every two adjacent first 301 outgoing packets, revealing the response time after the request is sent and the activity level of the interaction.
[0053] (d) Packet size characteristics:
[0054] The packet size characteristic analyzes the size distribution of each packet in the data flow and is an indicator for evaluating traffic characteristics.
[0055] Size interval division: The packet size is divided into intervals from 2^(s-1) to 2^s bytes (s is an integer between [6, 11]). By counting the number of packets in each size interval, the size distribution characteristics of the flow are plotted.
[0056] Statistical feature calculation: For the entire data flow and all subsequences of outgoing or incoming packets, the mean, median, standard deviation, third quartile, and sum of packet sizes are calculated. These statistical features can comprehensively reflect the distribution of packet sizes and reveal the change pattern of traffic.
[0057] (e) Header and tailer features:
[0058] Header Packet Selection: The first 30 packets in a flow are selected as header packets. These packets typically form the beginning of a flow and contain information about the initial request. Analyzing header packets can help understand the flow establishment process and its initial interactive characteristics.
[0059] Tail Packet Selection: The last 30 packets in a stream are selected as the tail packets. These packets typically contain response information at the end of the stream. Analyzing the tail packets can help understand the end characteristics of the stream and the final interaction content.
[0060] Inbound and Outbound Packet Statistics: Counts the number of outbound and inbound data packets in the header and trailer packets. This can reveal the interactive characteristics of the flow and help analyze the relationship between requests and responses.
[0061] All features are derived based on packet sequence statistics to illustrate the differences between data flows requesting different resources from different perspectives. The biggest differences between data flows requesting different resources are the total number of packets, the number of request packets, and the number of response packets. The proportion of packets in different directions to the total number of packets may also vary. Furthermore, the order of request and response interactions during the entire communication process varies, as does the rate of packet growth. Different resources are requested, the MSS sizes negotiated for TCP connections vary, and the size distribution of packets also varies. In this embodiment, for each data flow, a feature vector is extracted based on twelve feature descriptions: overall packet statistics for the proportion of packets in different transmission directions, packet burst patterns, packet growth rate and growth trend, packet time distribution, packet time interval information, packet distribution information, first and last packet information, packet transmission direction distribution information, packet request and response distribution statistics, packet size statistics, and packet size distribution statistics. In other embodiments, the number of statistical features can be reduced or increased as needed.
[0062] (2) Contextual semantic recognition
[0063] In this embodiment, contextual semantic recognition mainly includes data flow spatial relationship characterization based on contextual semantic recognition and data flow temporal relationship characterization based on contextual semantic recognition. Among them, the data flow spatial relationship characterization based on contextual semantic recognition proposes a new scheme for generating fingerprints for data flows, that is, measuring the similarity of data flows between websites from different angles. These fingerprints can quantify to what extent the appearance of a certain data flow represents the appearance of the website to which it belongs (data flow spatial relationship). The data flow temporal relationship characterization based on contextual semantic recognition further considers the data flow subsequences that appear fixedly when visiting the same website multiple times. These common subsequences reveal the data flow sequence pattern (data flow temporal relationship) of visiting the website. Contextual semantic recognition for encrypted traffic is constructed in combination with the spatial and temporal correlation of data flows. The two parts are explained separately below:
[0064] ① Characterization of data flow spatial relationships based on contextual semantic recognition: Due to the presence of advertisements on website pages and changes in dynamic resources, the traffic generated by repeated visits to a website is not exactly the same. From the perspective of data flow, the data flow sets in each traffic sample are not exactly the same. In this embodiment, the statistical characteristics of data flows and the similarity of data flows between different websites are characterized. The traffic statistical characteristics are used to characterize representative data flows and quantify the extent to which the appearance of representative data flows can indicate that the website to which they belong is visited, that is, the data flow website matching index is calculated.
[0065] In this process, a fingerprint is generated for each data stream based on the all-round characterization of the data stream through the random forest algorithm, aiming to reflect the similarity between multiple data streams of different websites from different angles. In the random forest algorithm, each decision tree is trained based on different training data stream samples and statistical features. For each sample, the random forest model uses the results of all decision trees (rather than the final decision made by the decision tree through voting) to construct a t-dimensional vector as the fingerprint of the sample, where t is the number of decision trees in the random forest. In this embodiment, the fingerprint of a data stream sample f can be expressed as (T1(f), T2(f),…, T t (f)), where T k (k = 1, 2, …, t) is the result of the kth decision tree for the current sample (different from the final result of the random forest). The output of each decision tree is the index of a leaf, which identifies the flow into which the current decision tree classifies x. Data flow samples are further represented using fingerprints, converting the extracted traffic features into the proposed high-level and homogeneous fingerprint vectors. This fully utilizes the ingenuity of random forests and avoids feature complexity analysis, dimensionality reduction, and normalization. Because the fingerprint of each data flow sample is a homogeneous vector, calculating the distance (i.e., similarity) between different samples is much easier than by homogenizing the original feature vectors.
[0066] Then, we define the data flow website matching index FW_ij, which evaluates the distinguishability of the data flow by calculating the proportion of samples with the same label among the neighboring samples. Specifically, the website matching index FW_ij is equal to the number of decision tree outputs for website I divided by the total number of decision trees in the random forest; the FW_ij value of the data flow flow(i,j) represents the probability that the user is currently visiting website I, where flow(i,j) represents the jth data flow of sample i under website I. If the FW_ij of the data flow flow(i,j) is small, it can only be guessed to a small extent that the user is currently visiting website I. Conversely, if the FW_ij of flow(i,j) is large, it can largely indicate that the user is visiting website I. The method for calculating the data flow website fingerprint is to count the proportion of votes of the decision tree pointing to website I in the fingerprint vector.
[0067] ② Characterization of data flow temporal relationships based on contextual semantic recognition: First, based on the data flow extracted by multivariate feature analysis, the website indication coefficient of the data flow is calculated; then, the traffic samples visiting the same website are split into data flow instance sequences, and the LCS is calculated to generate a serialized fingerprint; finally, the importance score of the serialized fingerprint is calculated using the website indication coefficient of the data flow.
[0068] In this embodiment, the temporal relationship of data streams is mainly characterized by combining similar information of data streams between different websites, mining common subsequences of data streams that frequently appear when accessing websites, and exploring the temporal correlation of data streams within the same website. The process is divided into two steps: serialized fingerprint generation based on LCS (longest common subsequence) and importance scoring of serialized fingerprints. Serialized fingerprint generation based on LCS focuses on the study of the relevant order of valid data stream instances, and generates corresponding serialized fingerprints by combining the labels of access data stream samples into sequences. The importance scoring of serialized fingerprints is evaluated by calculating the frequency of occurrence of each fingerprint in all data stream sequences and the Score (L) value of the corresponding data stream.
[0069] In order to quantify the extent to which the existence of a data flow can indicate that the website to which it belongs is visited, this embodiment defines an indicator called the data flow website indication coefficient (Flow Website-Indication Index, W2I). The website indication coefficient W2I shows the spatial correlation of multiple data flows from different websites, and represents the extent to which the existence of the current data flow sample can indicate the website to which it belongs. Specifically, for the j-th data flow flow (i, j) of sample i under website I, its website indication coefficient is defined as W2I ij, W2I shows the spatial relationship between data flows of different websites. The calculation method is as follows: bi-babel(i,j) is used as the label of the data flow, indicating that the data flow comes from website I, and among all the data flows of website I, it is the data flow identified as j. The data flow sample labeled bi-label(i,j) is called an instance of flow(i,j). For the data flow sample labeled bi-label(i,j), calculate the proportion of samples labeled bi-label(i,j) in its K nearest neighbor samples, and then calculate the average of the proportion values calculated for all data flow samples labeled bi-label(i,j) as the W2I of flow(i,j). ij The following formula is the calculation formula for W2I, where K is the parameter in the K nearest neighbor algorithm, I (i,j) is the total number of data stream samples with all labels bi-label(i,j).
[0070]
[0071] In actual applications, although the order of establishing TCP connections in data streams in multiple access records is not exactly the same, the order of some data stream instances can still be followed. The traffic samples generated by a visit to a website can be split into multiple data streams, and the labels of these data stream samples are combined into a sequence from early to late. Indicates that the length of the a-th visit to the current website is n a A data stream sequence, where Represents the i-th data stream. Similarly, Indicates that the length generated by the b-th visit to the current website is n b For every two data stream sequences, such as S a He He S b , generate a serialized fingerprint for it. Assume It's S a A subsequence of the reciprocal p elements of a sequence, It's S b The subsequence of the reciprocal q elements of the sequence, S a and S b The LCS is expressed as L(n a ,n b ), which is defined as S a and S b The serialized fingerprint of . Among them, L(p,q) is and LCS, Indicates that Append to the sequence L(p-1,q-1), for .
[0072] The calculation method of L(p,q) is as follows:
[0073]
[0074] In other words, traffic samples generated by visiting the same website are split into data stream instance sequences, and a serialized fingerprint is obtained by calculating the LCS between each sequence. Because the data stream sequences generated by each visit are not exactly the same, multiple serialized fingerprints are generated after pairwise processing. Some fingerprints are identical, but the frequency of occurrence of serialized fingerprints in all data stream sequences varies.
[0075] The importance of the serialized fingerprint can be scored based on the calculated data flow website indication coefficient. The calculation formula of Score(L) is as follows:
[0076]
[0077] Among them, L is the currently scored serialized fingerprint, Score(L) is the importance score of L, #occur(L) represents the number of times the serialized fingerprint appears in all data stream sequences, W2I ij The index coefficient for the website. The serialized fingerprints are sorted from highest to lowest according to their scores. The top 10 serialized fingerprints with the highest scores are selected. The serialized fingerprints containing only one data stream are then taken into account to construct a website fingerprint representing the current website visit behavior.
[0078] Through refined spatiotemporal correlation analysis of data streams, it is possible to accurately identify whether users are accessing websites of interest to analysts through encrypted proxies within a specific time window, improving the ability to detect and identify encrypted traffic.
[0079] (3) Deep identification of website fingerprints associated with multiple streams
[0080] In this embodiment, the results obtained from the spatial relationship characterization and temporal relationship characterization of the data flow during the contextual semantic recognition process are mainly input into the constructed random forest classification model to complete the iterative training of the classification model. Traffic identification can be directly completed through the trained classification model.
[0081] Given the success of the random forest algorithm in identifying network behavior for encrypted traffic, the use of random forest as the smallest unit of network behavior analysis—the classifier of data streams—belongs to the random forest algorithm under the ensemble learning branch. Its basic unit is the decision tree. A large number of decision trees (the origin of the forest) independently learn and make predictions, and then vote together to determine the final classification result. The introduction of two randomnesses makes each decision tree independent. Voting on the classification results of multiple weak classifiers (decision trees) forms a strong classifier (random forest). This decision-making method gives the random forest algorithm good noise resistance and is not prone to overfitting. There are two important parameters when constructing a random forest model, namely the number of decision trees t and the number of features m selected. The construction process of each decision tree is independent, and the steps are as follows:
[0082] (1) Assuming the number of training samples is N, randomly select n samples as the training sample set of the current decision tree;
[0083] (2) Assuming that the sample feature dimension is M, randomly select m features as the attributes based on which the current decision tree splits;
[0084] (3) The current decision tree is completely split until all training samples are accurately classified or all features are used up;
[0085] Repeat the steps of constructing a single decision tree t times to obtain a random forest model with t number of decision trees. The classification effect of random forest is related to two factors: (1) the correlation between any two decision trees in the forest, that is, the greater the correlation, the higher the error rate; (2) the classification ability of each decision tree in the forest, that is, the stronger the classification ability, the lower the error rate of the entire forest. The size of the number of features m of the decision tree determines the correlation between different decision trees and the classification ability of a single decision tree. If m is too small, the correlation between the decision trees is weak, but the classification ability of a single decision tree is weak; conversely, both will increase. Therefore, the key issue in constructing a random forest is how to choose the optimal m. The random forest algorithm is evaluated by calculating the out-of-bag error rate of the decision tree. When constructing a decision tree, the training set is randomly selected with replacement. For each decision tree, about 1 / 3 of the training samples do not participate in the generation of the current decision tree. These samples are called out-of-bag data of the current decision tree. By classifying the out-of-bag data on the generated decision tree, the out-of-bag error rate of each decision tree can be obtained to estimate the generalization error, without the need for cross-validation to evaluate the classification ability of the decision tree.
[0086] In a real network environment, a traffic segment captured over a given period of time is largely a mixture of traffic generated by users visiting target website I and many other websites. Whether each serialized fingerprint under the website fingerprint of target website I can be found in this segment may be affected by the interference of traffic from websites other than target website I, because data flow samples from other websites may be misclassified into the data flow category under target website I. In fact, understanding the traffic samples of the website fingerprint of target website I in a real network environment, that is, the performance of the website fingerprint in mixed traffic samples, is the key to building a network behavior classifier for identifying website I. Therefore, in this embodiment, by characterizing the data flow of mixed traffic samples in space and time, a website feature vector of the target website in the mixed traffic samples can be generated.
[0087] Specifically, the data stream's fingerprint vector, target website matching index, serialized fingerprint, and importance score are fed into a random forest classification model for iterative training. During this iterative training process, the training results are adjusted using the target website matching index and importance score. Mixed traffic samples that have visited target website I are defined as positive samples, while mixed traffic samples that have not visited target website I are defined as negative samples. Determining whether target website I has been visited now becomes a classic binary classification problem.
[0088] Finally, the actual mixed traffic is input into the trained random forest classification model, and the random forest classification model can determine whether the target website has been visited.
[0089] The purpose of this invention is to provide a method for identifying proxy encrypted traffic based on contextual multi-stream correlation. This method, which is targeted at encrypted proxies, studies machine learning-based contextual multi-stream correlation traffic identification technology to address the practicality of network behavior identification technology for encrypted traffic. From an information perspective, this technology will help regulators effectively identify and block illegal encrypted traffic and combat non-compliant network behavior. From an application perspective, through in-depth analysis of daily user encrypted traffic, it will provide enterprises and institutions with precise data protection and monitoring solutions, ensuring the overall security and stability of the network environment.
[0090] The technical advantages of the present invention are as follows:
[0091] ① We propose multi-flow fusion and correlation identification to improve fingerprint recognition accuracy: Traditional methods typically perform website fingerprint recognition based on a single network flow, which is susceptible to network traffic congestion and insufficient feature similarity. This method establishes a fine-grained mapping relationship between website resources and network flows through multi-flow fusion and correlation analysis. This method can explore the similarity of flow features caused by shared resources between different websites, thereby improving fingerprint recognition accuracy in complex network environments.
[0092] ② Solve the training-test asymmetry problem and enhance model robustness: Traditional fingerprint recognition methods often suffer from performance degradation during training and testing due to inconsistent traffic feature distribution, especially in complex network environments. This method, through fine-grained spatiotemporal analysis, proposes a new multi-stream fusion method that effectively addresses the training-test asymmetry problem caused by traffic congestion, significantly improving the model's generalization and robustness in real-world testing.
[0093] Those skilled in the art will understand the specific meanings of the above terms in the present invention in specific circumstances. The drawings in the embodiments are used to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0094] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A proxy encrypted traffic identification method based on contextual multi-flow association, characterized in that: include: Multivariate feature analysis: Collect mixed traffic samples generated by each round of website visits and perform data flow extraction; Perform dual-label calibration and statistical feature extraction on each data stream; Characterizing spatial relationships of data streams based on contextual semantic recognition: Based on the results of multivariate feature analysis, a random forest algorithm is used to generate fingerprint vectors for each data stream and calculate the target website matching index; Characterizing the temporal relationship of data streams based on contextual semantic recognition: This involves calculating the website indicator coefficient of data streams extracted through multivariate feature analysis. Traffic samples accessing the same website are split into data stream instance sequences, and the LCS is calculated to generate a serialized fingerprint. The importance score of the serialized fingerprint is calculated using the website indicator coefficient of the data stream. Deep identification of website fingerprints associated with multiple streams: Build a random forest classification model and perform iterative training of the random forest classification model based on the fingerprint vector of the data stream, the target website matching index, the serialized fingerprint of the data stream, and the importance score. Traffic identification can be completed through the trained classification model.
2. The method for identifying proxy encrypted traffic based on contextual multi-flow association according to claim 1, characterized in that: The double-label marking includes: the first label is the URL of the visited website, and the second label is the identifier of each data flow in all data flows of the website to which it belongs.
3. The method for identifying proxy encrypted traffic based on contextual multi-flow association according to claim 1 or 2, characterized in that: The statistical feature extraction includes: overall statistical features, data packet sequence features, data packet time features, data packet size features, head packet and tail packet features; wherein, the overall statistical features provide a comprehensive description of each flow; the data packet sequence features involve the interaction sequence between requests and responses and the density distribution of data packets in the time dimension; the data packet time features involve the time characteristics of the arrival of data packets in the data flow; the data packet size features involve the size distribution of each data packet in the data flow.
4. The method for identifying proxy encrypted traffic based on contextual multi-flow association according to claim 1, characterized in that: The fingerprint vector generation process includes: for each data stream sample, using the results of all decision trees in the random forest algorithm to construct a t-dimensional vector as the fingerprint vector of the data stream sample, where t is the number of decision trees in the random forest; the decision tree is trained based on different training data stream samples and statistical features.
5. The method for identifying proxy encrypted traffic based on contextual multi-flow association according to claim 1, characterized in that: The website matching index is equal to the number of decision trees whose output is the target website divided by the total number of decision trees in the random forest.
6. The method for identifying proxy encrypted traffic based on contextual multi-flow association according to claim 1, characterized in that: The calculation process of the website indication coefficient of the data stream includes: For the jth data flow flow(i,j) of sample i under website I, its website indication coefficient is defined as W2I ij The calculation method is as follows: the data flow sample labeled bi-label(i,j) is called an instance of flow(i,j). For the data flow sample labeled bi-label(i,j), calculate the proportion of samples labeled bi-label(i,j) in its K nearest neighbor samples. Then, the average of the proportion values calculated for all data flow samples labeled bi-label(i,j) is taken as the W2I of flow(i,j). ij .
7. The method for identifying proxy encrypted traffic based on contextual multi-flow association according to claim 1, characterized in that: The serialized fingerprint generation process includes: The traffic samples generated by visiting the same website are split into data stream instance sequences, and a serialized fingerprint is obtained by calculating the LCS between each sequence.
8. The method for identifying proxy encrypted traffic based on contextual multi-flow association according to claim 1, characterized in that: The importance score calculation process of the serialized fingerprint includes: Among them, Score(L) is the importance score of the serialized fingerprint, L is the currently scored serialized fingerprint, #occur(L) represents the number of times the serialized fingerprint appears in all data stream sequences, W2I ij Indicates the coefficient of the website.
9. The method for identifying proxy encrypted traffic based on contextual multi-flow association according to claim 1, characterized in that: The iterative training of the random forest classification model includes: The fingerprint vector of the data stream and the target website matching index, the serialized fingerprint of the data stream and the importance score are input into the random forest classification model for iterative training. During the iterative training process, the training results are adjusted according to the target website matching index and the importance score.
10. The method for identifying proxy encrypted traffic based on contextual multi-flow association according to claim 1, characterized in that: The method of completing traffic identification by using a trained classification model includes: inputting actual mixed traffic into a trained random forest classification model, and the random forest classification model can determine whether a target website has been visited.
Citation Information
Patent Citations
Encrypted website fine grit classification method and device based on different HTTP versions
CN111382780A
Encrypted traffic website identification method based on time-space association website fingerprints
CN116208506A