A Flow Association Method and System for Tor Anonymous Network Based on VAE and Subset Sum

Through the traffic association method based on the variational autoencoder and subset and algorithm, the accuracy and efficiency of traffic association in the Tor network are solved, and the accurate correlation of inlet and outlet traffic in the Tor anonymous network is achieved, which improves the accuracy and reliability of network security monitoring.

CN120128507BActive Publication Date: 2025-07-08NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510601475.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-08
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

After the existing Tor network traffic association technology faces the latest Tor protocol updates, especially the application of fill schemes, traffic packet count estimation becomes more difficult, resulting in a decrease in the accuracy and efficiency of traffic associations, making it difficult to effectively identify and analyze Tor egress traffic.

Method used

Using a method based on variational autoencoder (VAE) and subset and algorithm, the packet absolute time characteristics and packet size characteristics of inlet and exit traffic in Tor anonymous network are collected, and the sliding window is used to correlate traffic with the subset and algorithm of dynamic programming, and the hidden state is extracted in combination with the long and short-term memory network (LSTM) to perform feature extraction and association analysis.

Benefits of technology

It improves the accuracy and efficiency of the association of communication subjects in the Tor network, can quickly and accurately identify and associate incoming traffic and egress traffic, enhances network traceability capabilities, and prevents the risk of anonymous network abuse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128507B_ABST
    Figure CN120128507B_ABST
Patent Text Reader

Abstract

The present invention discloses a flow association method and system for the Tor anonymous network based on VAE and subset sum. The method includes: deploying traffic collectors to collect Tor anonymous network traffic and perform traffic splitting according to five-tuple components; extracting the packet absolute time feature and packet size feature of the ingress and egress traffic; initially screening flow pairs through the packet absolute time to obtain potential associated flow pairs; performing time slot division on the potential associated flow pairs and calculating the number of packets and packet size within each time slot to obtain the sequence of the number of data packets within the time slot and the sequence of packet sizes, constructing a feature matrix from the sequence of the number of packets and the sequence of packet sizes, and using a variational autoencoder for feature extraction to obtain the ingress flow feature sequence and the egress flow feature sequence; inputting the feature sequences into an association analyzer to complete the flow association of the Tor anonymous network. The present invention can quickly and accurately complete the association of Tor anonymous network traffic, thereby providing strong technical support for network traceability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer network security, and particularly relates to a flow association method and system for a Tor anonymous network based on VAE and subset sum. Background Art

[0002] Tor (The Onion Router) is a network protocol for anonymous communication, which is widely used in the fields of privacy protection and information security. Its core goal is to hide the user's IP address and communication content through multi-layer encryption and relay forwarding technologies, making it difficult to trace and analyze data traffic. This anonymity makes Tor an important tool for privacy protection and is also used to circumvent censorship and access restricted resources. However, this anonymous mechanism has also been exploited by some malicious actors, such as illegal transactions, data breaches, and malware propagation. Therefore, the analysis and tracking of Tor network traffic have become an important research direction in the field of network security.

[0003] Flow association is a traffic analysis technique where an attacker monitors and compares the traffic characteristics between different nodes in the Tor network to match the traffic entering (inflow) and leaving (outflow) the Tor network, which may lead to the de-anonymization of the IP addresses of communication endpoints. With the development of deep learning, researchers have started to explore deep learning-based flow association attacks. Nasr et al. proposed a deep learning-based flow association attack system Deepcorr: Strong flow correlation attacks on tor using deep learning in the top conference ACM CCS 2018 in the field of computer and communication security in 2018, with page numbers 1962 - 1976. It uses a convolutional neural network (CNN) to implement the flow association function specific to the Tor network, and can achieve an accuracy of 96% for traffic with a short observation length. Oh et al. proposed an improved flow association attack DeepCoFFEA: Improved Flow Correlation Attacks on Tor via Metric Learning and Amplification in the IEEE Symposium on Security and Privacy, 2022 in 2022, with page numbers 1915 - 1932. It uses deep learning to train a feature embedding network to map Tor flows and exit flows to a low-dimensional space for comparison, thereby reducing the computational cost. However, due to the latest updates of the Tor protocol, especially the application of padding schemes, it has become more difficult to estimate Tor packet counts. Therefore, there is a need to further improve the traffic association technology for the Tor anonymous network. Summary of the Invention

[0004] In view of the above problems, the purpose of the present invention is to provide a flow association method and system for Tor anonymous network based on VAE and subset sum, which aims to solve the traffic association problem existing in the current Tor network, especially for the analysis and identification of traffic in Tor exits. By effectively applying passive traffic analysis technology, the present invention aims to improve the accuracy and efficiency of communication entity association in the Tor network, so as to better address the anonymity and security challenges mentioned in the background art.

[0005] The method includes the following steps:

[0006] Step 1, collect the traffic on the specified relay node in the Tor anonymous network, and initially identify the types of traffic, where the types include ingress traffic and egress traffic. The ingress traffic and egress traffic respectively represent the traffic entering the Tor network and the traffic leaving the Tor network. Divide the network flow according to the five-tuple form for the ingress traffic and egress traffic and save them as F I and F E respectively, where the five-tuple refers to: source IP address, destination IP address, source port, destination port, and protocol type. Determine that the traffic with the same five-tuple is the same traffic;

[0007] Step 2, extract the packet absolute time feature T I and the packet size feature S I of the ingress traffic F I divided according to the five-tuple, as well as the packet absolute time feature T E and the packet size feature S E of the egress traffic F E divided according to the five-tuple;

[0008] Step 3, use the packet absolute time features T I and T E to preliminarily screen the ingress flow and egress flow pairs. For each ingress flow, according to the packet absolute time feature, match the packet absolute time feature with the egress flows starting and ending within the same time period, so as to obtain potential associated flow pairs P IE ;

[0009] Step 4, divide the potential associated flow pairs P IE into time slots, set the time slot length to T, and divide the original associated flow pairs P IE with a duration of t into time slots according to the time slot T. Count the number and size of data packets in each time slot to obtain the ingress flow time slot packet number sequence as , the ingress flow time slot packet size sequence as , and the egress flow time slot packet number sequence as , the size sequence of the egress flow time slot packets is ; construct the ingress flow time slot packet feature matrix = and the egress flow time slot packet feature matrix = , and use a variational autoencoder for feature extraction to obtain the ingress flow feature sequence and the egress flow feature sequence ;

[0010] Step 5, input the ingress flow feature sequence and the egress flow feature sequence into the association analyzer, and complete the association between the ingress traffic and the egress traffic of the Tor anonymous network through a sliding window and a subset sum algorithm based on dynamic programming.

[0011] Step 1 includes: using a collection tool to capture the original traffic, dividing the original traffic in the form of a five-tuple, and identifying the type of traffic according to one or more of the size, IP address, protocol, and port of the traffic passing through the specified relay node.

[0012] In Step 2, use a feature extraction tool to extract the packet arrival time, packet departure time, and packet size to obtain the ingress flow packet absolute time feature T I , the packet size feature, and the egress flow packet absolute time feature T E and the packet size feature S E .

[0013] Step 3 includes: analyzing the absolute timestamps of each packet in the ingress flow and the egress flow to obtain:

[0014] = ,

[0015] = ,

[0016] where is the nth packet absolute timestamp of the ingress flow packet absolute time feature, is the nth packet absolute timestamp of the egress flow packet absolute time feature;

[0017] Through the comparison of the absolute difference between the first packet absolute timestamps of the ingress flow and the egress flow packet absolute time features and the absolute difference between the nth packet absolute timestamps and the fixed threshold , if both absolute differences are less than or equal to , it is determined as a potential associated flow pair P IE , otherwise it is not a potential associated flow pair, and the formula is:

[0018] ,

[0019] ,

[0020] wherein, is a fixed threshold for determining whether two flows may be the same flow.

[0021] Step 4 includes: padding the feature matrix to ensure that each feature matrix has a fixed length x, where x is the maximum value of all time slot numbers , , and the feature matrix is expressed as:

[0022] ,

[0023] ,

[0024] wherein, 、 respectively represent the number of packets and the packet size of each time slot of the ingress flow, 、 respectively represent the number of packets and the packet size of each time slot of each egress flow.

[0025] Step 4 also includes:

[0026] Denote any one of 、 as X, and use the variational autoencoder VAE to extract features from X and learn the posterior distribution of the latent variable :

[0027] ,

[0028] where X is the input, Z is the latent variable, representing the traffic features after dimensionality reduction, are the parameters of the encoder, which are learned through training, is a normal distribution with a mean of and a variance of ;

[0029] Calculate the mean and variance according to the hidden state extracted by the long short-term memory network LSTM (LSTM, Long Short-Term Memory):

[0030] ,

[0031] ,

[0032] ,

[0033] ,

[0034] Among them, is the hidden state of the long short-term memory network, representing the feature information from the input sequence to the current time step t, is the data input at the current time step t, is the hidden state at the previous time step t - 1, is the final hidden state (feature at the last moment) after the LSTM processes all time steps, 、 is the weight matrix of the fully connected layer, used to map the hidden state to the parameters of the latent variable distribution, 、 is the bias term, and the function is used to calculate the standard deviation to ensure is a positive value;

[0035] Calculate the reparameterization trick to generate the latent variable Z:

[0036] ;

[0037] Among them, is the standard normal distribution random noise (mean 0, variance 1); is the random noise sampled from the standard normal distribution;

[0038] Adopt Poisson sampling to make the generated latent variable Z output an integer :

[0039] ,

[0040] Among them, is to perform Poisson sampling on the absolute value of the latent variable Z;

[0041] The VAE decoder reconstructs the traffic data:

[0042] Restore the input features from the latent variable Z:

[0043] ,

[0044] Among them, is the probability distribution of the decoder generating the data X given the latent variable Z; is a normal distribution, used to describe the distribution of the reconstructed data output by the decoder; is the predicted value of the VAE decoder, and the variance represents the fluctuation range of the reconstructed value;

[0045] The VAE encoder is optimized by minimizing the reconstruction error and the KL divergence:

[0046] ,

[0047] ,

[0048] ,

[0049] where L is the total loss function of the VAE encoder, is the reconstruction error, is the KL divergence, is the latent variable 's prior distribution, which is a standard normal distribution , and respectively represent the mean and variance of the latent variable distribution output by the VAE encoder, is the latent space dimension, used to control the dimension of the latent variable, is the weight of the KL divergence;

[0050] Obtain the inlet flow feature sequence and the outlet flow feature sequence :

[0051] ,

[0052] ,

[0053] where and respectively represent the d-th data of the absolute value of the latent variable generated by the inlet flow feature matrix and the d-th data of the absolute value of the latent variable generated by the outlet flow feature matrix; and are the inputs of the correlation analyzer.

[0054] Step 5 includes: For the inlet flow feature sequence and the outlet flow feature sequence , based on the obtained latent correlation flow pair P IE , associate each inlet flow feature sequence with the potentially correlated outlet flow feature sequence in the outlet flow feature sequence , given a sliding window of size , , calculate the scores for the inlet flow features and outlet flow features within each window in a sliding window manner;

[0055] Within each sliding window, for the first window, obtain:

[0056] ,

[0057] ,

[0058] Among them, nums is the feature set of the inlet flow feature sequence within each window, and target is the sum of the data of the potentially associated outlet flow feature sequence within each window. and are respectively the j-th data of the inlet flow feature sequence and the j-th data of the potentially associated outlet flow feature sequence within each window.

[0059] For the inlet flow feature set within each sliding window , determine whether there exists a subset such that the sum of the elements of the subset is equal to the total sum of the potentially associated outlet flow features . If it exists, return the vector , as the vector composed of the elements of the subset . If it does not exist, no return value is required:

[0060] ,

[0061] ,

[0062] Convert whether there exists a subset such that the sum of the elements of the subset is equal to the total sum of the potentially associated outlet flow features into: whether there exists a subset such that the sum of the elements of the subset is within ( . If it exists, return the vector , as the vector composed of the elements of the subset . If it does not exist, no return value is required:

[0063] ,

[0064] ,

[0065] where is a constant, generally set to 5, representing the allowable error range;

[0066] When there is a solution for the inlet flow and the potentially associated outlet flow within the sliding window, the assigned score is 1, indicating the similarity of the flow pair within the sliding window; if no solution is found, the assigned score is -1, indicating the dissimilarity of the flow pair within the sliding window.

[0067] When neither the entrance nor the potentially associated exit flow has sent or received data packets within the sliding window, or if the packets received by the entrance were not sent by the exit flow, the assigned score is 0, indicating that the flow pair similarity within the sliding window cannot be determined. The scoring formula is as follows:

[0068] ,

[0069] where is the window score and i is the window index;

[0070] After calculating the scores for all windows, a score adjustment process is carried out, emphasizing the contribution of consecutive windows with the same score and reducing the impact of isolated score windows, thereby enhancing the accuracy of the correlation analysis. Finally, the weights are recalculated based on the adjusted scores, and the following formula is used for score weighted sum adjustment:

[0071] ,

[0072] where is the window score after weighting, K is a parameter affecting the importance of consecutive windows with the same score, generally set to 0.1, is an auxiliary function used to count the consecutive score values s to the left of the window index i;

[0073] After obtaining the weighted scores, the final score S is obtained based on the number of windows:

[0074] ,

[0075] where is the window score after weighting, is the number of windows;

[0076] After obtaining the final scores for all windows, the average window score is taken to obtain the final similarity score, and the entrance flow with the highest score and the potentially associated exit flow pair are selected as the final associated flow pairs; for the final associated flow pairs, by comparing with the preset threshold which is generally set to 0.9. If the score is greater than or equal to the threshold, it is determined as a valid associated flow pair; specifically, if the score exceeds the threshold, it indicates a significant correlation between the entrance flow and the potentially associated exit flow, and thus is determined as a valid associated flow pair; otherwise, if the score is lower than the threshold, it indicates insufficient correlation between the flow pairs, and thus is excluded and cannot be determined as an associated flow pair.

[0077] This process ensures that only flow pairs with a high degree of matching are selected as associated flow pairs, thereby improving the accuracy and reliability of the traffic correlation analysis.

[0078] The present invention also provides a flow association system for a Tor anonymous network of VAE and subset sum implemented based on the above method, including:

[0079] A traffic collection module, configured to collect traffic on a specified relay node in the Tor anonymous network, initially identify whether it is ingress traffic or egress traffic, where ingress traffic and egress traffic respectively represent traffic entering and leaving the Tor network, and divide the ingress traffic and egress traffic into network flows in the form of a five-tuple and save them as F I and F E , where the five-tuple refers to: source IP address, destination IP address, source port, destination port, and protocol type, and traffic with the same five-tuple is considered to be the same traffic;

[0080] A feature extraction module, configured to extract the packet absolute time feature T I and the packet size feature S I of the ingress traffic F I divided according to the five-tuple, and the packet absolute time feature T E and the packet size feature S E of the egress traffic F E divided according to the five-tuple;

[0081] A preliminary screening module, configured to, using the packet absolute time features T I and T E , conduct a preliminary screening on the ingress flow and egress flow pairs. For each ingress flow, according to the packet absolute time feature, match the packet absolute time feature with the egress flows starting and ending within the same time period, so as to obtain potential associated flow pairs P IE ;

[0082] A time slot division and feature extraction module, configured to divide the obtained potential associated flow pairs P IE into time slots, set the time slot length to T, and divide the original duration t associated flow pairs P IE into time slots according to the time slot T, count the number and size of data packets in each time slot, and obtain the ingress flow time slot packet number sequence as , the packet size sequence as , the egress flow time slot packet number sequence as , the packet size sequence as ; construct a feature matrix: = and = , use a variational autoencoder for feature extraction, and obtain the ingress flow feature sequence , the egress flow feature sequence ;

[0083] An association analysis module, configured to use the obtained ingress flow characteristics and egress flow characteristics and input them into an association analyzer, and complete the association of the ingress traffic and egress traffic of the Tor anonymization network through a sliding window and a subset sum algorithm based on dynamic programming.

[0084] The present invention also provides an electronic device, including a processor and a memory, where the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method.

[0085] The present invention also provides a storage medium, storing a computer program or instruction, and when the computer program or instruction runs on a computer, the steps of the method are executed.

[0086] Beneficial effects: By collecting the traffic at the ingress and egress of the relay nodes in the Tor anonymization network and splitting the traffic; respectively extracting the absolute packet time characteristics and packet size characteristics of the ingress and egress traffic; initially screening the flow pairs through the absolute packet time to obtain potential associated flow pairs; performing time slot division on the potential associated flow pairs and calculating the number of packets and packet size within each time slot to obtain the packet number sequence and packet size sequence within the time slot, constructing a matrix for the packet number sequence and packet size sequence, and using a variational autoencoder to extract features to obtain the ingress flow feature sequence and egress flow feature sequence; inputting the obtained feature sequences into the association analyzer, and using a sliding window and a dynamic programming subset sum algorithm to complete the flow association in the Tor anonymization network. The present invention can quickly and accurately associate the ingress traffic and egress traffic in the Tor anonymization network, thereby providing strong technical support for network traceability. Description of the Drawings

[0087] Figure 1 is a flowchart of the flow association method for the Tor anonymization network according to the present invention.

[0088] Figure 2 is a schematic diagram of collecting the ingress traffic and egress traffic in the Tor anonymization network according to the present invention.

[0089] Figure 3 is a schematic diagram of the structure of the variational encoder used in the present invention.

[0090] Figure 4 is a schematic diagram of the working process of the association analyzer according to the present invention.

[0091] Figure 5 is a schematic diagram of the flow association effect according to the present invention. Detailed Embodiments

[0092] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings.

[0093] Reference Figure 1 A flow association method for a Tor anonymous network based on VAE and subset sum provided by the present invention includes the following steps:

[0094] Step 1: Deploy a traffic collector to collect the ingress traffic and egress traffic on specified relay nodes of the Tor anonymous network and perform preliminary traffic splitting.

[0095] In the Tor anonymous network, traffic capture is performed by deploying a traffic collector on key relay nodes in the Tor network. The relay nodes act as data transfer stations, responsible for delivering data from the ingress to the target server or egress. These relay nodes can be ordinary computers, routers, or servers and play an important role in the Tor network. By deploying a traffic collector on one or more relay nodes, all traffic data passing through these nodes can be effectively captured.

[0096] According to an embodiment of the present invention, as Figure 2 shown, the traffic collector is deployed on key nodes of the Tor routing system, specifically including network ingress nodes and egress nodes. These core relay nodes undertake the two-way traffic forwarding function of Tor anonymous communication, handling both the user request traffic from the ingress direction and managing the data transfer task leading to the egress node. The collector deployed at this location can completely capture the two-way communication data flowing through the node and store the original network traffic in a standardized PCAP (Packet Capture) format, forming a time series data set containing complete packet header information, providing a standardized data basis for subsequent traffic feature extraction and association analysis.

[0097] For the captured traffic, initially distinguish whether the traffic is ingress traffic or egress traffic of the Tor anonymous network according to one or more of the information such as the IP address, traffic direction, protocol, and port of the traffic.

[0098] Differentiation method based on IP address and traffic direction: For incoming traffic, the source IP address is usually an IP address outside the Tor network, while the destination IP address is the IP address of an exit within the Tor network. For outgoing traffic, the source IP address is the IP address of an exit within the Tor network, and the destination IP address is still an IP address outside the Tor network. The Tor network usually uses ".onion addresses", which are only generated and used within the Tor network. Therefore, by determining whether the destination address of a data packet is an.onion address, it is possible to identify whether it is outgoing traffic within the Tor network. For external traffic, it can be differentiated by determining whether the destination address of the data packet belongs to a known public IP address range. In addition, since the Tor network forwards traffic through relay nodes, the source and destination addresses of external traffic usually change. Therefore, it is possible to identify whether it is external traffic by analyzing the relay node information in the data packet.

[0099] Differentiation method according to protocol and port: For incoming traffic, it usually enters the network through the Tor default ports (such as TCP 9050 or 9150), which are used to establish an encrypted link between the Tor client and the entry node. The protocol type is TLS encrypted transmission, and the data packet content is multi-layer encrypted ciphertext without clear text protocol features (such as HTTP headers). For outgoing traffic, the source port is usually a high port dynamically assigned by the Tor exit node (ranging from 49152 to 65535), and the protocol may be plain text HTTP or end-to-end encrypted HTTPS.

[0100] Through the above method, the collected traffic is divided into incoming traffic and outgoing traffic. The collected traffic is generally in the PCAP format. The incoming traffic and outgoing traffic are divided into network flows in the form of a five-tuple and saved as F I and F E . The five-tuple refers to: source IP address, destination IP address, source port, destination port, and protocol type. A network flow with the same five-tuple data is considered the same flow. By doing so, the data packets in the incoming flow F I and the outgoing flow F E are split according to whether the five-tuples are the same, and the processed data packets are repackaged and written into the pcap file.

[0101] Step 2: Use a feature extractor to extract the absolute packet time feature T and the packet size feature S of the incoming traffic F I and the outgoing traffic F E respectively.

[0102] According to the embodiments of the present invention, for the incoming traffic F I and the outgoing traffic F E , use the feature extraction tool CICFlowMeter to extract FI and F E The absolute packet time T and packet size feature S of. CIC FlowMeter is a tool for traffic measurement and analysis. It can extract and analyze features from pcap files. Use this tool to extract features from the shunted pcap files in step 1 to generate corresponding pickle feature files. Select the packet arrival time and packet size feature data in the pickle file. Finally, obtain the absolute packet time feature T I of the packets of Tor network entry traffic F I and packet size feature S I , as well as the absolute packet time feature T E of the packets of Tor network exit traffic F E and packet size feature S E .

[0103] Step 3: Use the preliminary screener to pair each entry flow with the exit flows that start and end within the same period using the absolute packet time feature, and screen out potential associated flow pairs P IE .

[0104] Analyze the absolute packet timestamps of each packet in the entry flow and the exit flow to obtain:

[0105] = ,

[0106] = ,

[0107] where is the nth absolute packet timestamp of the absolute packet time feature of the entry flow packet, is the nth absolute packet timestamp of the absolute packet time feature of the exit flow packet;

[0108] By comparing the absolute difference between the first absolute packet timestamps of the absolute packet time features of the entry flow and the exit flow and the absolute difference between the nth absolute packet timestamps with a fixed threshold If both absolute differences are less than or equal to , it is determined as a potential associated flow pair P IE , otherwise it is not a potential associated flow pair. The formula is as follows:

[0109] ,

[0110] ,

[0111] where, is the fixed threshold, which is set to 500ms in the present invention and is used to determine whether two flows may be one flow;

[0112] Step 4: Use a time slot divider to perform time slot division on potential associated flow pairs, obtain the packet quantity and packet size features, construct a feature matrix, and use a variational autoencoder to extract features from the feature matrix to obtain the input of the association analyzer.

[0113] By performing time slot division on the obtained potential associated flow pair P IE First, set the length of the time slot as T = 500 ms. Based on the original duration t of each associated flow, divide it according to the time slot length T to obtain n time slots, where , indicating rounding to the nearest integer. According to the time attributes of the data packets, count the number and size of the data packets in each time slot to obtain the entrance flow time slot packet quantity sequence as , the packet size sequence as , the exit flow time slot packet quantity sequence as , and the packet size sequence as . Construct the feature matrix = and = . The feature matrix is constructed as follows:

[0114] ,

[0115] ,

[0116] where represents the entrance flow matrix, represents the exit flow matrix. The first column represents the packet quantity sequence in each time slot, and the second column represents the packet size sequence in each time slot.

[0117] The present invention uses VAE to extract features from the feature matrix. The variational autoencoder (VAE) can remove noise and redundant information while retaining the core features of the traffic data. Taking the traffic feature matrix as an example, the following are the operation steps for using VAE to reduce the dimension and extract features of the traffic features:

[0118] (a) Data filling

[0119] Since the number of time slots n of the potential associated flow pairs is different, fill the feature matrix to ensure that each n has a fixed length x, where x is the maximum value of all the number of time slots. In the present invention, the maximum traffic duration is 8 min, that is . For n < x, perform zero padding. For the sake of easy description, hereinafter is used to represent , in any one of them, is expressed as follows:

[0120] ;

[0121] (b)Data standardization

[0122] The original data is mapped into a normal distribution with a mean of 0 and a standard deviation of 1 for each column using Z-score standardization, so that the distribution of all data approaches a normal distribution. The result is shown in the following formula:

[0123] ;

[0124] Where:

[0125] , ,

[0126] , ;

[0127] (c)Encoder feature extraction:

[0128] The encoder is used to extract features and learn the posterior distribution of the latent variable :

[0129] ,

[0130] where X is the input, Z is the latent variable, representing the dimensionality-reduced flow characteristics, are the parameters of the encoder, which are learned through training, is a normal distribution with a mean of and a variance of , the mean and variance are calculated from the hidden states extracted by LSTM:

[0131] ,

[0132] ,

[0133] ,

[0134] ,

[0135] Where, is the LSTM hidden state, representing the feature information of the input sequence up to the current time step t, is the data input at the current time step t, is the hidden state at the previous time step t-1, is the final hidden state (the feature at the last time step) after the LSTM processes all time steps, 、 is the weight matrix of the fully connected layer, used to map the hidden state to the parameters of the latent variable distribution, , is the bias term, Calculate the standard deviation to ensure that it is positive;

[0136] (d) The reparameterization trick generates the latent variable Z:

[0137] ,

[0138] where, is the standard normal distribution random noise (mean 0, variance 1); is the random noise sampled from the standard normal distribution;

[0139] Use Poisson sampling to make the output an integer: ;

[0140] where, is the Poisson sampling of the absolute value of the latent variable Z;

[0141] (e) The decoder of the VAE reconstructs the traffic data:

[0142] Restore the input features from the latent variable Z:

[0143] ,

[0144] where, is the probability distribution of the decoder generating the data X given the latent variable Z, is a normal distribution, used to describe the distribution of the reconstructed data output by the decoder, is the predicted value of the decoder, and the variance represents the fluctuation range of this reconstructed value;

[0145] (f) The VAE is optimized by minimizing the reconstruction error and the KL divergence:

[0146] ,

[0147] ,

[0148] ,

[0149] where, L is the total loss function of the VAE, is the reconstruction error, is the KL divergence, is the prior distribution of the latent variable which is the standard normal distribution , and are the mean and variance of the latent variable distribution output by the encoder, , where is the latent space dimension, used to control the dimension of the latent variable, is the weight of the KL divergence, set to 0.1;

[0150] Such as Figure 3 As shown, the VAE used in the present invention is built based on LSTM, and the model structure is: input layer - LSTM encoder - sampling layer - LSTM decoder - output layer. Specifically, the LSTM encoder contains 128 LSTM units, which are used to capture the dependencies of time series data and extract hidden states ; Subsequently, the dimensions of the fully connected layers FC_μ and FC_σ are 64, and the mean and variance of the latent variable are calculated respectively, and then the latent variable is generated by using the reparameterization trick in the sampling layer , The dimension of is d = 64, and Poisson sampling is used to make the output an integer Z′, ensuring that the model is differentiable and enhancing the generalization ability. In the LSTM decoder part, first, the dimension of the fully connected layer FC_1 is 64, and is mapped to the initial hidden state of the LSTM, and then 128 LSTM units are used to generate the reconstructed time series. Finally, the dimension of the fully connected layer FC_2 is 2, and the output of the decoder is mapped to the original data dimension to generate a reconstructed time matrix X′ that matches the input. Among them, the role of the LSTM encoder is to extract the temporal features of the traffic data and learn the key patterns; the sampling layer realizes the regularization of the data distribution through the VAE mechanism, enhancing the generalization ability of the model; the role of the LSTM decoder is to generate the traffic time series by using the latent variable Z, thus ensuring the effectiveness of feature extraction and removing noise during the reconstruction process, improving the representation ability of traffic features.

[0151] The dimensionality reduction and feature extraction in the above steps (a) to (f) are to extract features from the constructed feature matrix through the VAE encoder, learn the latent variable distribution, and learn an effective latent feature representation by minimizing the reconstruction error and the KL divergence. The encoder of the trained VAE is used to obtain the inlet flow feature sequence and the outlet flow feature sequence , and are the inputs of the correlation analyzer. and The format of is as follows:

[0152] ,

[0153] ,

[0154] Among them, and respectively represent the d-th data of the absolute value of the latent variable generated by the inlet flow feature matrix and the outlet flow feature matrix;

[0155] Step 5: Use the sliding window and the dynamic programming subset sum algorithm to calculate the score and judge the threshold of the input features, and complete the association between the inlet flow and the outlet flow of the Tor anonymous network.

[0156] As Figure 4 shown, the working process of the association analyzer constructed by the present invention is: the inlet flow features obtained in step 4 and the outlet flow features are input into the association analyzer. The association analyzer processes the flow data in the window by using the sliding window and the subset sum algorithm based on dynamic programming.

[0157] For the inlet flow feature sequence and the outlet flow feature sequence , based on the potentially associated flow pair P IE obtained above, each inlet flow feature sequence is associated with the potentially associated outlet flow feature sequence in the outlet flow feature sequence. Given a sliding window of size , the sliding window moves with a step size of 2 on the data in each sequence of the inlet flow feature sequence and the potentially associated outlet flow feature sequence. Each time it moves, the scores of the inlet flow features and the outlet flow features in each window are calculated:

[0158] ,

[0159] ,

[0160] Among them, is the outlet flow feature sequence, is the d-th data of the inlet flow feature sequence, is the potentially associated outlet flow feature sequence, is the d-th data of the potentially associated outlet flow feature sequence;

[0161] For each sliding window, taking the first window as an example, we get:

[0162] ,

[0163] ,

[0164] ​For the set of ingress flow features within each sliding window , it is necessary to determine whether there exists a subset such that the sum of the elements of this subset is equal to the total of the egress flow features . If a subset exists, return the vector , which is the vector composed of the elements in the subset . If not, there is no need to return a value:

[0165] ,

[0166] ,

[0167] Transform the determination of whether there exists a subset such that the sum of the elements of the subset is equal to the total of the egress flow features into: whether there exists a subset such that the sum of the elements of the subset is between ( . Here, =5. If a subset exists, return the vector , which is the vector composed of the elements in the subset . If not, there is no need to return a value:

[0168] ,

[0169] ,

[0170] where is a constant, generally set to 5, representing the allowable error range;

[0171] When there is a solution for the ingress flow and the potentially associated egress flow within the sliding window, assign a score of 1, indicating that the flow pair within this window has similarity; if no solution is found, assign a score of -1, indicating that the flow pair within this window does not have similarity. Further, when neither the ingress nor the potentially associated egress flow has sent or received data packets within this window, or if the packets received by the ingress were not sent by the egress flow, assign a score of 0, indicating that the similarity of the flow pair within this window cannot be determined. The scoring formula is as follows:

[0172] ,

[0173] After calculating the scores of all windows, perform score adjustment processing, emphasizing the contribution of consecutive identical-score windows and reducing the impact of isolated score windows, thereby enhancing the accuracy of correlation analysis. Finally, recalculate the weights based on the adjusted scores, and use the following formula for score weighting and adjustment to improve the recognition accuracy of the flow similarity. The formula for recalculating the score weights is as follows:

[0174] ,

[0175] where, is the window score after weighting, is the initial score, i is the score number, and K is a parameter that affects the importance of consecutive windows with the same score, set to 0.1; is an auxiliary function, and respectively represent the number of windows with consecutive scores of 1 and -1 before the current window. In this way, greater weights are given to windows with consecutive similar or dissimilar features, making the final score better reflect the true relationship between flow pairs.

[0176] After obtaining the weighted scores, obtain the final score S according to the number of windows:

[0177] ,

[0178] where, is the window score after weighting, i is the window index, is the number of windows. In the present invention ;

[0179] After obtaining the final scores of all windows, select the entry flow and the potentially associated exit flow pair with the highest score as the final associated flow pair. Compare the score of this flow pair with the preset threshold . If the score is greater than or equal to the threshold, it is determined as a valid associated flow pair. This indicates that there is a significant correlation between the entry flow and the potentially associated exit flow; conversely, if the score is lower than the threshold, it means that the correlation between the flow pair is insufficient, and it is excluded and not recognized as an associated flow pair. Through such processing, it is ensured that only flow pairs with a high matching degree are selected as associated flow pairs, thereby improving the accuracy and reliability of the correlation analysis of the entry traffic and the exit traffic in the Tor anonymous network.

[0180] Based on the same technical concept as the method embodiment, the present invention also provides a flow association system for the Tor anonymous network based on VAE and subset sum, including:

[0181] The traffic collection module is used to collect the traffic on a specified relay node in the Tor anonymous network, initially identify whether it is incoming traffic or outgoing traffic, where the incoming traffic and the outgoing traffic respectively represent the traffic entering and leaving the Tor network, and divide the incoming traffic and the outgoing traffic into network flows in the form of a five-tuple and save them as F I and F E , where the five-tuple refers to: source IP address, destination IP address, source port, destination port, and protocol type. Traffic with the same five-tuple is considered to be the same traffic;

[0182] The feature extraction module is used to extract the absolute packet time feature T I and the packet size feature S I of the incoming traffic F I divided according to the five-tuple, as well as the absolute packet time feature T E and the packet size feature S E of the outgoing traffic F E divided according to the five-tuple;

[0183] The preliminary screening module is used to preliminarily screen the incoming and outgoing traffic pairs by using the absolute time features T I and T E . Specifically, for each incoming flow, according to its absolute time feature, it is matched with the outgoing flows that start and end within the same time period to obtain potential associated flow pairs P IE .

[0184] The time slot division and feature extraction module is used to divide the obtained potential associated flow pairs P IE into time slots, set the time slot length to T, and divide the original associated flow pairs P IE with a duration of t into time slots according to the time slot T, count the number and size of the data packets in each time slot, and obtain the incoming flow time slot packet number sequence as , the packet size sequence as , the outgoing flow time slot packet number sequence as , the packet size sequence as ; construct the feature matrix: = and = , use the variational autoencoder for feature extraction to obtain the incoming flow feature sequence , the outgoing flow feature sequence ;

[0185] The association analysis module is used to associate the obtained incoming flow features and outgoing flow features and An input correlation analyzer completes the correlation of the entry traffic and the exit traffic of the Tor anonymous network through a sliding window and a subset sum algorithm based on dynamic programming.

[0186] It should be understood that a flow correlation system of a Tor anonymous network based on VAE and subset sum in an embodiment of the present invention can implement all the technical solutions in the above method embodiments. The functions of its respective functional modules can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can refer to the relevant descriptions in the above embodiments and will not be elaborated here.

[0187] As Figure 5 shown, by using a flow correlation method and system of a Tor anonymous network based on VAE and subset sum of the present invention, it is ultimately possible to achieve the precise correlation of the anonymous entry traffic and the exit traffic in the Tor network. Specifically, by inputting the collected entry traffic and exit traffic of the Tor anonymous network, after being processed by the method or system of the present invention, it is possible to effectively identify and correlate the entry traffic and the exit traffic.

[0188] In addition, through efficient traffic analysis technology, the method of the present invention can not only effectively prevent the risk of possible abuse of the anonymous network, but also provide a more accurate and reliable monitoring means for the field of network security.

[0189] The present invention also provides a computer device, including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors. When the program is executed by the processor, it implements the steps of the method as described above.

[0190] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method as described above.

[0191] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, ID-REM, optical memories, etc.) containing computer-usable program codes.

[0192] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device create means for implementing the functions specified in the flowchart Figure 1 one flow or more flows and / or blocks Figure 1 one block or more blocks.

[0193] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flowchart Figure 1 one flow or more flows and / or blocks Figure 1 one block or more blocks.

[0194] These computer program instructions may also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flowchart Figure 1 one flow or more flows and / or blocks Figure 1 one block or more blocks.

Claims

1. A flow association method for Tor anonymous network based on VAE and subset sum, characterized in that It includes the following steps: Step 1: Collect the traffic on the specified relay nodes in the Tor anonymous network and initially identify the types of traffic. The types include incoming traffic and outgoing traffic, where incoming traffic and outgoing traffic respectively represent the traffic entering the Tor network and the traffic leaving the Tor network. Divide the incoming traffic and outgoing traffic into network flows in the form of five-tuples and save them as F I and F E , where the five-tuple refers to: source IP address, destination IP address, source port, destination port, and protocol type. Determine that traffic with the same five-tuple is the same traffic flow; Step 2: Extract the packet absolute time feature T I and the packet size feature S I of the incoming traffic F divided according to the five-tuple, I as well as the packet absolute time feature T E and the packet size feature S E of the outgoing traffic F divided according to the five-tuple; E ​ Step 3, utilize the packet absolute time feature T I and T E , perform a preliminary screening on the inlet flow and outlet flow pairs. For each inlet flow, according to the packet absolute time feature, match the packet absolute time feature with the outlet flows that start and end within the same time period, so as to obtain potential associated flow pairs P IE ; Step 4, the potential associated flow pair P IE is divided into time slots with the time slot length set to T, and the original associated flow pair P with a duration of t IE is divided according to the time slot T into time slots, and the number and size of data packets in each time slot are counted to obtain the sequence of the number of packets in the inlet flow time slots as the sequence of the size of packets in the inlet flow time slots as the sequence of the number of packets in the outlet flow time slots as the sequence of the size of packets in the outlet flow time slots as Construct the feature matrix of packets in the inlet flow time slots and the feature matrix of packets in the outlet flow time slots Feature extraction is performed using a variational autoencoder to obtain the inlet flow feature sequence F I and the outlet flow feature sequence F E ; Step 5, input the inlet flow feature sequence F I and the outlet flow feature sequence F E into the association analyzer, and complete the association between the inlet traffic and the outlet traffic of the Tor anonymous network through a sliding window and a subset sum algorithm based on dynamic programming, including: within each sliding window, for the first window, obtain: Among them, nums is the feature set of the entry flow feature sequence within each window, and target is the sum of the data of the potentially associated exit flow feature sequence within each window. and are respectively the j-th data of the entry flow feature sequence and the j-th data of the potentially associated exit flow feature sequence within each window. For the set of ingress flow features within each sliding window Determine whether there exists a subset such that the sum of the elements of the subset nums′ is equal to the total target of the potentially associated egress flow features. If it exists, return the vector F is the vector composed of the elements in the subset nums′. If it does not exist, there is no need to return a value: When Whether there exists a subset such that the sum of the elements of the subset nums′ is equal to the total target of the potentially associated exit flow characteristics is transformed into: Whether there exists a subset such that the sum of the elements of the subset nums′ is between (target - Δ, target + Δ). If it exists, return the vector T′ which is the vector composed of the elements in the subset nums′. If it does not exist, there is no need to return a value: When where Δ is a constant.

2. The method according to claim 1, wherein Step 1 includes: capturing the original traffic, dividing the original traffic in the form of quintuples, and identifying the type of traffic according to one or more of the size, IP address, protocol, and port of the traffic passing through the specified relay node obtained.

3. The method according to claim 2, wherein In step 2, the arrival time of the data packet, the packet sending time, and the data packet size are extracted to obtain the absolute time feature T of the ingress flow packet I , the packet size feature, and the absolute time feature T of the egress flow packet E and the packet size feature S E .

4. The method according to claim 3, wherein Step 3 includes: analyzing the absolute timestamps of each packet in the ingress flow and egress flow to obtain: T I = [T I1 , T I2 ··· T In , T E = [T E1 , T E1 ··· T En , where T In is the nth packet absolute timestamp of the absolute time feature of the ingress flow packet, and T En is the nth packet absolute timestamp of the absolute time feature of the egress flow packet; By comparing the absolute difference between the absolute timestamps of the first packet of the packet absolute time characteristics of the inlet flow and the outlet flow and the absolute difference between the absolute timestamps of the nth packet and the fixed threshold Δt, if both absolute differences are less than or equal to Δt, it is determined as a potentially associated flow pair P IE , otherwise it is not a potentially associated flow pair, and the formula is: |T I1 -T E1 |≤Δt, |T In -T En |≤Δt, where Δt is a fixed threshold for determining whether two flows may be the same flow.

5. The method according to claim 4, wherein Step 4 includes: padding the feature matrix to ensure that each feature matrix has a fixed length x, where x is the maximum value n of the number of all time slots max , x = n max , the feature matrix is represented as: wherein, respectively represent the number of packets and the packet size of each time slot of the ingress flow, respectively represent the number of packets and the packet size of each time slot of each egress flow.

6. The method according to claim 5, characterized in that Step 4 also includes: Let X represent X I and X E Among any one of them, use the variational autoencoder VAE to extract features from X and learn the posterior distribution of the latent variable where X is the input and Z is the latent variable, are the parameters of the encoder, and N(μ, σ 2 ) is a normal distribution with mean μ and variance σ 2 . Calculate the mean μ and variance σ based on the hidden state extracted by the long short-term memory network LSTM 2 : h t = LSTM(X t, h t-1 ), μ = W μ h T + b μ , logσ 2 = W σ h T + b σ , σ = exp(0.5·logσ 2 ) Among them, h t is the hidden state of the long short-term memory network, representing the feature information of the input sequence up to the current time step t, X t is the data input at the current time step t, h t-1 is the hidden state at the previous time step t - 1, h T is the final hidden state after the LSTM processes all time steps, W μ and W σ are the weight matrices of the fully connected layer, b μ and b σ are the bias terms, and the function exp(0.5·logσ 2 ) is used to calculate the standard deviation; calculating the reparameterization trick to generate the latent variable Z: Z = μ + σ·∈, ∈~N(0,I); where ∈ is a standard normal distribution random noise; σ is a random noise sampled from the standard normal distribution; using Poisson sampling to make the generated latent variable Z output as an integer Z′: Z′ = Poisson(|Z|), where Poisson(|Z|) is Poisson sampling of the absolute value of the latent variable Z; the VAE decoder reconstructs the traffic data: restoring the input features from the latent variable Z: p θ (X|Z) = N(X; g θ (Z), σ 2 ) where p θ (X|Z) is the probability distribution of the decoder generating data X given the latent variable Z; N(X; g θ (Z), σ 2 ) is a normal distribution; g θ (Z) is the predicted value of the VAE decoder, and the variance σ 2 represents the fluctuation range of the reconstructed value; the VAE encoder is optimized by minimizing the reconstruction error and KL divergence: L = L rec + βD KL , L rec = ||X - X'|| 2 , where \(L\) is the total loss function of the VAE encoder, \(L\) rec is the reconstruction error, is the KL divergence, \(p(Z)\) is the prior distribution of the latent variable \(Z\), which is the standard normal distribution \(N(0, I)\), \(\mu\) i and respectively represent the mean and variance of the latent variable distribution output by the VAE encoder, \(d\) is the latent space dimension, and \(\beta\) is the weight of the KL divergence; Obtain the inlet flow characteristic sequence F I and the outlet flow characteristic sequence F E : Among them, and respectively represent the d-th data of the absolute value of the latent variable generated by the inlet flow characteristic matrix and the d-th data of the absolute value of the latent variable generated by the outlet flow characteristic matrix; F I and F E are the inputs of the association analyzer.

7. The method according to claim 6, characterized in that, Step 5 includes: for the inlet flow feature sequence F I and the outlet flow feature sequence F E , based on the obtained potential associated flow pair P IE , associate the inlet flow feature sequence F I with the potentially associated outlet flow feature sequence FE′ in the outlet flow feature sequence F E for correlation analysis. Given a sliding window of size j, where j < d, calculate the scores for the inlet flow features and outlet flow features within each window in a sliding window manner; When there is a solution for the ingress flow and the potentially associated egress flow within the sliding window, the assigned score is 1, indicating that the flow pairs within the sliding window have similarity; if no solution is found, the assigned score is -1, indicating that the flow pairs within the sliding window do not have similarity; When neither the ingress nor the potentially associated egress flow has sent or received data packets within the sliding window, or if the packets received by the ingress are not sent by the egress flow, the assigned score is 0, indicating that the similarity of the flow pairs within the sliding window cannot be determined. The score formula is as follows: where scores[i] is the window score and i is the window index; After calculating the scores of all windows, perform score adjustment processing. Finally, recalculate the weights based on the adjusted scores and use the following formula for score weighted sum adjustment: where scores′[i] is the window score after weighting, K is a parameter affecting the importance of consecutive windows with the same score, and c(s,i) is an auxiliary function for counting the consecutive score values s on the left side of the window index i; After obtaining the weighted scores, obtain the final score S according to the number of windows: where scores′[i] is the window score after weighting and p is the number of windows; After obtaining the final scores of all windows, take the average window score to obtain the final similarity score, and select the ingress flow and the potentially associated egress flow pair with the highest score as the final associated flow pair; for the final associated flow pair, by comparing with the preset threshold thr, if the score is greater than or equal to the threshold, it is determined as a valid associated flow pair; specifically, if the score exceeds the threshold, it indicates that there is a significant correlation between the ingress flow and the potentially associated egress flow, so it is determined as a valid associated flow pair; otherwise, if the score is lower than the threshold, it indicates that the correlation between the flow pairs is insufficient, so it is excluded and cannot be determined as an associated flow pair.

8. A flow association system for a Tor anonymous network of VAE and subset sum implemented based on the method according to any one of claims 1 to 7, characterized in that, It includes: A traffic collection module is used to collect the traffic on a specified relay node in the Tor anonymous network, initially identify whether it is incoming traffic or outgoing traffic, where the incoming traffic and the outgoing traffic respectively represent the traffic entering and leaving the Tor network, and divide the incoming traffic and the outgoing traffic into network flows in the form of five-tuples and save them as F I and F E , where the five-tuple refers to: source IP address, destination IP address, source port, destination port, and protocol type, and the traffic with the same five-tuple is considered to be the same traffic; A feature extraction module, which is used to extract the packet absolute time feature T I and the packet size feature S I of the incoming traffic F divided according to the five-tuple, as well as the packet absolute time feature T I and the packet size feature S E of the outgoing traffic F divided according to the five-tuple; E E ​​ The primary screening module is used to utilize the packet absolute time features T I and T E , to conduct a primary screening on the inlet flow and outlet flow pairs. For each inlet flow, according to the packet absolute time features, match the packet absolute time features with the outlet flows that start and end within the same time period, so as to obtain potential associated flow pairs P IE ; Time slot division and feature extraction module, which is used to obtain the potential associated flow pair P IE Perform time slot division, set the time slot length to T, and divide the original associated flow pair P with a duration of t IE According to the time slot T into Time slots, count the number and size of data packets in each time slot, and obtain the entrance flow time slot packet number sequence as The packet size sequence is The exit flow time slot packet number sequence is The packet size sequence is Construct a feature matrix: And Use a variational autoencoder for feature extraction to obtain the entrance flow feature sequence F I , the exit flow feature sequence F E ; The association analysis module is used to obtain the entry flow characteristics and the exit flow characteristics F I and F E Input the association analyzer, and complete the association of the entry traffic and the exit traffic of the Tor anonymous network through a sliding window and a subset sum algorithm based on dynamic programming.

9. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores program code which, when executed by the processor, causes the processor to perform the steps of the method according to any one of claims 1 to 8.

10. A storage medium, characterized in that, It stores a computer program or instructions which, when run on a computer, perform the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Anonymous network traffic identification method and device based on traffic reconstruction and inheritance learning

    CN114615093A

  • Flow association method and system for Tor anonymous network

    CN116319086A