Tor hidden service traffic identification method based on multi-view comparative learning

By employing a multi-view comparative learning method, combined with TLS layer and TCP/IP header feature extraction, cue learning, and text encoder, the problems of insufficient feature extraction and unreasonable perspective fusion in Tor hidden service traffic identification are solved, achieving efficient and accurate traffic identification.

CN120856375APending Publication Date: 2025-10-28SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510883260.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-28
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies for identifying hidden service traffic on Tor suffer from problems such as insufficient feature extraction, unreasonable perspective fusion, and lack of semantic guidance, resulting in poor recognition performance.

Method used

By employing a multi-view comparative learning approach, and constructing TLS layer and TCP/IP header feature extraction modules, combined with cue learning and text encoder, we achieve deep fusion and alignment of traffic and text features, thereby improving recognition accuracy.

Benefits of technology

It significantly improves the accuracy of identifying hidden service traffic on Tor, achieving an accuracy of 99.42% in closed-world scenarios and maintaining a true accuracy of 97.63% in open-world scenarios, demonstrating good robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856375A_ABST
    Figure CN120856375A_ABST
Patent Text Reader

Abstract

The invention designs a Tor hidden service traffic identification method based on multi-view comparative learning, which comprises multi-view-based traffic feature extraction, prompt learning-based text feature extraction and comparative learning-based model training and testing, and is characterized in that firstly, multi-view-based traffic feature extraction and fusion are utilized; the problem of insufficient feature extraction under a single view angle is overcome, and richer information is provided for model training; secondly, by means of the technology in the field of natural language processing, a prompt text is constructed by means of statistical information of flow, and semantic information contained in the prompt text is mined through a text encoder; and finally, spatial alignment of the traffic features and the text features is completed by using a comparative learning framework, cross-view knowledge complementation and collaborative optimization are realized, and the identification performance of the Tor hidden service traffic is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of cyberspace security supervision and governance, and anonymous communication network traffic identification in a big data environment. Specifically, it is applied to the identification of hidden service traffic in anonymous communication networks represented by The Onion Router (Tor), providing a technical foundation for governing anonymous communication networks and ensuring cyberspace security. Background Technology

[0002] The Onion Router (Tor) is an anonymous communication system built using Onion routing technology. It effectively hides user identities and communication relationships by encrypting data in multiple layers and transmitting it among volunteer relay nodes worldwide. Besides providing anonymous browsing of surface-level websites, the Onion Router also supports hiding services, allowing users to access sites that don't appear in regular search engines through Onion domains ending in ".onion". Unfortunately, this advantage of the Onion Router has been exploited by criminals who use Tor hiding services to build various illegal websites for drug trafficking, pornography distribution, and other illegal activities. These hiding services, due to their multi-layered encryption and distributed network transmission technologies, make them extremely difficult to track and locate. Government agencies worldwide consider regulating the Tor Network a top priority for maintaining cybersecurity. To address this challenge, academia and industry currently agree that detecting traffic from Tor hiding services is the most effective way to combat dark web crime.

[0003] Researchers have proposed various analysis methods for traffic accessing hidden services on Tor. Early detection methods relied primarily on obvious features in network traffic, such as domain name detection and server address matching, to determine the websites users visited. However, the Tor protocol, through multiple iterations, has effectively hidden key features of the original traffic through multi-layered encryption and onion routing technology, rendering traditional detection methods based on fixed features ineffective. With the rapid development of artificial intelligence, traffic identification methods based on deep learning have gradually become a core technology for monitoring Tor networks. By training neural networks to automatically learn subtle differences and data distributions in traffic data, accurate classification of traffic access to different websites can be achieved. Even in anonymous network environments, regulators can still leverage the powerful data classification capabilities of artificial intelligence to extract key features from massive amounts of traffic data, thereby accurately identifying the hidden service sites accessed by users. Currently, related research has achieved preliminary results, but existing work has not yet provided an accurate and efficient identification solution. The single-view methods used in existing research, such as using packet direction sequences or time series, are difficult to comprehensively represent the complex traffic of hidden Tor services, easily leading to large fluctuations in classification performance. Multi-perspective recognition methods fuse heterogeneous information from multiple sources, including direction, time, packet length, protocol fields, and byte payloads, extracting rich features through parallel networks or multi-channel structures. However, because their fusion strategies are mostly simple concatenation or averaging, lacking dynamic and adaptive weight allocation of the contributions of each perspective, the recognition performance has not been significantly improved. Therefore, it is necessary to construct an efficient multi-perspective data fusion mechanism to mine the deep semantic relationships hidden in the network. In summary, designing an accurate and efficient multi-perspective recognition method is an urgent problem to be solved for Tor hidden service traffic identification.

[0004] To address the aforementioned problems, this invention proposes a method for identifying Tor hidden service traffic based on multi-view contrastive learning. Compared to previous work, the innovations of this invention are: First, by utilizing multi-view traffic feature extraction and fusion, it overcomes the problem of insufficient feature extraction under a single view, providing richer information for model training; second, by leveraging techniques from the field of natural language processing, it constructs prompt text using traffic statistics, and then mines the semantic information contained therein through a text encoder; finally, it uses a contrastive learning framework to achieve spatial alignment between traffic features and text features, realizing cross-view knowledge complementarity and collaborative optimization, significantly improving the recognition performance of Tor hidden service traffic. Summary of the Invention

[0005] This invention aims to address the problems of insufficient feature extraction, unreasonable perspective fusion, and lack of semantic guidance in existing Tor hidden service traffic identification methods. It provides a Tor hidden service traffic identification method based on multi-view comparative learning, which can achieve accurate and efficient classification of encrypted dark web traffic.

[0006] To achieve the above objectives, this invention comprises three steps: multi-view traffic feature extraction, cue-based text feature extraction, and contrastive learning-based model training and testing. Multi-view traffic feature extraction involves constructing TLS layer input matrices and TCP / IP header input matrices separately, and designing corresponding feature extraction and fusion modules to achieve in-depth mining of traffic features. Cue-based text feature extraction utilizes traffic statistics to construct cue text and employs a text encoder to extract semantic information, achieving full mining of text features. Contrastive learning-based model training and testing aligns traffic features with text features by constructing positive and negative sample pairs and designing a contrastive learning loss function, thereby improving the model's recognition accuracy; specifically as follows...

[0007] S1: Multi-perspective traffic feature extraction. To achieve accurate identification of Tor hidden service traffic, information contained in the raw traffic is extracted from multiple perspectives. This process consists of three steps: TLS layer traffic feature extraction, TCP / IP header feature extraction, and feature fusion module construction. The specific structure is as follows... Figure 1 As shown.

[0008] S11: TLS Layer Traffic Feature Extraction. Features are extracted from the raw traffic data from the perspective of TLS layer communication interactions. This process consists of two steps: constructing the TLS layer input matrix and designing the TLS layer feature extraction module. The specific steps are as follows:

[0009] (1) Constructing the TLS layer input matrix. The payload size and packet direction information are extracted from the TLS layer header to construct the TLS layer input matrix. This process consists of three steps: extracting size information, extracting direction information, and constructing the input matrix.

[0010] a) Extract size information. Extract the size information of the TLS layer payload from the TLS layer header.

[0011] b) Extracting direction information. The sending direction of a data packet is determined based on its source and destination IP addresses. "+1" indicates the packet was sent from the client to the Tor network's entry node, and "-1" indicates the packet was returned from the Tor network's entry node to the client. In this way, the direction of all data packets in a flow is extracted to form a direction sequence.

[0012] c) Construct the input matrix. Organize the payload size information and packet direction information extracted in steps a) and b) into a 2×len time sequence matrix, where len is the number of packets extracted from the stream. If the number of packets in the stream is less than len, fill the empty spaces in the time sequence matrix with 0; if the number of packets in the stream is greater than len, truncate the current stream after the len-th packet and discard subsequent packets.

[0013] (2) Design of the TLS layer feature extraction module: A feature extraction module is constructed using the Temporal Convolutional Network (TCN) basic block to extract TLS layer traffic features. This process consists of two steps: constructing the TCN basic block and constructing the feature extraction module based on the TCN basic block. The specific structure of the module is as follows: Figure 2 As shown.

[0014] a) Constructing the TCN base block. The TCN base block contains two sets of dilated convolutional blocks and one residual connection to improve the model's capture of long-range dependencies and focus on small mappings. Each set of dilated convolutional blocks contains one one-dimensional dilated convolutional layer, one weight normalization layer, one ReLU activation function, and one Dropout layer.

[0015] b) Construct a feature extraction module based on TCN base blocks. The feature extraction module contains a total of 4 TCN base blocks, and the kernel size of the one-dimensional dilated convolutional layer in each TCN base block is tls. ks However, the expansion rates are tls in order. ds 2×tls ds 4×tls ds 8×tls ds The number of channels is tls fn 2×tls fn 4×tls fn 8×tls fn A max-pooling layer is connected after each TCN base block to expand the receptive field and extract higher-dimensional features.

[0016] S12: TCP / IP Header Feature Extraction. Based on the TCP / IP protocol, extract valid byte information from the raw traffic. This process consists of two steps: constructing the TCP / IP header input matrix and designing the TCP / IP header feature extraction module.

[0017] (1) Constructing the TCP / IP header input matrix. A subset of bytes is extracted from the TCP / IP header to construct the input matrix. This process consists of two steps: determining the bytes to extract and constructing the input matrix.

[0018] a) Determine the bytes to extract. Discard some header fields that affect model training, such as IP address and port number. Select fields from the TCP and IP headers that cover a total of 24 bytes to construct the model input, including protocol version, header length, time to live, etc.

[0019] b) Construct the input matrix. Organize the extracted field information into a 24×len time series matrix, where len is the number of packets extracted from a stream. If the number of packets in a stream is less than len, fill the empty spaces in the time series matrix with 0; if the number of packets is greater than len, truncate the current stream after the len-th packet and discard subsequent packets.

[0020] (2) Design of TCP / IP header feature extraction module: A feature extraction module is constructed using Convolutional Neural Network (CNN) basic blocks to extract TCP / IP header features. This process consists of two steps: constructing CNN basic blocks and constructing a feature extraction module based on CNN basic blocks. The specific structure of the module is as follows: Figure 3 As shown.

[0021] a) Constructing the basic CNN block. The basic CNN block contains two sets of one-dimensional convolutional blocks, one max pooling layer, and one Dropout layer. Each one-dimensional convolutional block contains one one-dimensional convolutional layer, one batch normalization layer, and one ReLU activation function.

[0022] b) Construct a feature extraction module based on CNN base blocks. The feature extraction module contains three CNN base blocks, with the batch normalization layer removed from the first CNN base block. The kernel size of the one-dimensional convolutional layer in each CNN base block is tcp. ks However, the number of channels is in the order of TCP. fn 2×tcp fn 4×tcp fn .

[0023] S13: Feature Fusion Module Design. The feature fusion module is designed to fuse the TLS layer traffic feature f1 and TCP / IP header feature f2 extracted in steps S11 and S12. This process specifically consists of three steps: designing the mutual attention layer, designing the information supplementation layer, and designing the alternating overlay structure.

[0024] (1) Design the mutual attention layer. Align and merge the TLS layer traffic feature f1 and the TCP / IP header feature f2. This process consists of three steps: calculating the mutual attention matrix, calculating the weighted features, and calculating the merged features.

[0025] a) Calculate the mutual attention matrix. The mutual attention matrix is ​​obtained by adjusting the features from different perspectives using a linear transformation layer. The mutual attention mam1 and mam2 corresponding to f1 and f2 are obtained according to formula (1). Where W1 is the linear transformation matrix corresponding to the linear transformation layer.

[0026]

[0027] b) Calculate the weighted features. The importance of each viewpoint feature is calculated using a normalized exponential function to generate weighted features. The weighted features att1 and att2 corresponding to mam1 and mam2 are obtained according to formula (2). Where W2 is the linear transformation matrix corresponding to the linear transformation layer.

[0028] att i =softmax(mam) i (f) i W2) (2)

[0029] c) Calculate the fusion features. Using formula (3), the weighted features from each perspective are concatenated and the original information is added to obtain the fusion feature output. Here, concat means concatenating by dimension.

[0030] output=concat(att1,att2)+input (3)

[0031] (2) Design an information supplementation layer. An information supplementation layer is designed after the mutual attention layer to fill in key information from the original modality. The information supplementation layer has two inputs: the output of the mutual attention layer and the TCP / IP header features from the previous layer. This process consists of three steps: calculating the attention matrix, calculating the features of the reassembled TCP / IP header, and calculating the output features.

[0032] a) Calculate the attention matrix. This involves considering the TCP / IP header features from the previous layer. The input is fed into the linear layer, and the TCP / IP header features of the current layer are generated using the linear transformation matrix W3. The attention matrix is ​​calculated using the normalized exponential function according to formula (4). in ReLU is the activation function.

[0033]

[0034] b) Calculate the features of the reassembled TCP / IP header. Based on formula (5), calculate the attention matrix... As input, calculate the reassembled TCP / IP header feature F. n Where W4 is the linear transformation matrix corresponding to the linear transformation layer.

[0035]

[0036] c) Calculate the output features. According to formula (6), calculate the output features F of the reassembled TCP / IP header. n The output of the mutual attention layer is used as input, and a batch normalization layer (BN) and a learnable matrix w are used. i The output feature F is calculated. out .

[0037] F out =BN(output+w) i F n (6)

[0038] (3) Design an alternating overlay structure. By alternating overlay of mutual attention layers and information supplementation layers n times, feature depth fusion between different perspectives is achieved.

[0039] S2: Text Feature Extraction Based on Cue Learning. Cue text is constructed using traffic statistics features, and text features are extracted using a text encoder. This process specifically includes three steps: traffic statistics feature extraction, cue text generation, and feature extraction based on a text encoder.

[0040] S21: Traffic Statistical Feature Extraction. This involves extracting multiple statistical features from the raw traffic data. The process consists of two steps: traffic statistical feature calculation and traffic statistical feature selection.

[0041] (1) Calculation of traffic statistics features. As shown in Table 1, 51 different types of traffic statistics features were selected for Tor hidden service access traffic, including the total number of bidirectional data packets, the total number of bidirectional transmitted bytes, and the data packet arrival interval.

[0042] Table 1. Complete Set of Traffic Characteristics

[0043]

[0044] (2) Flow Statistical Feature Filtering. For the 51 statistical features extracted in step (1), the Max-Relevance and Min-Redundancy (mRMR) algorithm is used to filter out num features. In the mRMR algorithm, feature x... i Score(x) i The calculation process of x is shown in formula (7), where x i Let represent the i-th feature, y represent the label corresponding to this set of features, S represent the feature subset, and I represent the calculation of mutual information.

[0045]

[0046] S22: Prompt Text Generation. To help the model learn the statistical characteristics of traffic, these characteristics are first converted into prompt text describing the traffic flow in natural language. This process consists of two steps: designing a prompt text template and generating the text by applying the template.

[0047] (1) Design prompt text template. Construct a concise and easily understood prompt text template. This process consists of two steps: setting the traffic label template and setting the statistical feature template.

[0048] a) Set the traffic label template. In the traffic label template "This is a traffic flow of {URL}.", set the traffic label field {URL} to indicate that the current flow belongs to the {URL} label.

[0049] b) Set statistical feature templates. For the num features selected in step S21, describe their meaning in words and concatenate their values ​​in the current stream {value}. i}. Concatenate the templates of all num features after the traffic label template to obtain the complete prompt text template.

[0050] (2) Generate text using a template. For each bidirectional flow, calculate statistical features and fill these features into the template designed in step (1) to obtain the complete prompt text. This process is specifically divided into two steps: calculating feature values ​​and filling the template.

[0051] a) Calculate the feature values. For each bidirectional input stream, select num statistical features according to step S21 and calculate the corresponding feature values.

[0052] b) Template filling. The calculated feature values ​​are filled into the template designed in step (1) to obtain the corresponding prompt text.

[0053] S23: Feature Extraction Based on Text Encoder. Construct a BERT-based text encoder to extract features from the prompt text. For example... Figure 4 As shown, this process is divided into four steps: word segmentation preprocessing, BERT encoding, output extraction, and dimension adjustment.

[0054] (1) Word segmentation preprocessing. The prompt text generated in step S22 is preprocessed by word segmentation to construct an input suitable for BERT. This process consists of four steps: word segmentation, adding special tags, vocabulary mapping, and text length processing.

[0055] a) Tokenization. The prompt text is input into the BERT tokenizer and segmented into sub-words to obtain a sub-word sequence.

[0056] b) Add special markers. Add a special marker [CLS] at the beginning of the sub-word sequence for overall semantic representation, and add a marker [SEP] between sentences and at the end of the sub-word sequence to separate text blocks, thus obtaining the complete sub-word sequence.

[0057] c) Vocabulary mapping. Map each subword in the subword sequence to the corresponding tokenID in the vocabulary to obtain the token sequence.

[0058] d) Text length processing. The token sequence is padded and truncated according to a fixed input length l to obtain a fixed-length token sequence, ensuring that different samples can be processed in parallel in the same batch without exceeding the limit.

[0059] (2) BERT Encoding. The token sequence is fed into a network with BERT as its backbone to obtain a text feature vector containing global context information. This process consists of two steps: generating the feature matrix and Transformer encoding.

[0060] a) Generate the feature matrix. The token sequence is embedded sequentially through the embedding layer to complete word embedding, position embedding and paragraph embedding. The results of the three embeddings are added together to obtain the feature matrix M.

[0061] b) Transformer Encoding. The feature matrix M is fed into a multi-layer Transformer encoder to generate an output vector containing global contextual information. In each Transformer encoder layer, contextual information is first aggregated through self-attention, and then the features are refined using a feedforward network. This process is repeated layer by layer to capture more distant contextual relationships.

[0062] (3) Output extraction. After BERT encoding is completed, the output vector corresponding to the [CLS] token at the beginning of the token sequence is extracted as the summary representation of the entire text and output to the next layer for dimensional adjustment.

[0063] (4) Dimension Adjustment. After the output is extracted, a fully connected layer is set up to adjust the dimensions, ensuring that the text and traffic features are comparable in the same metric space, so as to improve the effect of cross-view alignment and contrastive learning.

[0064] S3: Model Training and Testing Based on Contrastive Learning. The cosine similarity between the output vectors of steps S1 and S2 is calculated using contrastive learning, and this similarity is used to train and test the model. This process consists of three steps: designing the contrastive learning framework, training the model, and testing with samples.

[0065] S31: Design a comparative learning framework. For example... Figure 5As shown, the output vectors of the two feature extraction modules designed in steps S1 and S2 are fed into the contrastive learning framework, and the cosine similarity is calculated to measure the relationship between them. This process is specifically divided into two steps: constructing positive and negative sample pairs and designing a loss function.

[0066] (1) Construct positive and negative sample pairs. The flow characteristics TF in the same flow sample i are... i and text features TE i As a positive sample pair, i.e. (TF) i ,TE i ), i∈{1,2,...,N}. Transform the TF of sample i. i Text features TE compared to other traffic samples j j As a negative sample pair, i.e. (TF) i ,TE j ), i≠j, i,j∈{1,2,...,N}. Where N is the number of traffic samples sent to the model for training in each batch.

[0067] (2) Design the loss function. Based on formula (8), a bidirectional information maximization loss function is designed to achieve bidirectional comparison, allowing the model to learn more effective features from both perspectives. Here, sim() is used to calculate the cosine similarity between two vectors, and sim(TF) is used to calculate the cosine similarity between two vectors. i ,TE i ) represents the similarity between the traffic feature and its corresponding text feature in the i-th positive sample, sim(TF) i ,TF j ) represents the similarity between traffic features and text features in a negative sample pair constructed from sample i and sample j. τ is a temperature coefficient used to control the discriminative power between positive and negative samples. exp represents an exponential operation with base e.

[0068]

[0069] S32: Training the model. The model constructed in steps S1 and S2 is trained using the training set data. This process consists of four steps: hyperparameter selection, weight initialization, batch training of the model, and storing the model weights.

[0070] (1) Hyperparameter selection. Based on the definition and search range of hyperparameters described in Table 2, the values ​​are determined using a hyperparameter random selection algorithm.

[0071] (2) Weight Initialization. A random weight initialization algorithm is used to initialize all trainable weights and biases of the model designed in steps S1 and S2, including W1, W2, W3, W4, ... wait.

[0072] (3) Train the model in batches. Train the model in batches according to the size N determined in step (1). Repeat the training epochs on the training set until the model loss L stabilizes and no longer decreases.

[0073] (4) Store model weights. Store the weight file of the model trained in step (3).

[0074] Table 2 Hyperparameter Definitions

[0075]

[0076] S33: Test Samples. The proposed method is tested on the test set. This process consists of four steps: model weight invocation, test metric selection, model sample inference, and test metric calculation.

[0077] (1) Model weight call. Call the model weight file stored in step S32, initialize the model structure and load the weights.

[0078] (2) Selection of Test Metrics. Test metrics were selected for Tor hidden service traffic identification in two scenarios: closed world and open world. For the closed world scenario, accuracy (ACC) was selected, as shown in formula (9); for the open world scenario, true positive rate (TPR) and false positive rate (FPR) were selected, as shown in formulas (10) and (11). TP represents true positives, which is the number of samples that are actually positive correctly identified as positive by the model; TN represents true negatives, which is the number of samples that are actually negative correctly identified as negative by the model; FP represents false positives, which is the number of samples that are actually negative incorrectly identified as positive by the model; and FN represents false negatives, which is the number of samples that are actually positive incorrectly identified as negative by the model.

[0079]

[0080] (3) Model Sample Inference. Inference is performed on the model using samples from the test set. This process consists of five steps: traffic data processing, setting label text templates, generating label text feature vectors, similarity determination, and determining sample classification labels. The specific model is as follows: Figure 6 As shown.

[0081] a) Traffic data processing. Information is extracted from the raw traffic data to construct the TLS layer input matrix and the TCP / IP header input matrix, which are then input into the trained model to generate feature vectors (TF).

[0082] b) Set the tag text template. Use the tag template "This is a traffic flow of {URL}." to generate corresponding text descriptions for all tags.

[0083] c) Generate labeled text feature vectors. Input all the text descriptions obtained in step b) into the trained model and convert them into a set of text feature vectors TE1,…,TE2. K , where K is the number of tags.

[0084] d) Similarity determination. Calculate the similarity between the traffic feature vector TF and the feature vectors of all K tag texts.

[0085] e) Sample classification label determination. In a closed-world scenario, the label corresponding to the label text feature vector with the highest similarity is selected as the classification result for the sample. In an open-world scenario, the label corresponding to the label text feature vector with the highest similarity is selected, and its confidence score is further checked to determine whether it belongs to a hidden service or a surface network.

[0086] (4) Calculation of test indicators. Compare the traffic classification labels of the test samples with their actual labels in both closed and open world scenarios, and calculate the indicators selected in step (2).

[0087] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the Tor hidden service traffic identification method based on multi-view contrastive learning.

[0088] A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the Tor hidden service traffic identification method based on multi-view contrastive learning.

[0089] Compared with the prior art, the advantages of the present invention are as follows:

[0090] This invention proposes a method for identifying Tor hidden service traffic based on multi-view contrastive analysis. Compared with other existing methods for identifying Tor hidden service traffic, the advantages of this method are as follows:

[0091] (1) This method realizes traffic feature extraction based on multiple perspectives. By introducing multi-perspective data input, it effectively improves the comprehensiveness of data representation and the model's discriminative ability. First, in steps S11 and S12, the input matrices under the two perspectives are constructed, and corresponding feature extraction modules are designed to fully explore the potential key information in the traffic data under each perspective. Second, through step S13, a feature fusion module is designed to achieve deep fusion of multi-perspective features, which not only retains the complementary information of each perspective but also avoids information redundancy.

[0092] (2) This method implements text feature extraction based on cue learning. By constructing cue text, a text encoder is used to extract text features from the traffic, thereby improving the model's discriminative ability. First, in step S21, a complete set of features that can effectively represent Tor hidden service access traffic is selected. The mRMR algorithm is used to filter the features, ensuring the diversity and complementarity of the selected features in the semantic space. Second, in step S22, a general cue text template is designed by combining traffic labels and feature definitions, enabling the model to capture key semantic information from the cue text. Finally, in step S23, a BERT-based text encoder is designed to extract text features from Tor hidden service access traffic.

[0093] (3) This method achieves Tor hidden service traffic identification based on contrastive learning. By calculating the cosine similarity between the traffic feature vector of the test sample and the text feature vector of the known category, the classification performance is improved. First, in step S31, a contrastive learning framework is designed to construct positive and negative sample pairs to deeply align traffic data features from different perspectives. Second, in step S31, a bidirectional information maximization loss function is designed to achieve the comparison of the "traffic → text" and "text → traffic" dual paths, effectively reducing the information offset under a single path. This enables the model to achieve an accuracy of 99.42% and a true accuracy of 97.63% in closed-world and open-world scenarios, respectively, which is superior to existing state-of-the-art methods. As the number of surface site point samples increases from 1000 to 6000, the true accuracy of the model in the open world can be maintained above 94.3%, demonstrating good robustness. Attached Figure Description

[0094] Figure 1 Schematic diagram of the traffic feature extraction module based on multiple perspectives;

[0095] Figure 2 Schematic diagram of the TLS layer feature extraction module;

[0096] Figure 3 Schematic diagram of the TCP / IP header feature extraction module;

[0097] Figure 4 A schematic diagram of a text feature extraction module based on cue-based learning;

[0098] Figure 5 A schematic diagram of a Tor hidden service traffic identification framework based on contrastive learning;

[0099] Figure 6 Schematic diagram of the model sample testing process;

[0100] Figure 7 Hyperparameter selection flowchart;

[0101] Figure 8 Model training flowchart. Detailed Implementation

[0102] The technical solutions in the embodiments will be described in detail below with reference to the accompanying drawings. Obviously, the embodiments described below are merely one embodiment of the method of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the following embodiments without creative effort are within the scope of protection of the present invention.

[0103] Example 1: A method for identifying Tor hidden service traffic based on multi-view comparative learning. Figure 5 This embodiment demonstrates the overall framework for multi-view Tor hidden service traffic identification. The implementation process of this invention consists of three main steps: multi-view traffic feature extraction, cue-based text feature extraction, and contrastive learning-based model training and testing.

[0104] S1: Multi-perspective traffic feature extraction. In this embodiment, the specific process consists of three steps: TLS layer traffic feature extraction, TCP / IP header feature extraction, and feature fusion module construction. This embodiment utilizes the PyTorch deep learning framework to construct the traffic feature extraction module. By inheriting the nn.Module module class, and defining the TLS layer traffic feature extraction module, TCP / IP header feature extraction module, and feature fusion module sequentially according to the specific module design in step S1, the multi-perspective traffic feature extraction module designed in step S1 is finally realized.

[0105] S11: TLS layer traffic feature extraction. For example... Figure 2 As shown, in this embodiment, the specific process is divided into two steps: constructing the TLS layer input matrix and designing the TLS layer feature extraction module.

[0106] (1) Constructing the TLS layer input matrix. In this embodiment, the specific process consists of three steps: extracting size information, extracting direction information, and constructing the input matrix.

[0107] a) Extract size information. Install Wireshark software on the local host and the pyshark toolkit in the Python environment. Call the get_field_by_showname("Length") function to extract the size information of the TLS layer payload from the TLS header. Divide the size information by 530 and round up. Set all values ​​greater than 10 to 10. Finally, divide all values ​​by 10 and keep the decimal part to normalize the size information to the [0,1] range.

[0108] b) Extracting direction information. Use the pyshark tool to obtain the src attribute of TLS packets and extract the source IP address information from the header. If the source IP address is equal to the Tor proxy node IP, use "-1" to indicate the direction of the packet; otherwise, use "+1" to indicate the direction of the packet. Extract the direction of each packet in sequence to form a direction sequence, which serves as the direction information of the traffic.

[0109] c) Construct the input matrix. Set the length len of the matrix to 1400, and organize the size and direction information extracted in steps a) and b) into a timing matrix of shape 2×len. Fill the empty spaces in the timing matrix with 0 for portions where the number of data packets is less than 1400, and truncate and discard portions where the number of data packets is greater than 1400.

[0110] (2) Design the TLS layer feature extraction module. In this embodiment, the specific process is divided into two steps: constructing the TCN basic block and constructing the feature extraction module based on the TCN basic block.

[0111] a) Construct the TCN base block. Using the temporal matrix generated in step (1) as input, perform convolution on it using the nn.Conv1d function, then normalize the weights using a WeightNorm layer, and then use nn.Relu to import the Relu function for non-linear transformation of the data. Finally, call the nn.Dropout function to perform a Dropout operation, randomly deactivating some hidden layer neurons with Dropout = 0.2. Repeat the above process once, and finally output the data to the next TCN base block.

[0112] b) Construct a feature extraction module based on TCN base blocks. This embodiment uses the nn.Module class, which inherits from four TCN base blocks, and sets the kernel size of each TCN base block to tls. ks Set all to 4, and set tls fn Set to 32, the number of channels corresponding to the 4 TCN base blocks are [32, 64, 128, 256], and tls ds The expansion rates for the four TCN base blocks are set to 1, and are [1, 2, 4, 8]. After each TCN base block, the nn.MaxPooling1D function is called for max pooling. After the data has been processed by four TCN base blocks, the output is the TLS layer traffic feature f1.

[0113] S12: TCP / IP header feature extraction. In this embodiment, the specific process is divided into two steps: constructing the TCP / IP header input matrix and designing the TCP / IP header feature extraction module.

[0114] (1) Constructing the TCP / IP header input matrix. In this embodiment, the specific process consists of two steps: determining the extracted bytes and constructing the input matrix.

[0115] a) Determine the bytes to extract. Use the pyshark tool to select a total of 24 bytes from the IP and TCP headers as input to the model, and divide all bytes by 256 to map them to the [0,1] range.

[0116] b) Construct the input matrix. Using len equal to 1400, organize the field information extracted in step a) into a time series matrix of shape 24×1400. Pad the streams with less than 1400 data packets with 0, and truncate the streams with more than 1400 data packets. Finally, input the matrix into the subsequent extraction module.

[0117] (2) Design a TCP / IP header feature extraction module. The specific results of the module are as follows: Figure 3 As shown. In this embodiment, the specific process consists of two steps: constructing the CNN basic block and constructing a feature extraction module based on the CNN basic block.

[0118] a) Constructing the CNN base block. Using the temporal matrix generated in step (1) as input, the `nn.Conv1d` function is called to perform convolution on the matrix. Then, the `BatchNormalization` function is used for normalization, making the mean of the input to each layer 0 and the variance 1. Then, `nn.Relu` is called to import the Relu function to perform non-linear transformations on the data. This process is repeated once. Finally, the `nn.MaxPooling1D` function is called to perform dimensionality reduction on the data. Finally, the `nn.Dropout` function is called to perform a Dropout operation, randomly deactivating some hidden layer neurons with a Dropout value of 0.1 before outputting the data to the next CNN base block.

[0119] b) Construct a feature extraction module based on CNN basic blocks. Use the `nn.Module` class, which inherits from three CNN basic blocks, and remove the batch normalization layer from the first CNN basic block. Use TCP... ks The value is equal to 6 as the kernel size for each CNN basic block, and TCP... fn The value is set to 32, resulting in channel sizes of [32, 64, 128] for the three CNN base blocks. After processing by the three CNN base blocks, the output is the TCP / IP header feature f2.

[0120] S13: Feature Fusion Module Design. The TLS layer traffic feature f1 and TCP / IP header feature f2 extracted in steps S11 and S12 are fused together using multi-perspective features. In this embodiment, the specific process consists of three steps: designing a mutual attention layer, designing an information supplementation layer, and designing an alternating overlay structure.

[0121] (1) Design the mutual attention layer. In this embodiment, the specific process is divided into three steps: calculating the mutual attention matrix, calculating the weighted features, and calculating the fusion features.

[0122] a) Calculate the mutual attention matrices. Calculate the mutual attention matrices mam1 and mam2 by applying the linear transformation matrix W1 to f1 and f2.

[0123] b) Calculate weighted features. Call the softmax function to calculate the importance of each viewpoint feature and generate weighted features. Then, use the linear transformation matrix W2 to obtain the weighted features att1 and att2 corresponding to mam1 and mam2.

[0124] c) Calculate the fused features. Use the concat function to concatenate att1 and att2, and add the original information input to obtain the fused feature output.

[0125] (2) Design of the information supplementation layer. In this embodiment, the specific process is divided into three steps: calculating the attention matrix, calculating the reassembled TCP / IP header features, and calculating the output features.

[0126] a) Calculate the attention matrix. The input is fed into a linear layer and generated using the linear transformation matrix W3. Finally, according to formula (4), the softmax function and ReLU function are used to... The attention matrix is ​​obtained by calculating the output.

[0127] b) Calculate the characteristics of the reassembled TCP / IP header. Use the ReLU function and the linear transformation matrix W4 to... and Calculations are performed to generate the reassembled TCP / IP header feature F. n .

[0128] c) Calculate the output features. Use the BatchNormalization function to evaluate F. n The output of the mutual attention layer is processed using the learnable matrix w. i The output feature F is calculated. out .

[0129] (3) Design an alternating overlay structure. In this embodiment, three mutual attention layers and information supplementation layers are alternately overlaid. After data processing, the output is a flow feature vector. The specific structure is as follows: Figure 1 As shown.

[0130] S2: Text feature extraction based on cue learning. In this embodiment, the specific process includes three steps: traffic statistics feature extraction, cue text generation, and feature extraction based on a text encoder.

[0131] S21: Traffic Statistical Feature Extraction. In this embodiment, the specific process consists of two steps: traffic statistical feature calculation and traffic statistical feature filtering.

[0132] (1) Calculation of traffic statistics features. Using Wireshark software and pyshark toolkit, a total of 51 traffic statistics features were extracted as shown in Table 1, which were used as input for the next step.

[0133] (2) Flow statistics feature selection. The mRMR algorithm was used to select the 51 statistical features extracted in step (1). Finally, the 12 statistical features with the highest scores were selected as the model input.

[0134] S22: Prompt Text Generation. In this embodiment, the process consists of two steps: designing a prompt text template and applying the template to generate the text.

[0135] (1) Designing a prompt text template. This embodiment constructs a concise and easily understood prompt text template, which is divided into two steps: setting a traffic label template and setting a statistical feature template.

[0136] a) Set traffic label templates. Design label prompt text templates according to the labels to which different types of traffic belong, such as "This is a traffic flow of {Google}." to indicate that the traffic comes from Google.

[0137] b) Set up statistical feature templates. Design feature prompt text templates based on the specific meaning of statistical features, such as "The average packet size is {} bytes.", where the curly braces {} can be filled with the average packet size statistical feature.

[0138] (2) Generate text using a template. In this embodiment, the process consists of two steps: calculating feature values ​​and filling the template.

[0139] a) Calculate feature values. For each input stream, select 12 statistical features according to step S21 and calculate the corresponding feature values.

[0140] b) Template Filling. The calculated feature values ​​are filled into the prompt template according to the rules designed in step (1). Each flow corresponds to one traffic label prompt text and 12 statistical feature prompt texts. For example, "This is a trafficflow of {Google}." indicates that the traffic comes from Google, while "The average packet size is {pktLenAvg} bytes." indicates the average packet size of the traffic, and "{pktPerSecAvg} packets are transmitted per second." indicates the average number of packets per second of the traffic. The traffic label prompt text and the 12 statistical feature prompt texts are concatenated to obtain the complete traffic prompt text, which is then sent to the model.

[0141] S23: Feature extraction based on a text encoder. This embodiment selects the BERT-base model based on the transformers library to construct a text encoder, the specific structure of which is as follows: Figure 4 As shown, it mainly includes four steps: word segmentation preprocessing, BERT encoding, output extraction, and dimension adjustment.

[0142] (1) Word segmentation preprocessing. In this embodiment, the specific process consists of four steps: word segmentation, adding special tags, vocabulary mapping, and text length processing.

[0143] a) Tokenization. The BertTokenizer class from the transformers library is called to segment the prompt text generated in step S22 into a sequence of subwords that conform to the vocabulary of the BERT-base model.

[0144] b) Add a special token. Call the `build_inputs_with_special_tokens()` function in the `BertTokenizer` class to add `[CLS]` to the beginning of the sequence.

[0145] Used for overall semantic representation, add [SEP] delimiters between sentences and at the end of the text.

[0146] c) Vocabulary mapping. Use the convert_tokens_to_ids() function in the BertTokenizer class.

[0147] The function maps each subword in the word sequence obtained in step a) to the corresponding token ID in the vocabulary, thus obtaining a token sequence.

[0148] d) Text length processing. Using a fixed length l equal to 256, the encode_plus() function in the BertTokenizer class is used to pad and truncate the sequence to obtain the final token sequence.

[0149] (2) BERT encoding. In this embodiment, the specific process consists of two steps: generating the feature matrix and Transformer encoding.

[0150] a) Generate the feature matrix. Input the token sequence obtained in step (1) into the BERT embedding layer to complete word embedding, position embedding and paragraph embedding respectively, and add them together to obtain the feature matrix M.

[0151] b) Transformer Encoding. The feature matrix M is fed into a 12-layer Transformer encoder for iterative processing. Each token, after being processed by all Transformer layers, will obtain a representation containing global context information.

[0152] (3) Output extraction. The [CLS] marker at the beginning of the token sequence represents the semantic information of the entire input sequence.

[0153] (4) Dimension adjustment. The nn.Linear function is called to introduce a fully connected layer to reduce the dimension of the original text vector to 512, and the final output is the text feature vector.

[0154] S3: Model Training and Testing Based on Contrastive Learning. In this embodiment, the specific process consists of three steps: designing a contrastive learning framework, training the model, and testing samples.

[0155] S31: Design the contrastive loss function. In this embodiment, this process is specifically divided into two steps: constructing positive and negative sample pairs and designing the loss function.

[0156] (1) Construct positive and negative sample pairs. The flow characteristics TF in the same sample i are... i and text features TE i As a positive sample pair, i.e. (TF) i TE i ), i∈{1,2,...,n}. Transform the TF of sample i. i Text features TE of other traffic samples j j As a negative sample pair, i.e. (TF) i TE j ), i≠j,i,j∈{1,2,...,n}.

[0157] (2) Design the loss function. For each batch of 64 samples, the similarity between "traffic → text" and "text → traffic" is calculated simultaneously. The traffic feature vector and text feature vector output from steps S1 and S2 are input into the contrastive learning module, and the loss function L is used to maximize the similarity of positive sample pairs while minimizing the similarity of negative sample pairs.

[0158] S32: Training the model. The `data.Data Loader` function is called to load the training, validation, and test sets. The model built in steps S1 and S2 is trained using the training data. In this example, the training process is as follows: Figure 8 As shown, the process consists of the following four steps: hyperparameter selection, weight initialization, batch training of the model, and storage of model weights.

[0159] (1) Hyperparameter selection. A hyperparameter random selection algorithm is implemented using scikit-learn, with accuracy (ACC) as the benchmark metric. When accuracy no longer increases, all hyperparameter values ​​are determined. The specific process is as follows: Figure 7 As shown in Table 3, the hyperparameter values ​​determined in this embodiment are as follows.

[0160] (2) Weight initialization. Call the Xavier weight initialization function to initialize all trainable weights and biases from steps S1 and S2 according to a uniform distribution.

[0161] (3) Train the model in batches. Train the model in batches of 64 samples each, and repeat the training for 40 rounds until the model loss L stabilizes and no longer decreases.

[0162] (4) Store model weights. After the model training is complete, call the torch.save(model, savepath) function to store the weight file, where the parameter model is the trained model and savepath is the target storage address.

[0163] Table 3 shows the values ​​of the hyperparameters in this embodiment.

[0164]

[0165] S33: Test Samples. In this embodiment, metric testing is performed on the test set. This process specifically consists of four steps: model weight invocation, test metric selection, model sample inference, and test metric calculation.

[0166] (1) Model weights are called. The model.load_state_dict(torch,load(savepath)) function is used to call the saved weights file and load the model weights saved in the savepath path into the test model.

[0167] (2) Selection of test metrics. This embodiment evaluates the final performance of the embodiment from two dimensions: closed-world scenarios and open-world scenarios. In the closed-world scenario, the classification performance of this embodiment is evaluated by accuracy (ACC), while in the open-world scenario, the classification performance is evaluated by true positive rate (TPR) and false positive rate (FPR).

[0168] (3) Model Sample Inference. In this embodiment, the specific process consists of five steps: traffic data processing, setting label text templates, generating label text feature vectors, similarity determination, and sample classification label determination. The specific structure is as follows: Figure 6 As shown.

[0169] a) Traffic data processing. Process the data according to steps S11 and S12 to obtain f1 and f2, and then input them into the trained model to obtain the feature vector TF.

[0170] b) Set the label text template. Use the label template "This is a traffic flow of {URL}." to generate the corresponding text description for all traffic.

[0171] c) Generate label text feature vectors. All the prompt texts obtained in step b) are fed into the trained model and transformed into a set of text feature vectors TE1,…,TE2. 50 .

[0172] d) Similarity determination. Calculate TF and TE1,…,TE 50 The similarity between them.

[0173] e) Sample classification label determination. The label corresponding to the label text feature vector with the highest similarity is selected as the classification result of the sample. In the open-world scenario, labels with a confidence threshold greater than 0.35 are further classified as hidden services, and those with a confidence threshold less than 0.35 are classified as surface networks.

[0174] (4) Calculation of test metrics. After testing, the model in the closed-world scenario achieved an accuracy of 99.42% with 10 labels and maintained an accuracy of 97.63% with 50 labels. In the open-world scenario, as the sample size increased from 1000 to 6000, the TPR decreased from 97.63% to 94.3%, and the FPR decreased from 0.89% to 0.22%. The model consistently maintained a high TPR and a low FPR. These metrics fully demonstrate that the model in this embodiment can operate efficiently in a real network environment, exhibiting high robustness and adaptability to complex and diverse traffic characteristics.

[0175] The above description is merely a preferred embodiment of the method of the present invention. The present invention is not limited to the above embodiments. It should be noted that any equivalent substitutions or obvious modifications made by those skilled in the art under the guidance of this specification fall within the scope of this specification and should be protected by the present invention.

Claims

1. A method for identifying Tor hidden service traffic based on multi-view comparative learning, characterized in that, The method includes the following steps: S1: Traffic feature extraction based on multiple perspectives S2: Text feature extraction based on cue learning S3: Model training and testing based on contrastive learning; S1 includes S11: TLS layer traffic feature extraction S12: TCP / IP header feature extraction S13: Feature fusion module design; S2 includes S21: Flow statistics feature extraction S22: Prompt text generation, S23: Feature extraction based on text encoder; S3 includes S31: Design a comparative learning framework S32: Training the model. S33: Test sample.

2. The Tor hidden service traffic identification method based on multi-view comparative learning according to claim 1, characterized in that, S11: TLS layer traffic feature extraction, as detailed below. From the perspective of TLS layer communication interaction, features are extracted from the raw traffic data in two steps: constructing the TLS layer input matrix and designing the TLS layer feature extraction module. The specific steps are as follows: (1) Construct the TLS layer input matrix. Extract the payload size and packet direction information from the TLS layer header to construct the TLS layer input matrix. This involves three steps: extracting size information, extracting direction information, and constructing the input matrix. (2) Design of TLS layer feature extraction module: The feature extraction module is constructed using the Temporal Convolutional Network (TCN) basic block to extract TLS layer traffic features. Specifically, it consists of two steps: constructing the TCN basic block and constructing the feature extraction module based on the TCN basic block. a) Construct the TCN base block. The TCN base block contains two sets of dilated convolutional blocks and one residual connection to improve the model's capture of long-range dependencies and focus on small mappings. Each set of dilated convolutional blocks contains one one-dimensional dilated convolutional layer, one weight normalization layer, one ReLU activation function, and one Dropout layer. b) Construct a feature extraction module based on TCN base blocks. The feature extraction module contains a total of 4 TCN base blocks. The kernel size of the one-dimensional dilated convolutional layer in each TCN base block is tls. ks The expansion rates are tls ds 2×tls ds 4×tls ds 8×tls ds The number of channels is tls fn 2×tls fn 4×tls fn 8×tls fn A max-pooling layer is connected after each TCN base block to expand the receptive field and extract higher-dimensional features; S12: TCP / IP header feature extraction. Based on the TCP / IP protocol, extract valid byte information from the raw traffic. This is divided into two steps: constructing the TCP / IP header input matrix and designing the TCP / IP header feature extraction module. (1) Construct the TCP / IP header input matrix. Extract some bytes from the TCP / IP header to construct the input matrix. This involves two steps: determining the bytes to be extracted and constructing the input matrix. (2) Design TCP / IP header feature extraction module: Use the basic blocks of Convolutional Neural Network (CNN) to build a feature extraction module to extract TCP / IP header features. Specifically, it is divided into two steps: building the basic blocks of CNN and building the feature extraction module based on the basic blocks of CNN.

3. The Tor hidden service traffic identification method based on multi-view comparative learning according to claim 2, characterized in that, S13: Feature Fusion Module Design. The feature fusion module integrates the TLS layer traffic feature f1 and TCP / IP header feature f2 extracted in steps S11 and S12. This is achieved through three steps: designing the mutual attention layer, designing the information supplementation layer, and designing the alternating overlay structure. (1) Design a mutual attention layer to align and fuse the TLS layer traffic feature f1 and the TCP / IP header feature f2. This involves three steps: calculating the mutual attention matrix, calculating the weighted features, and calculating the fused features. a) Calculate the mutual attention matrix. Adjust the features from different perspectives using a linear transformation layer to obtain the mutual attention matrix. According to formula (1), obtain the mutual attention mam1 and mam2 corresponding to f1 and f2, where W1 is the linear transformation matrix corresponding to the linear transformation layer. mam i =f i T W1f j (i≠j) (1) b) Calculate the weighted features. Use the normalized exponential function to calculate the importance of each view feature and generate weighted features. According to formula (2), obtain the weighted features att1 and att2 corresponding to mam1 and mam2, where W2 is the linear transformation matrix corresponding to the linear transformation layer. to i =softmax(mam i )(f i W2) (2) c) Calculate the fusion feature. Using formula (3), the weighted features from each perspective are concatenated and the original information is added to obtain the fusion feature output, where concat means concatenating by dimension. output=concat(att1,att2)+input (3) (2) Design an information supplementation layer. After the mutual attention layer, design an information supplementation layer to fill in the key information of the original modality. The information supplementation layer has two inputs: the output of the mutual attention layer and the TCP / IP header features of the upper layer. Specifically, it consists of three steps: calculating the attention matrix, calculating the features of the reassembled TCP / IP header, and calculating the output features. a) Calculate the attention matrix, taking the TCP / IP header features from the previous layer... The input is fed into the linear layer, and the TCP / IP header features of the current layer are generated using the linear transformation matrix W3. The attention matrix is ​​calculated using the normalized exponential function according to formula (4). in ReLU is an activation function. b) Calculate the features of the reassembled TCP / IP header, and apply the attention matrix according to formula (5). As input, calculate the reassembled TCP / IP header feature F. n Where W4 is the linear transformation matrix corresponding to the linear transformation layer. c) Calculate the output features. According to formula (6), calculate the reassembled TCP / IP header features F. n The output of the mutual attention layer is used as input, and a batch normalization layer (BN) and a learnable matrix w are used. i The output feature F is calculated. out , F out =BN(output+w i F n ) (6) (3) Design an alternating superposition structure, and achieve feature deep fusion between different perspectives by alternating superposition of mutual attention layer and information supplementation layer n times.

4. The Tor hidden service traffic identification method based on multi-view comparative learning according to claim 3, characterized in that, S21: Traffic Flow Statistical Feature Extraction. This involves extracting multiple statistical features from the raw traffic flow data, and consists of two steps: traffic flow statistical feature calculation and traffic flow statistical feature filtering. (1) Traffic statistics feature calculation: For Tor hidden service access traffic, 51 different types of traffic statistics features were selected, including the total number of bidirectional data packets, the total number of bidirectional transmitted bytes, and the data packet arrival interval. (2) Flow statistics feature selection: For the 51 statistical features extracted in step (1), the Max-Relevance and Min-Redundancy (mRMR) algorithm is used to select num features. In the mRMR algorithm, feature x i Score(x) i The calculation process of x is shown in formula (7), where x i Let represent the i-th feature, y represent the label corresponding to this feature set, S represent the feature subset, and I represent the mutual information calculation.

5. The Tor hidden service traffic identification method based on multi-view comparative learning according to claim 4, characterized in that, S22: Prompt text generation, which consists of two steps: designing a prompt text template and generating text by applying the template. (1) Design prompt text templates. Construct concise prompt text templates that are easy for the model to understand. This involves two steps: setting traffic label templates and setting statistical feature templates. (2) Apply the template to generate text. For each bidirectional flow, calculate the statistical features and fill these features into the template designed in step (1) to obtain the complete prompt text. Specifically, it is divided into two steps: calculate the feature value and fill the template.

6. The Tor hidden service traffic identification method based on multi-view comparative learning according to claim 5, characterized in that, S23: Feature extraction based on text encoder. Construct a text encoder based on BERT to extract features from the prompt text. Specifically, it is divided into 4 steps: word segmentation preprocessing, BERT encoding, output extraction, and dimension adjustment. (1) Word segmentation preprocessing: Perform word segmentation preprocessing on the prompt text generated in step S22 to construct an input suitable for BERT. Specifically, it is divided into 4 steps: word segmentation, adding special tags, vocabulary mapping, and text length processing. a) Word segmentation: Input the prompt text into the BERT word segmenter and segment it into sub-words to obtain a sub-word sequence. b) Add special markers: Add the special marker [CLS] at the beginning of the sub-word sequence for overall semantic representation, and add the marker [SEP] between sentences and at the end of the sub-word sequence to separate text blocks, thus obtaining the complete sub-word sequence. c) Vocabulary mapping: Map each subword in the subword sequence to its corresponding token ID in the vocabulary, resulting in a token sequence. d) Text length processing: The token sequence is padded and truncated according to the fixed input length l to obtain a fixed-length token sequence, so as to ensure that different samples can be processed in parallel in the same batch without going out of bounds. (2) BERT encoding: The token sequence is fed into the network with BERT as the backbone to obtain a text feature vector containing global context information. Specifically, it is divided into two steps: generating feature matrix and Transformer encoding. a) Generate the feature matrix. The token sequence is sequentially embedded through the embedding layer to perform word embedding, position embedding, and paragraph embedding. The results of the three embeddings are added together to obtain the feature matrix M. b) Transformer encoding: The feature matrix M is fed into a multi-layer Transformer encoder to generate an output vector containing global context information. In each layer of the Transformer encoder, the context information is first aggregated through self-attention, and then the features are refined using a feedforward network. The process is deepened layer by layer to capture more distant contextual relationships. (3) Output extraction: After BERT encoding is completed, the output vector corresponding to the [CLS] tag at the beginning of the token sequence is extracted as a summary representation of the entire text and output to the next layer for dimensional adjustment. (4) Dimension adjustment: After the output is extracted, a fully connected layer is set up to adjust the dimensions to ensure that the text and traffic features are comparable in the same metric space, so as to improve the effect of cross-view alignment and comparative learning.

7. The Tor hidden service traffic identification method based on multi-view comparative learning according to claim 6, characterized in that, S31: Design a contrastive learning framework. The output vectors of the two feature extraction modules designed in steps S1 and S2 are fed into the contrastive learning framework. Cosine similarity is calculated to measure the relationship between the two. This involves two steps: constructing positive and negative sample pairs and designing a loss function. (1) Construct positive and negative sample pairs, and combine the flow features TF in the same flow sample i. i and text features TE i As a positive sample pair, i.e. (TF) i ,TE i Given that i ∈ {1, 2, ..., N}, set the TF of sample i. i Text features TE compared to other traffic samples j j As a negative sample pair, i.e. (TF) i ,TE j ), i≠j, i,j∈{1,2,...,N}, where N is the number of traffic samples sent to the model for training in each batch. (2) Design the loss function. According to formula (8), design a bidirectional information maximization loss function to achieve bidirectional comparison, so that the model can learn more effective features from both perspectives. Here, sim() is used to calculate the cosine similarity between two vectors, and sim(TF) is used to calculate the cosine similarity between two vectors. i ,TE i ) represents the similarity between the traffic feature and its corresponding text feature in the i-th positive sample, sim(TF) i ,TE j ) represents the similarity between traffic features and text features in a negative sample pair constructed from sample i and sample j, τ is a temperature coefficient used to control the discriminative power between positive and negative samples, and exp represents an exponential operation with base e.

8. The Tor hidden service traffic identification method based on multi-view contrastive learning according to claim 7, characterized in that, S32: Training the model. The model constructed in steps S1 and S2 is trained using the training set data. Specifically, it consists of four steps: hyperparameter selection, weight initialization, batch training of the model, and storing model weights. S33: Test Samples. The proposed method is tested on the test set, which involves four steps: model weight invocation, test metric selection, model sample inference, and test metric calculation. (1) Model weight call: The model weight file stored in step S32 is called to initialize the model structure and load the weights. (2) Test metrics selection: Test metrics were selected for Tor hidden service traffic identification from two scenarios: closed world and open world. For the closed world scenario, accuracy (ACC) was selected, as shown in formula (9); for the open world scenario, true positive rate (TPR) and false positive rate (FPR) were selected, as shown in formulas (10) and (11). TP represents true positives, which is the number of samples that the model correctly identifies as positive; TN represents true negatives, which is the number of samples that the model correctly identifies as negative; FP represents false positives, which is the number of samples that the model incorrectly identifies as positive; and FN represents false negatives, which is the number of samples that the model incorrectly identifies as negative. (3) Model sample inference: Inference is performed on the model using samples from the test set. This involves five steps: traffic data processing, setting label text templates, generating label text feature vectors, similarity determination, and determining sample classification labels. (4) Calculate the test index. Compare the traffic classification labels of the test samples with their actual labels in both closed and open world scenarios, and calculate the index selected in step (2).

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the Tor hidden service traffic identification method based on multi-view contrastive learning as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, the computer instructions implement the Tor hidden service traffic identification method based on multi-view contrastive learning as described in any one of claims 1-8.