A multi-granularity method, device, and medium for identifying associations between user behaviors on the bright and dark webs
By deploying traffic watermark detectors and sparse signal screening at the entrances and exits of the Tor network, combining time correlation marking and wavelet analysis, the accurate identification and cross-domain correlation problems of anonymous user behavior in large-scale network environments are solved, and efficient tracking and identification of monitored websites and dark webs are achieved.
Patent Information
- Application Number
- CN202510624823.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The existing technology is difficult to effectively identify and track the behavior of anonymous network users in a large-scale high-speed network environment, especially in the Tor network. The lack of efficient traffic compression perception algorithms leads to poor scalability, insufficient accuracy in the recognition of fine-grained anonymous user behavior, and cross-domain anonymous user association technology is difficult to achieve efficient association and precise tracking and positioning of dark websites.
By deploying a traffic watermark detector at the entrance and exit of the Tor network, embed traffic watermarks using pseudo-random number generator and shared keys, and combining time correlation marks, we identify monitored websites and dark web websites visited by users, use sparse signal filtering and compression perception processing, combine wavelet analysis and multi-fractal feature extraction, and use the dark web fingerprint library for real-time updates and multi-dimensional association rules to identify user behavior.
It realizes accurate correlation recognition of monitored websites and dark webs accessed by users in a large-scale network environment, improves the accuracy and robustness of behavior recognition, enhances cross-domain tracking continuity, improves monitoring and identification efficiency, and reduces computing resource requirements.
Smart Images

Figure CN120151117B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet technology, and in particular to a multi-granularity method, device, and medium for identifying associations between user behaviors on the bright and dark webs. Background Art
[0002] Anonymous communication networks are widely used on the internet due to their strong privacy protection and data security capabilities. The Tor network, a typical example, provides anonymous browsing services to millions of users daily. However, due to their concealment, anonymity, and anti-traceability, these networks are also used for illegal activities, posing a serious threat to social stability. Currently, anonymous networks primarily maintain user anonymity through multi-layer encryption and random path selection mechanisms. However, traffic analysis techniques targeting these networks, such as web fingerprinting attacks, can now identify the websites visited by users by analyzing encrypted traffic patterns, posing new challenges to user privacy and security. To effectively combat illegal activities exploiting anonymous networks while protecting the privacy rights of legitimate users, network operators must actively implement technical measures to continuously improve their monitoring and analysis capabilities for anonymous network traffic. Existing technologies primarily use traditional machine learning and deep learning for traffic analysis. The former relies on experts to manually extract features such as packet size and time interval to train models such as support vector machines (SVMs) and k-NNs to identify website access patterns. The latter utilizes deep neural networks (DNNs), such as CNNs and transformers, to automatically learn complex features from traffic data, eliminating the need for manual feature engineering. Both approaches aim to address the problems of identifying, tracking, and locating encrypted traffic.
[0003] However, existing technologies are unable to cope with traffic analysis in large-scale high-speed network environments and lack efficient traffic compression sensing algorithms, resulting in poor scalability. At the same time, they are not accurate enough in identifying fine-grained anonymous user behavior and are greatly affected by factors such as missing attributes and signal attenuation. In addition, current cross-domain anonymous user association technologies are difficult to achieve efficient association of large-scale online anonymous network traffic and precise tracking and positioning of dark web sites, which limits the rapid response and effective crackdown on illegal online activities and hinders the construction of a secure and trustworthy network environment. Summary of the Invention
[0004] The present invention provides a multi-granularity method, device and medium for identifying the association between user behavior on the light and dark webs, so as to solve the problem of difficulty in effectively identifying the association between monitored websites and dark webs visited by users in a large-scale network environment.
[0005] To achieve the above objectives, this application provides a multi-granularity method for identifying associations between user behaviors on the bright and dark webs, including:
[0006] Obtaining network traffic from the Internet according to a preset period, classifying and predicting the network traffic according to a set confidence level, and obtaining a first traffic of users accessing the monitored website;
[0007] Performing watermark embedding and time-correlation marking on the first traffic to obtain a traffic watermark; wherein the session corresponding to the marked first traffic is used for the user to perform all subsequent website access behaviors;
[0008] Within a preset time period, if the traffic watermark detector identifies user traffic containing the traffic watermark, the dark web website accessed by the user traffic is defined as the dark web visited by the user; wherein, the traffic watermark detector is deployed on a controlled guard node at the entrance and exit of the Tor network.
[0009] The present invention performs user behavior identification on the first network traffic on the Internet, and utilizes a classification prediction method based on a set confidence level to effectively distinguish the traffic of users accessing monitored websites. This method can improve the accuracy and robustness of behavior identification through strict confidence judgment. Embedding a watermark into the first traffic is equivalent to labeling the traffic with a unique tag. This watermark technology is concealed and not easily tampered with or removed, providing a reliable basis for subsequent tracking and identification. Traffic is associated with a specific session through time-related tagging, which means that even if the user switches websites or uses different network paths in subsequent visits, as long as the session remains active, the traffic watermark can continue to work. Furthermore, as a cross-domain identifier, watermarks can maintain consistency across different network domains and access paths, thereby supporting cross-domain access behavior association. Time correlation tags further help determine the temporal order and logical relationship of these associated behaviors. This combined tagging method can enhance recognition accuracy and improve tracking continuity. Therefore, once the watermark detector identifies that user traffic containing a traffic watermark has accessed a dark web website, the information in the watermark can be used to effectively associate the dark web access behavior with previously monitored website access behavior. This method breaks down data silos and enables cross-domain tracking. In addition, the traffic watermark detector is deployed on a controlled guard node at the entrance and exit of the Tor network. The Tor network is one of the main channels for accessing the dark web, so this deployment location is of strategic significance. The watermark detector can cover a large amount of potential dark web access traffic, improving the efficiency of monitoring and identification.
[0010] Compared to existing technologies, this invention uses a combination of watermark embedding and time-correlation tagging to uniquely and persistently identify user traffic. This allows for accurate tracking and correlation of monitored websites and dark web traffic visited by users, even in large-scale network environments. Furthermore, using traffic watermark detectors deployed at the entrances and exits of the Tor network, it enables precise capture and identification of dark web traffic, thus resolving the difficulty of effectively correlating and identifying monitored websites visited by users with the dark web in large-scale network environments.
[0011] As a preferred solution, watermark embedding and time correlation marking are performed on the first traffic to obtain a traffic watermark, specifically:
[0012] The watermark is embedded into the first traffic according to a pseudo-random number generator and a shared key, and the first traffic is time-correlatedly marked according to a timestamp to obtain the traffic watermark.
[0013] In this preferred solution, the combination of watermarks and timestamps enables traffic to be efficiently tracked and associated in subsequent processing, and accurate traffic identification and correlation analysis can be achieved whether within the same session or across sessions and network domains.
[0014] As a preferred solution, the watermark is embedded into the first traffic according to a pseudo-random number generator and a shared key, specifically:
[0015] Generate a binary sequence according to a pseudo-random number generator and a shared key, and convert the binary sequence into a time interval sequence of Begin cells;
[0016] Begin cells are inserted into the transmission process of the first traffic according to the time interval sequence.
[0017] In this preferred solution, watermark generation relies on a shared key, ensuring that only legitimate recipients can interpret the timing pattern of the Begin cell. Even attackers who intercept traffic cannot recover the watermark information. Furthermore, the sequence generated by the pseudo-random number generator is time-sensitive, making it impossible for attackers to reproduce historical sequences, thus preventing watermark forgery through traffic replay.
[0018] As a preferred solution, the first traffic is time-correlatedly marked according to the timestamp, specifically:
[0019] Obtaining a timestamp of a first data packet of the first traffic, and generating a watermark sequence according to the timestamp of the first data packet;
[0020] The watermark sequence is used as the time fingerprint of the HS-RP circuit and is embedded into all associated traffic by directionally adjusting the transmission delay of adjacent data packets. The associated traffic refers to the traffic generated when users access different services through the same HS-RP circuit. The different services include monitored websites and subsequent dark web services.
[0021] This preferred solution generates a watermark sequence based on the timestamp of the first data packet, ensuring high synchronization between the watermark and the start time of the traffic, and providing a precise time reference for subsequent traffic analysis and time-correlation tracking. Embedding watermarks by directionally adjusting the delays of adjacent data packets is a relatively covert method that is difficult for ordinary users or malicious attackers to detect, reducing the risk of detection and evasion. Furthermore, the watermark sequence is used as the time fingerprint of the HS-RP circuit and embedded in all associated traffic, enabling cross-service traffic tracking. Even if the user switches between different services during the access process, the watermark will continue to exist and function, ensuring the continuity of tracking.
[0022] As a preferred solution, within a preset time period, if the traffic watermark detector identifies user traffic containing the traffic watermark, the dark web website accessed by the user traffic is defined as the dark web visited by the user, specifically:
[0023] If the traffic watermark detector on the controlled guard node does not detect user traffic containing the traffic watermark within a preset time period, the user traffic of the unselected controlled nodes is intercepted, so that the user traffic passes through the controlled guard node; wherein the controlled guard node is deployed at the entrance and exit of the Tor network;
[0024] If, within the preset time period, the traffic watermark detector on the controlled guard node detects user traffic containing the traffic watermark, a three-hop circuit is established based on the guard node, the intermediate node, and the exit node of the Tor network; wherein the three-hop circuit is used to encrypt and anonymize communications between the user and the hidden service;
[0025] If the exit node through which the three-hop circuit passes is a controlled relay node, the dark web website visited by the user traffic and the IP of the dark web website are determined according to the traffic watermark detector on the relay node, and the dark web website visited by the user traffic is defined as the dark web visited by the user.
[0026] This preferred solution intercepts and reroutes traffic when no watermark is detected, ensuring that all target traffic passes through controlled guard nodes, thereby improving monitoring coverage and accuracy. When the exit node of a three-hop circuit is a controlled relay node, the traffic watermark detector on the relay node can be used to accurately locate the dark web website visited by the user and its IP address.
[0027] As a preferred solution, the controlled guard node includes a controlled entry node and a controlled exit node;
[0028] The controlled ingress node is deployed with a traffic watermark generator, and the controlled egress node is deployed with a traffic watermark detector;
[0029] The dark web websites visited by the same user are identified based on the traffic watermark generator and the traffic watermark detector.
[0030] In this preferred solution, due to the unique nature of the Tor network, user access behavior often spans multiple network domains and different access paths. By deploying appropriate watermark processing equipment at the entry and exit nodes, cross-domain traffic correlation can be supported. Even if the user switches between different network paths or uses different hidden services during the access process, as long as the traffic carries the same watermark, it can be accurately correlated and identified by the system.
[0031] As a preferred solution, the network traffic is classified and predicted according to the set confidence level to obtain the first traffic of users visiting the monitored website, specifically:
[0032] Performing sparse signal screening and compressed sensing processing on the network traffic to obtain compressed traffic;
[0033] The compressed traffic is classified and predicted according to a set confidence level to obtain the first traffic of the user visiting the monitored website.
[0034] This preferred solution effectively reduces network traffic volume through sparse signal screening and compressed sensing processing. When identifying user behavior, introducing a confidence threshold ensures the reliability and stability of the recognition results; only when the confidence level of the recognition result reaches or exceeds the set threshold will it be considered a valid user behavior, which helps reduce false positives and missed negatives.
[0035] As a preferred solution, sparse signal screening and compressed sensing processing are performed on the network traffic to obtain compressed traffic, specifically:
[0036] According to a k-sparsity constraint, sparse feature vectors are screened from the network traffic to construct an anonymous traffic feature set; wherein the k-sparsity constraint is established by performing sparse distribution statistics on historical anonymous traffic data;
[0037] Compressed sensing is performed on the anonymous traffic feature set according to a perception matrix to obtain the compressed traffic; wherein the perception matrix is obtained by mapping a generation matrix in a preset manner, and the generation matrix is established according to a selected error correction code.
[0038] In this preferred solution, the k-sparsity constraint requires selecting no more than k non-zero elements from the dataset, which are considered to be the most representative. This selection process avoids considering all possible feature combinations, thereby greatly improving the efficiency of feature extraction. Using compressed sensing technology combined with a specific perception matrix to compress anonymous traffic feature sets can significantly reduce the amount of data while retaining sufficient information for subsequent user behavior identification or traffic analysis. Furthermore, by reducing the amount of data and optimizing the feature extraction process, this process can save computing resources and increase processing speed, especially when processing large-scale network traffic data.
[0039] As a preferred solution, the compressed traffic is classified and predicted according to a set confidence level to obtain the first traffic of the user visiting the monitored website, specifically:
[0040] Extracting onion service traffic from the compressed traffic in a predetermined manner;
[0041] Perform feature screening on the onion service traffic to obtain a traffic representation vector;
[0042] The traffic representation vector is classified and predicted according to a set confidence level to obtain the first traffic of the user visiting the monitored website.
[0043] This preferred solution performs feature screening on onion service traffic to obtain a traffic representation vector. This process can focus on the most critical feature information, which helps to make subsequent classification predictions more accurate.
[0044] As a preferred solution, the onion service traffic is subjected to feature screening to obtain a traffic representation vector, specifically:
[0045] Performing wavelet analysis on the onion service traffic to obtain multifractal features;
[0046] Calculating information leakage amounts of several features in the multifractal features according to conditional entropy to obtain an information leakage amount set; taking features corresponding to the N largest information leakage amounts in the information leakage amount set from the multifractal features to form a first feature set;
[0047] In the current dimensional space and the preset low-dimensional space, calculating the conditional probability between each pair of data points in the first feature set, to obtain a first conditional probability set and a second conditional probability set respectively;
[0048] With the goal of minimizing the KL divergence loss between the first conditional probability set and the second conditional probability set, mapping the first feature set from the current dimensional space to the low-dimensional space to obtain a second feature set;
[0049] The second feature set is subjected to a nonlinear transformation according to an encoder to obtain the flow representation vector; wherein the encoder is established based on an attention mechanism and a multi-layer perceptron.
[0050] This preferred solution processes onion service traffic through wavelet analysis to obtain multifractal features. These features can capture the complexity and nonlinear characteristics of traffic data, providing a rich information basis for subsequent analysis. The high leakage features in the multifractal features are screened using the amount of information leakage as a quantitative indicator. This process ensures that the selected features are closely related to user privacy and information leakage risks, helping subsequent analysis to focus more on key features and improve the accuracy and efficiency of the analysis. By minimizing the KL divergence loss between the first conditional probability set and the second conditional probability set for dimensionality reduction, the dimensionality of the data can be significantly reduced while retaining key information.
[0051] As a preferred solution, the traffic representation vector is classified and predicted according to a set confidence level to obtain the first traffic of the user visiting the monitored website, specifically:
[0052] In the traffic representation vector, the distance between each traffic embedding and the centroid of the monitored website set is calculated to obtain a number of distance sets;
[0053] For a distance set corresponding to a first flow embedded in the plurality of distance sets, if the nearest centroid distance is less than the radius of the website corresponding to the nearest centroid, and the difference between the second nearest centroid distance and the nearest centroid distance is greater than a preset threshold, then it is determined that the website corresponding to the nearest centroid is the first monitored website visited by the user, and the flow of the user visiting the first monitored website is defined as the first flow of the user visiting the monitored website;
[0054] The closest centroid distance is the distance between the first flow embedding and the closest centroid, and the second closest centroid distance is the distance between the first flow embedding and the second closest centroid.
[0055] This preferred solution calculates the distance between each traffic embedding and the centroid of the monitored website set, generating a distance set that reflects the similarity between traffic and each website. By combining absolute and relative distance constraints, a high-confidence classification framework is constructed in the embedding space. Only traffic that clearly matches the characteristics of the target website is accepted for prediction, making it particularly suitable for scenarios that require strict filtering of false positives.
[0056] As a preferred solution, after obtaining the first traffic of users visiting the monitored website, the method further includes:
[0057] By matching the first traffic with known dark web service features in a dark web fingerprint library, the hidden dark web accessed by the user is identified; wherein the dark web fingerprint library is updated in real time based on the dark web and its real-time access information.
[0058] This preferred solution uses a dark web fingerprint library for feature matching, enabling efficient dark web identification without adding excessive system overhead. Compared to in-depth analysis of all traffic, this feature matching approach is more lightweight and suitable for deployment and application in large-scale network environments.
[0059] As a preferred solution, the dark web fingerprint library is updated in real time based on the dark web and its real-time access information, including:
[0060] Extract the traffic data of users accessing the dark web according to the high-frequency time period to obtain comprehensive traffic data; wherein the high-frequency time period is the time window in which the number of users accessing the dark web is higher than the preset number of dark web accesses;
[0061] Extracting target features of users accessing the dark web from the comprehensive traffic data to obtain a target feature set; performing user access behavior pattern recognition on the comprehensive traffic data according to a clustering algorithm to obtain a behavior pattern recognition result;
[0062] The target feature set and the behavior pattern recognition result form a fingerprint feature set;
[0063] According to the preset update mechanism, the fingerprint feature set is stored in the dark web fingerprint library to obtain the updated dark web fingerprint library.
[0064] This optimal solution accurately identifies user activity windows by extracting the highest frequency of dark web access, thus avoiding wasting resources on inactive periods during data collection. Clustering algorithms automatically group similar traffic data together, allowing for rapid identification of distinct user behavior patterns and improving identification efficiency.
[0065] As a preferred solution, after defining the dark web website accessed by the user traffic as the dark web accessed by the user, the method further includes:
[0066] After a preset time interval, obtain the current dark web fingerprint database;
[0067] Detecting traffic containing a preset tag information set in the current dark web fingerprint library, and defining traffic that successfully matches as traffic to be tested;
[0068] Scoring the traffic to be tested according to multi-dimensional association rules to obtain a user suspicion score; wherein the multi-dimensional association rules are established based on the number of occurrences of preset tags in dark web traffic, the category of dark web websites, dark web access time, and historical behavior deviation thresholds;
[0069] The traffic to be tested whose user suspicion score is higher than a preset threshold is defined as suspicious traffic, and the traffic with the same preset tag information in the suspicious traffic is classified to obtain the dark web traffic classification result of the user's suspicious access behavior.
[0070] This preferred solution can accurately locate traffic data that may be related to suspicious user behavior by detecting traffic containing preset tag information sets in the dark web fingerprint library, avoiding the tedious process of large-scale data screening required in traditional methods. The multi-dimensional association rules not only consider the number of occurrences of preset tags in dark web traffic, but also combine multiple dimensions such as the category of dark web websites, dark web access time, and historical behavior deviations, providing a more comprehensive and accurate basis for user suspicion scoring; while existing classification methods may rely more on a single or limited feature dimension for classification. By classifying suspicious traffic with the same preset tag information, the classification results of the user's suspicious access behavior can be obtained, which helps network security personnel to have a deeper understanding of the user's suspicious behavior patterns and find several dark webs visited by the current user.
[0071] As a preferred solution, the hidden dark web accessed by the user is identified by matching the first traffic with known dark web service features in a dark web fingerprint library, specifically:
[0072] Performing enhanced flow marking on the first flow to obtain enhanced flow marking information; wherein the enhanced flow marking includes a dynamic watermark mark composed of a composite watermark sequence and a multi-protocol mark composed of an encrypted watermark segment, the composite watermark sequence is obtained by performing a hash operation on a timestamp and a data packet size, and the encrypted watermark segment is obtained by inserting a custom data field during a handshake process of a transport layer security protocol;
[0073] After a preset time interval, the current dark web fingerprint library is obtained, and the first dark web traffic having the same enhanced traffic marking information as the first traffic is obtained from the dark web fingerprint library according to the enhanced index, and the website corresponding to the first dark web traffic is defined as the hidden dark web visited by the user; wherein the enhanced index is obtained by establishing an index containing the enhanced traffic marking information for each dark web traffic record in the dark web fingerprint library.
[0074] In this preferred solution, due to the customization of multi-protocol tags and the irreversibility of hash operations, the enhanced traffic tag information is highly unique and concealed. Therefore, even in complex network environments and with large amounts of traffic data, it is still possible to accurately identify the hidden dark web that the user is visiting. The enhanced index is established based on the enhanced traffic tag information, which contains highly unique dynamic watermark tags and multi-protocol tags. This means that each dark web traffic record has a corresponding, unique enhanced traffic tag. Therefore, when it is necessary to retrieve dark web traffic with the same tag information as specific traffic, the enhanced index can provide accurate matching capabilities to ensure the accuracy of the retrieval results.
[0075] The present application also provides a multi-granularity device for identifying the association of user behavior on the bright and dark web, including a monitoring module, a marking module, and an identification module;
[0076] The monitoring module is configured to obtain network traffic from the Internet according to a preset period, classify and predict the network traffic according to a set confidence level, and obtain a first traffic volume of users accessing the monitored website;
[0077] The marking module is configured to perform watermark embedding and time-correlation marking on the first traffic to obtain a traffic watermark; wherein the session corresponding to the marked first traffic is used for the user to perform all subsequent website access behaviors;
[0078] The identification module is used to define the dark web website accessed by the user traffic as the dark web visited by the user if the traffic watermark detector identifies the user traffic containing the traffic watermark within a preset time period; wherein the traffic watermark detector is deployed on a controlled guard node at the entrance and exit of the Tor network.
[0079] As a preferred solution, the marking module is specifically:
[0080] The watermark is embedded into the first traffic according to a pseudo-random number generator and a shared key, and the first traffic is time-correlatedly marked according to a timestamp to obtain the traffic watermark.
[0081] As a preferred solution, the watermark is embedded into the first traffic according to a pseudo-random number generator and a shared key, specifically:
[0082] Generate a binary sequence according to a pseudo-random number generator and a shared key, and convert the binary sequence into a time interval sequence of Begin cells;
[0083] Begin cells are inserted into the transmission process of the first traffic according to the time interval sequence.
[0084] As a preferred solution, the first traffic is time-correlatedly marked according to the timestamp, specifically:
[0085] Obtaining a timestamp of a first data packet of the first traffic, and generating a watermark sequence according to the timestamp of the first data packet;
[0086] The watermark sequence is used as the time fingerprint of the HS-RP circuit and is embedded into all associated traffic by directionally adjusting the transmission delay of adjacent data packets. The associated traffic refers to the traffic generated when users access different services through the same HS-RP circuit. The different services include monitored websites and subsequent dark web services.
[0087] As a preferred solution, the identification module includes a truncation unit, a circuit unit and an association unit;
[0088] The truncation unit is configured to, if the traffic watermark detector on the controlled guard node does not detect user traffic containing the traffic watermark within a preset time period, truncate user traffic from unselected controlled nodes so that the user traffic passes through the controlled guard node; wherein the controlled guard node is deployed at the entrance and exit of the Tor network;
[0089] The circuit unit is configured to establish a three-hop circuit based on the guard nodes, intermediate nodes, and exit nodes of the Tor network if the traffic watermark detector on the controlled guard node detects user traffic containing the traffic watermark within the preset time period; wherein the three-hop circuit is used to encrypt and anonymize communications between the user and the hidden service;
[0090] The association unit is used to determine the dark web website visited by the user traffic and the IP address of the dark web website according to the traffic watermark detector on the relay node if the exit node passed by the three-hop circuit is a controlled relay node, and define the dark web website visited by the user traffic as the dark web visited by the user.
[0091] As a preferred solution, the controlled guard node includes a controlled entry node and a controlled exit node;
[0092] The controlled ingress node is deployed with a traffic watermark generator, and the controlled egress node is deployed with a traffic watermark detector;
[0093] The dark web websites visited by the same user are identified based on the traffic watermark generator and the traffic watermark detector.
[0094] As a preferred solution, the monitoring module includes a compression unit and an identification unit;
[0095] The compression unit is configured to perform sparse signal screening and compressed sensing processing on the network traffic to obtain compressed traffic;
[0096] The identification unit is used to classify and predict the compressed traffic according to a set confidence level to obtain the first traffic of the user visiting the monitored website.
[0097] As a preferred solution, the compression unit includes a screening subunit and a compression subunit;
[0098] The screening subunit is configured to screen out sparse feature vectors from the network traffic according to a k-sparsity constraint condition to construct an anonymous traffic feature set; wherein the k-sparsity constraint condition is established by performing sparse distribution statistics on historical anonymous traffic data;
[0099] The compression subunit is used to perform compressed sensing on the anonymous traffic feature set according to a sensing matrix to obtain the compressed traffic; wherein the sensing matrix is obtained by mapping the generation matrix in a preset manner, and the generation matrix is established according to a selected error correction code.
[0100] As a preferred solution, the recognition unit includes an onion subunit, a feature subunit and a prediction subunit;
[0101] The onion sub-unit is used to extract onion service traffic from the compressed traffic in a preset manner;
[0102] The feature subunit is used to perform feature screening on the onion service traffic to obtain a traffic representation vector;
[0103] The prediction subunit is used to perform classification prediction on the traffic representation vector according to a set confidence level to obtain the first traffic of the user visiting the monitored website.
[0104] As a preferred solution, the characteristic subunit is specifically:
[0105] Performing wavelet analysis on the onion service traffic to obtain multifractal features;
[0106] Calculating information leakage amounts of several features in the multifractal features according to conditional entropy to obtain an information leakage amount set; taking features corresponding to the N largest information leakage amounts in the information leakage amount set from the multifractal features to form a first feature set;
[0107] In the current dimensional space and the preset low-dimensional space, calculating the conditional probability between each pair of data points in the first feature set, to obtain a first conditional probability set and a second conditional probability set respectively;
[0108] With the goal of minimizing the KL divergence loss between the first conditional probability set and the second conditional probability set, mapping the first feature set from the current dimensional space to the low-dimensional space to obtain a second feature set;
[0109] The second feature set is subjected to a nonlinear transformation according to an encoder to obtain the flow representation vector; wherein the encoder is established based on an attention mechanism and a multi-layer perceptron.
[0110] As a preferred solution, the prediction subunit is specifically:
[0111] In the traffic representation vector, the distance between each traffic embedding and the centroid of the monitored website set is calculated to obtain a number of distance sets;
[0112] For a distance set corresponding to a first flow embedded in the plurality of distance sets, if the nearest centroid distance is less than the radius of the website corresponding to the nearest centroid, and the difference between the second nearest centroid distance and the nearest centroid distance is greater than a preset threshold, then it is determined that the website corresponding to the nearest centroid is the first monitored website visited by the user, and the flow of the user visiting the first monitored website is defined as the first flow of the user visiting the monitored website;
[0113] The closest centroid distance is the distance between the first flow embedding and the closest centroid, and the second closest centroid distance is the distance between the first flow embedding and the second closest centroid.
[0114] As a preferred solution, after obtaining the first traffic of users visiting the monitored website, the method further includes:
[0115] By matching the first traffic with known dark web service features in a dark web fingerprint library, the hidden dark web accessed by the user is identified; wherein the dark web fingerprint library is updated in real time based on the dark web and its real-time access information.
[0116] As a preferred solution, the dark web fingerprint library is updated in real time based on the dark web and its real-time access information, including:
[0117] Extract the traffic data of users accessing the dark web according to the high-frequency time period to obtain comprehensive traffic data; wherein the high-frequency time period is the time window in which the number of users accessing the dark web is higher than the preset number of dark web accesses;
[0118] Extracting target features of users accessing the dark web from the comprehensive traffic data to obtain a target feature set; performing user access behavior pattern recognition on the comprehensive traffic data according to a clustering algorithm to obtain a behavior pattern recognition result;
[0119] The target feature set and the behavior pattern recognition result form a fingerprint feature set;
[0120] According to the preset update mechanism, the fingerprint feature set is stored in the dark web fingerprint library to obtain the updated dark web fingerprint library.
[0121] As a preferred solution, after defining the dark web website accessed by the user traffic as the dark web accessed by the user, the method further includes:
[0122] After a preset time interval, obtain the current dark web fingerprint database;
[0123] Detecting traffic containing a preset tag information set in the current dark web fingerprint library, and defining traffic that successfully matches as traffic to be tested;
[0124] Scoring the traffic to be tested according to multi-dimensional association rules to obtain a user suspicion score; wherein the multi-dimensional association rules are established based on the number of occurrences of preset tags in dark web traffic, the category of dark web websites, dark web access time, and historical behavior deviation thresholds;
[0125] The traffic to be tested whose user suspicion score is higher than a preset threshold is defined as suspicious traffic, and the traffic with the same preset tag information in the suspicious traffic is classified to obtain the dark web traffic classification result of the user's suspicious access behavior.
[0126] As a preferred solution, the hidden dark web accessed by the user is identified by matching the first traffic with known dark web service features in a dark web fingerprint library, specifically:
[0127] Performing enhanced flow marking on the first flow to obtain enhanced flow marking information; wherein the enhanced flow marking includes a dynamic watermark mark composed of a composite watermark sequence and a multi-protocol mark composed of an encrypted watermark segment, the composite watermark sequence is obtained by performing a hash operation on a timestamp and a data packet size, and the encrypted watermark segment is obtained by inserting a custom data field during a handshake process of a transport layer security protocol;
[0128] After a preset time interval, the current dark web fingerprint library is obtained, and the first dark web traffic having the same enhanced traffic marking information as the first traffic is obtained from the dark web fingerprint library according to the enhanced index, and the website corresponding to the first dark web traffic is defined as the hidden dark web visited by the user; wherein the enhanced index is obtained by establishing an index containing the enhanced traffic marking information for each dark web traffic record in the dark web fingerprint library.
[0129] The present application also provides a storage medium on which a computer program is stored. The computer program is called and executed by a computer to implement the multi-granularity method for identifying associations between light and dark web user behaviors as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0130] Figure 1 This is a flowchart of a multi-granularity method for identifying associations between user behaviors on the bright and dark webs, provided in an embodiment of the present application;
[0131] Figure 2 This is a schematic diagram of flow representation generation provided by an embodiment of the present application;
[0132] Figure 3 This is the association traceability diagram provided by the embodiment of the present application;
[0133] Figure 4 This is a data processing flow chart provided in an embodiment of the present application;
[0134] Figure 5 This is the overall framework diagram provided by the embodiment of the present application;
[0135] Figure 6 This is a structural diagram of a multi-granularity device for identifying associations between light and dark web user behaviors provided in an embodiment of the present application. DETAILED DESCRIPTION
[0136] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0137] In the description of this application, it should be understood that the terms "first," "second," "third," and "fourth" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Thus, a feature defined as "first," "second," "third," and "fourth" may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, "several" means two or more.
[0138] The multi-granularity method for identifying the association between user behaviors on the light and dark webs provided in the embodiments of the present application is mainly used to design efficient and collaborative detection of anonymous traffic for legal, compliant and authorized platforms such as network operators in an open Internet environment, accurately identify encryption and obfuscation protocols, deeply explore the identities of anonymous users, and associate the communication relationship between access to the surface network and the dark web, thereby providing support for rapid response and effective crackdown on illegal network activities and helping to build a safe and trustworthy network environment.
[0139] Example 1:
[0140] See also Figure 1The embodiment of the present application provides a multi-granularity method for identifying associations between user behaviors on the bright and dark webs, including S1 to S3. The specific implementation steps are as follows:
[0141] S1. Obtain network traffic from the Internet according to a preset period, classify and predict the network traffic according to a set confidence level, and obtain the first traffic of users visiting the monitored website.
[0142] Step S1 of the embodiment of the present application includes S1.1 to S1.5, wherein S1.1 is a process of performing compressed sensing processing on the data, S1.2 is a process of performing wavelet analysis on the data, S1.3 is a process of screening the data based on the amount of leakage, S1.4 is a process of performing dimensionality reduction on the data based on the t-SNE technology, and S1.5 is a process of screening the traffic to the monitored website, specifically:
[0143] S1.1. Legally obtain anonymous network traffic for testing by collaborating with law enforcement agencies, utilizing publicly available data sets and platforms, and participating in legitimate cybersecurity research projects;
[0144] Using methods such as deep neural networks (DNNs), features are automatically extracted from anonymous network traffic to obtain several feature vectors. These feature vectors are then screened using mainstream machine learning models and algorithms such as k-NN and SVM to identify those that effectively represent anonymous network traffic, thereby forming a feature vector set. k-NN is a non-parametric, supervised learning classifier, and SVM is a support vector machine.
[0145] Based on the k-sparseness constraint, feature vectors that satisfy the k-sparse signal are further screened from the feature vector set. These feature vectors are sparse feature vectors. An anonymous traffic feature set is constructed based on the screened sparse feature vectors to characterize the anonymous network traffic. The k-sparseness constraint is established by analyzing the sparse distribution characteristics of historical anonymous traffic data.
[0146] Perform compressed sensing on the anonymous traffic feature set according to the sensing matrix to obtain compressed traffic;
[0147] The sensing matrix is obtained by mapping the generating matrix according to a preset method. The generating matrix is established according to the selected error correction code. This embodiment uses the classic error correction code - BCH code as an example to explain the specific generation and application process of the sensing matrix. The specific steps are as follows:
[0148] Step 1: Select an error-correcting code. First, a suitable error-correcting code must be selected. Given that BCH codes have a well-defined algebraic structure and meet the RIP (Restricted Isometric Property) requirement, this embodiment uses BCH codes as the basis for constructing the sensing matrix.
[0149] Step 2: Construct the perception matrix Φ. Next, take the BCH code as an example to construct the generator matrix, which will be used to generate the perception matrix Φ required for projection. The specific construction process is as follows: First, select the appropriate parameters of the BCH code, including the code length n, the number of information symbols k, and the error correction capability t; then, construct the generator matrix of the BCH code based on these parameters. , where each element is filled according to the coding rules of BCH code; finally, the bipolar mapping method is used to generate the matrix Convert to the perception matrix Φ, that is, map each 0 element in the matrix to " ", each 1 element in the matrix is mapped to " ";in, is the dimension of the perception matrix.
[0150] Step 3: Apply to anonymous traffic feature compression. Once the perception matrix Φ is constructed, it can be applied to anonymous traffic feature compression. Assume that given an N-dimensional sparse signal , through the perception matrix Map it to M-dimensional space to obtain a compressed M-dimensional feature vector This process can be described by the following mathematical expression:
[0151]
[0152] in, is the flow characteristic before compression, is the flow characteristic after compression, It is a perception matrix generated based on the error correction code.
[0153] It is worth noting that in addition to the BCH code-based method, there are other methods for generating perception matrices, such as Gaussian random projection, etc. These methods are similar to the BCH code method in principle, but differ in their specific implementation.
[0154] In this embodiment, S1.1, through sparse signal screening and compressed sensing processing, the amount of network traffic can be effectively reduced. When identifying user behavior, the introduction of a confidence threshold can ensure the reliability and stability of the recognition results; only when the confidence level of the recognition result reaches or exceeds the set threshold will it be considered as a valid user behavior, which helps to reduce false positives and false negatives.
[0155] Furthermore, the k-sparsity constraint requires that no more than k non-zero elements be selected from the dataset, which are considered the most representative. This selection process avoids considering all possible feature combinations, greatly improving the efficiency of feature extraction. Using compressed sensing technology combined with a specific perception matrix to compress anonymous traffic feature sets can significantly reduce the amount of data while retaining sufficient information for subsequent user behavior identification or traffic analysis. Furthermore, by reducing the amount of data and optimizing the feature extraction process, this process can save computing resources and increase processing speed, especially when processing large-scale network traffic data.
[0156] In addition, while previous studies have proposed various linear projection methods, this embodiment proposes a deterministic projection technique based on error-correcting codes. Based on this, it explores strategies for generating sensing matrix elements and studies the intrinsic relationship between the dimension of the compressed sensing matrix and signal sparsity and dimensionality. This effectively extracts useful information from traffic while minimizing the number of traffic samples, achieving efficient information extraction and traffic compression, thereby enabling efficient traffic analysis and processing in high-speed network environments. Furthermore, it ensures that the compressed data retains the key features of the original signal, improving the accuracy and scalability of traffic analysis.
[0157] S1.2. Extracting onion service traffic from the compressed traffic according to a predetermined method; wherein the predetermined method may be based on port and protocol identification, deep packet inspection (DPI), traffic feature analysis, behavioral pattern recognition, and machine learning models;
[0158] Wavelet analysis is performed on the onion service traffic to obtain multifractal characteristics. The specific processing method is as follows:
[0159] ① Basic function definition: Let is a basic function with compact support; the reconstruction error is minimized by L1 regularization to generate an adaptive wavelet basis function that matches the Tor traffic characteristics, satisfying the following conditions: and .
[0160] ② Traffic data preprocessing: Perform preprocessing on the onion service traffic, including formalization, segmentation and normalization, and assume that the traffic is a time series , The specific processing method is as follows: Convert the time series of traffic (expressed as a sequence of encrypted data packets) into a discrete time series to meet the requirements of wavelet analysis; Segment the traffic based on the time interval between packet arrivals or set a fixed window length (for example, every 1000 data packets as a group) to cope with the interference that may be brought by multi-label mixed traffic; Use methods such as Z-score normalization to eliminate the scale differences between traffic data, thereby preventing noise from having an adverse impact on the multifractal spectrum analysis; Among them, Z-score normalization, also known as standard score normalization or zero-mean normalization, is a commonly used data preprocessing technique.
[0161] ③ Discrete wavelet transform: Use an adaptive wavelet basis function to perform a discrete wavelet transform on the preprocessed traffic data to obtain wavelet transform coefficients ; Among them, , is the discrete input signal (traffic data) after preprocessing, usually expressed as a time series; is an integer, representing the dilation scale of the wavelet basis function, and the scale is inversely proportional to the frequency; is an integer, representing the translation position of the wavelet basis function on the time axis. and mentioned in the following text all follow this definition.
[0162] ④ Interval and neighborhood definition: Define the binary interval (dyadic interval) as , in order to lay a foundation for subsequent multifractal analysis; Among them, ;
[0163] Define the traditional three-neighborhood , and in view of the multi-label confusion characteristics of anonymous traffic, expand it into a cross-scale mixed neighborhood: , to capture the cross-label correlation characteristics; Among them, .
[0164] ⑤ Feature extraction: At all finer scales, in the case, calculate the local supremum of the wavelet coefficients located within the spatial neighborhood; Among them, ; And there exists j′ < j, representing all wavelet coefficients at scale j′; is a variable used to index the wavelet coefficients, for the k value at a given scale j′;
[0165] Structure function definition: Based on the local supremum of the wavelet coefficients, define the structure function to extract signal features; Among them, , q is a real number (positive or negative), which is used to adjust the sensitivity of the structure function to wavelet coefficients of different amplitudes. Its core function is to quantify the multi-scale statistical characteristics of the signal in different amplitude ranges;
[0166] Definition and calculation of scaling index: In view of the non-stationary characteristics of Tor traffic, such as the existence of extreme values, the scaling index Define and Truncation is performed according to quantiles (such as 95%) to avoid outliers dominating the scaling index estimate; .
[0167] ⑥ Multifractal spectrum calculation: The multifractal spectrum of signal X is obtained by Legendre transform of scaling exponent , which is a key step in multifractal analysis and is used to reveal the complexity and self-similarity of traffic data; , It is used to measure the local smoothness or roughness of a signal and quantify local singularities. When h > 0, the signal is smooth (low-frequency trend); when h < 0, the signal has sharp mutations (such as sudden peaks).
[0168] According to the above formula and data processing method, wavelet analysis is performed on the onion service traffic to obtain multi-fractal features. These features can reflect the characteristics of Tor traffic that are difficult to directly discover but meaningful, providing strong support for subsequent traffic analysis.
[0169] In this embodiment, S1.2 performs feature screening on onion service traffic to obtain a traffic representation vector. This process can focus on the most critical feature information, which helps to make subsequent classification predictions more accurate. By combining advanced technologies such as wavelet analysis and attention mechanism, meaningful features can be automatically extracted from complex network traffic, solving the problem of low accuracy of traffic association algorithms and models caused by factors such as missing attributes, signal attenuation, and traffic jitter.
[0170] S1.3. Using information leakage as a quantitative indicator, calculate the information leakage of all features in the multifractal feature to obtain an information leakage set; "using information leakage as a quantitative indicator" means using information leakage to measure the amount of information that the system can learn from a certain feature of the website, and "information leakage" reflects the characteristics of the feature. What it can reveal about the status of the website The amount of information;
[0171] Sort the information leakage amount of the information leakage amount set according to the data size, and select the top ten features with the largest leakage amount to form the first feature set;
[0172] Among them, the amount of information leakage in the open world scenario , specifically:
[0173]
[0174] in, is a random variable The entropy of is the monitored website, is a characteristic of a particular representation, Represents additional information or context in the open world, is known And consider In the case of Conditional entropy.
[0175] In this embodiment S1.3, by screening high-leakage features to form the first feature set, it not only reduces the interference of redundant data on the classifier and encoder, but also ensures that the scheme focuses on the dimensions that are most likely to leak information. It can improve analysis efficiency while enhancing the robustness against attacks, thereby avoiding problems such as overfitting caused by an excessive number of features, increased model complexity, reduced generalization ability, and increased computational cost.
[0176] S1.4. Based on the t-SNE technique, a Gaussian kernel function is used to calculate the similarity (conditional probability) between each pair of data points in the first feature set in the high-dimensional space (the current dimensional space). That is, the probability of a point selecting a point as a neighbor is calculated, resulting in a first conditional probability set (the high-dimensional space similarity matrix). In the low-dimensional space, the probability of a point selecting a point as a neighbor is also expressed in the form of conditional probability, but this time the t-distribution is used as the kernel function, resulting in a second conditional probability set (the low-dimensional space similarity matrix). t-SNE is a nonlinear dimensionality reduction technique used for high-dimensional data visualization.
[0177] With the goal of minimizing the Kullback-Leibler divergence (KL divergence) loss between the first conditional probability set and the second conditional probability set, the points of the first feature set in the low-dimensional space are continuously adjusted through optimization algorithms such as gradient descent, so that the distribution of data points in the low-dimensional space retains the relative proximity relationship in the high-dimensional space as much as possible, until convergence is achieved or the preset number of iterations is reached. Finally, the second feature set after dimensionality reduction is obtained;
[0178] The encoder is used to perform nonlinear transformation on the second feature set to obtain a traffic representation vector, and the traffic representation vector is processed into a specified format for input of the classifier and the encoder.
[0179] Among them, t-SNE technology uses Gaussian kernel function to calculate each pair of data points and The conditional probability between , indicating a point Select Point The probability of being a neighbor can be expressed as:
[0180]
[0181] In low-dimensional space, the points are calculated according to the t-SNE technique. and The conditional probability between , which can be expressed as:
[0182]
[0183] The goal of the t-SNE technique is to minimize the Kullback-Leibler divergence (KL divergence) between high-dimensional and low-dimensional data. Its cost function is for:
[0184]
[0185] Among them, the data points Refers to the original sample (such as eigenvector) in the high-dimensional space to be analyzed, and calculates the point Refers to the mapping points in the low-dimensional space generated by the optimization algorithm, represents the kth sample point in the original high-dimensional data, It is the low-dimensional representation (usually 2D or 3D) of x subscript k after t-SNE mapping; It is the joint probability distribution of similarity between the original high-dimensional data points, calculated by the Gaussian kernel function; It is the joint probability distribution of similarity between low-dimensional mapping points (calculation point y), calculated by t distribution (heavy-tailed distribution); and It is the specific probability value when i and j are determined, and i and j are the indexes used to traverse all data point pairs.
[0186] The encoder is constructed by combining the embedding vector generation method of the multi-head attention mechanism, the residual connection method and the multi-layer perceptron, and the encoder has been fully trained on the training set;
[0187] The encoder includes an input layer, a multi-head attention layer, a subsequent processing layer, and a residual connection. The input layer is used to receive the preprocessed flow representation vector. The multi-head attention layer is used to process the input vector with a self-attention mechanism to extract internal correlation information. The subsequent processing layer includes a multi-layer perceptron (MLP) layer, a 2D convolution layer, a 2D maximum pooling layer, a 1D convolution layer, a 1D maximum pooling layer, and an adaptive maximum convolution pooling layer, which are used to further extract features and transform the output of the multi-head attention layer. The residual connection is used to alleviate the gradient vanishing problem in deep networks and ensure smooth information transmission in the network. Specifically:
[0188] ① Multi-head attention layer (MHSA):
[0189] Composition: This layer contains multiple independent attention heads, each of which can process the input data independently, thereby capturing different representations of the input data in parallel;
[0190] Formula Usage:
[0191] Formula 1: This formula is used to define the output of the multi-head attention layer. It is a learnable weight matrix used to linearly combine the outputs of multiple attention heads to obtain the final output vector; is the number of heads, which determines the number of attention heads processed in parallel; It is The output of an attention head, It means concatenating the outputs of all attention heads along the feature dimension to form a longer vector.
[0192] Second formula: For each attention head, this formula calculates the linear transformation of query (Q), key (K) and value (V). 、 and It is aimed at The learnable weight matrix of each head, 、 and They are query, key and value, Represents the attention mechanism.
[0193] The third formula: This formula is used to calculate the output of the attention mechanism. is the dimension of the key vector, used to scale the dot product to avoid excessive values; Represents the transpose operation of the matrix;
[0194] Among them, the first formula is:
[0195]
[0196] The second formula is:
[0197]
[0198] The third formula is:
[0199]
[0200] ②Multi-layer Perceptron (MLP) layer:
[0201] Composition: This layer is a fully connected neural network, usually containing multiple hidden layers, used to further process the output of the multi-head attention layer.
[0202] Purpose: Perform nonlinear transformation on the output of the multi-head attention layer to extract higher-level features.
[0203] ③Convolutional Layers:
[0204] Composition: The encoder consists of 2D convolutional layers and 1D convolutional layers. The convolutional layers use convolution kernels to slide over the input data and extract local features through convolution operations.
[0205] use:
[0206] 2D Convolutional Layer (2D Conv): processes data with a two-dimensional structure (such as images) and can capture the spatial characteristics of the data.
[0207] 1D Convolutional Layer (1D Conv): processes one-dimensional data (such as time series data) and extracts local patterns in the time series.
[0208] ④Pooling Layers:
[0209] Composition: The encoder contains a 2D max pooling layer and a 1D max pooling layer.
[0210] Purpose: Reduce the spatial size or temporal length of data by downsampling while retaining important features and reducing the amount of computation and the number of parameters.
[0211] 2D Max Pooling: Slide a window on the two-dimensional feature map and select the maximum value in the window as the output. It is often used to reduce the spatial dimension.
[0212] 1D Max Pooling: Slide a window on a one-dimensional feature sequence and select the maximum value in the window as the output. It is often used to reduce the time dimension.
[0213] ⑤Adaptive Max Pooling:
[0214] Composition: This is a special pooling layer whose output size is fixed.
[0215] Purpose: Dynamically adjust the size of the pooling window according to the size of the input data to ensure that the output has a fixed size for subsequent processing.
[0216] ⑥Residual connection
[0217] Composition: Directly add input data to the output of certain layers through skip connections.
[0218] Formula Purpose: The fourth formula illustrates the role of residual connections. Here, F(x) is the output after a series of nonlinear transformations, and x is the input. Residual connections add the input x directly to F(x), helping to alleviate the vanishing gradient problem in deep networks.
[0219] Purpose: Improve the training effect of deep networks, allowing information to flow more smoothly in the network, and improving the training efficiency and performance of the model.
[0220] Among them, the fourth formula is:
[0221]
[0222] To apply this application example, please refer to Figure 2 , Figure 2 This is a flow representation generation diagram provided in an embodiment of the present application, which shows the process of performing nonlinear transformation on the second feature set according to the encoder to generate a flow representation vector.
[0223] S1.5. In the traffic representation vector, use cosine similarity to calculate the distance between each traffic embedding and the centroid of the monitored website set to obtain several distance sets;
[0224] For the distance set corresponding to the first flow embedding in several distance sets, if the nearest centroid distance is less than the radius of the website corresponding to the nearest centroid (and this condition can be defined as "condition 1", that is, the distance between the flow embedding and the website centroid is less than the radius of the website), and the difference between the second nearest centroid distance and the nearest centroid distance is greater than the preset threshold (Also, this condition can be defined as "Condition 2", that is, the difference between the second closest distance and the closest distance exceeds the preset threshold ), the website corresponding to the nearest centroid is determined to be the first monitored website visited by the user, and the traffic of the user visiting the first monitored website is defined as the first traffic of the user visiting the monitored website; and, according to this detection method, all traffic embeddings in the traffic representation vector are traversed, so that the first traffic finally obtained includes traffic from the user visiting several first monitored websites;
[0225] Among them, the radius of the first website is calculated based on the distance between the data point and the data mean in the first website using the standard deviation; the nearest centroid distance is the distance between the first traffic embedding and the nearest centroid, and the second nearest centroid distance is the distance between the first traffic embedding and the second nearest centroid; and "monitored websites" refer to those websites whose access traffic has been collected and analyzed in advance; by extracting features from these traffic, the encoder is trained, and the centroids of these websites are determined based on the position of these feature vectors in the embedding space. In contrast, "unmonitored websites" refer to those websites that have not undergone similar traffic collection and training processes, that is, their traffic features have not been used for encoder training; and the classification method based on conditions 1 and 2 can be defined as a classifier.
[0226] For example, suppose there is a traffic sample whose distance from website C is 0.8 units and its distance from the next closest website D is 1.6 units. At the same time, the difference threshold is set to 0.3 units.
[0227] Judgment process: Condition 1: Distance from traffic to website C (0.8) < Radius of website C (1.0) → Satisfied; Here, the distance between the traffic sample and website C is less than the radius of website C, so condition 1 is satisfied. Condition 2: Distance from the next closest website D (1.6) - Distance from website C (0.8) = 0.8 > Threshold (0.3) → Satisfied; Here, the difference between the distance from the traffic sample to the next closest website D and the distance to the closest website C is greater than the preset threshold, so condition 2 is also satisfied.
[0228] Conclusion: Based on the above conditions, it is determined that the traffic belongs to the behavior of visiting website C.
[0229] Among them, assuming that the website The embedding representation of … , where each is a vector, then the centroid formula is:
[0230]
[0231] Use the standard deviation to calculate the distance between the data point and the mean to get the radius of the website , and its calculation formula is:
[0232]
[0233] Use cosine similarity to calculate each flow embedding and website The distance between the centroids is calculated as:
[0234]
[0235] Condition 1 - The distance between the traffic embedding and the centroid of the website is less than the radius, which can be expressed as:
[0236]
[0237] Condition 2 - The difference between the second closest distance and the closest distance exceeds the threshold, which can be expressed as:
[0238]
[0239] in, is the mean of the embeddings, It is data points; " indicates a vector and centroid The dot product of is a vector The Euclidean norm (length) of is the dimension of the embedding vector, is the index variable to be summed, is the feature vector of the data point at a specific j and k; is the Euclidean norm of the centroid; is the distance between the traffic embedding and the second closest website centroid, is the distance between the traffic embedding and the nearest website centroid, is the threshold, is the embedding vector dimension.
[0240] In this embodiment, S1.5 calculates the distance between each traffic embedding and the centroid of the monitored website set, generating a distance set reflecting the similarity between traffic and each website. By combining absolute and relative distance constraints, a high-confidence classification framework is constructed in the embedding space. Only traffic that clearly matches the characteristics of the target website is considered for prediction, making it particularly suitable for scenarios requiring strict filtering of false positives.
[0241] In this embodiment S1, multifractal features are obtained by processing onion service traffic through wavelet analysis. These features can capture the complexity and nonlinear characteristics of traffic data, providing a rich information foundation for subsequent analysis. High leakage features in multifractal features are screened using information leakage as a quantitative indicator. This process ensures that the selected features are closely related to user privacy and information leakage risks, helping subsequent analysis to focus more on key features and improve the accuracy and efficiency of the analysis. By minimizing the KL divergence loss between the first conditional probability set and the second conditional probability set to perform dimensionality reduction processing, the dimensionality of the data can be significantly reduced while retaining key information.
[0242] S2. Perform watermark embedding and time-correlation marking on the first traffic to obtain a traffic watermark; wherein the session corresponding to the marked first traffic is used for the user to perform all subsequent website access behaviors.
[0243] Step S2 of the embodiment of the present application includes S2.1 to S2.2, wherein S2.1 is a process of embedding a watermark and marking time correlation on the first flow, and S2.2 is a process of performing enhanced flow marking on the first flow, specifically:
[0244] S2.1. Embed a watermark and perform time-correlation marking on the first flow to obtain a flow watermark. The session corresponding to the marked first flow is used by the user for all subsequent website access behaviors.
[0245] The watermark embedding refers to embedding the watermark into the first traffic according to the pseudo-random number generator and the shared key, specifically:
[0246] Generate a binary sequence according to a pseudo-random number generator and a shared key, and convert the binary sequence into a time interval sequence of Begin cells;
[0247] Begin cells are inserted into the transmission process of the first traffic according to a time interval sequence.
[0248] The time correlation marking refers to performing time correlation marking on the first traffic according to the timestamp, specifically:
[0249] Obtain the timestamp of the first data packet of the first flow, and generate a watermark sequence according to the timestamp of the first data packet;
[0250] The watermark sequence is used as the time fingerprint of the HS-RP circuit and is embedded into all associated traffic by directionally adjusting the transmission delay of adjacent data packets. Associated traffic refers to the traffic generated when users access different services through the same HS-RP circuit. Different services include monitored websites and subsequent dark web services.
[0251] The following is a detailed description of the marking method:
[0252] ① Watermark embedding. Its core strategy is this: during an access request, if the client knows the address of the hidden service, the system uses watermarking technology to secretly transmit the client's known identity (i.e., the hidden service address) to the controlled assembly point. When the hidden service receives a packet containing a Relay-Drop, it will directly ignore and discard it, ensuring that this operation does not cause any interference to upper-layer applications. "Relay-Drop" packets refer to packets that carry watermark signals during the network relay process and are designed to be directly discarded.
[0253] This example uses the encoding of communication cells using Relay-Drop packets (as the transmission medium for watermark signals) as an example to explain the watermark embedding method in detail. Specifically, the Tor protocol's Relay-Drop communication cells are used to covertly encode dark web service addresses, achieving highly concealed watermark embedding and detection. The detailed steps are as follows:
[0254] Preparation phase: A key is shared between the encoder and detector for subsequent watermark generation and detection.
[0255] Watermark Embedding: 1. Begin Cell Sequence Generation: Using a pseudo-random number generator (PRNG) and a shared key, a binary sequence B is generated based on the extended Gilbert model. In sequence B, "1" indicates that a Begin cell needs to be generated, and "0" indicates that it should not be generated. The Gilbert model is a statistical model used to describe data transmission errors in digital communication systems, and the Begin cell is a service response mechanism in the Tor protocol, used to identify and coordinate data transmission during network communication.
[0256] 2. Time Interval Calculation: Convert the binary sequence B into a sequence D of Begin cell time intervals. Specifically, iterate through the binary sequence B. Whenever a "1" is encountered, calculate the corresponding Begin cell transmission time based on the preset time interval parameter Δt. This yields a sequence D of time intervals containing the Begin cell transmission times.
[0257] 3. Watermark encoding: The dark web service address information is encoded in a covert manner into the time interval sequence D of the Begin cell without changing the other contents of the Relay-Drop packet. This ensures that the hidden service will directly discard the Relay-Drop packet upon receiving it, without affecting the upper-layer application.
[0258] Begin cells are inserted into the transmission process of the first flow according to the time interval sequence D. In addition to Begin cells, other service response mechanisms in the Tor protocol can also be combined to enhance the concealment and robustness of the watermark.
[0259] For watermark embedding in S2.1 of this embodiment, watermark generation relies on a shared key, ensuring that only legitimate recipients can interpret the time pattern of the Begin cell. Even if an attacker intercepts the traffic, they cannot recover the watermark information. Furthermore, the sequence generated by the pseudo-random number generator is time-sensitive, making it impossible for attackers to reproduce historical sequences, preventing watermark forgery through traffic replay.
[0260] In addition to the above-mentioned watermark embedding method based on the "Begin cell", this embodiment also provides a watermark embedding method using the "hash algorithm" as a supplement, as follows:
[0261] Record each data packet in the first flow Time of arrival at the network node , get the time when the first data packet in the traffic arrives , this time can be accurate to milliseconds or even higher precision to ensure its uniqueness;
[0262] Calculating delays between adjacent data packets in the first flow to obtain a first delay sequence;
[0263] Based on the watermark embedder, according to the watermark sequence The watermark bit in the first delay sequence is adjusted to the corresponding delay, and the data packet in the network traffic is sent according to the adjusted delay to complete the watermark embedding;
[0264] Among them, the watermark sequence It is obtained by capturing the arrival time of the first data packet in the network traffic and obtaining the timestamp; then mapping the timestamp to a hash value of a preset fixed length, performing base conversion and sequence division on the hash value, specifically: selecting a suitable hash function (such as SHA-256 hash function) as the basis for generating the watermark sequence; the arrival time of the first data packet in the first flow is used as the input of the hash function to calculate the initial hash value ,in, , It's time, is a hash function, Is the hash value; the initial hash value Convert to a binary sequence (every 4 bits correspond to 1 hexadecimal bit) to get a binary sequence ; Divide the binary sequence into groups of 8 bits , if the decimal value of a set of binary numbers is greater than 127, then let Otherwise, ; This way, a unique watermark sequence can be generated based on the time of the first data packet ;
[0265] In summary, the watermark embedding process is specifically described as follows: for each pair of adjacent data packets and , according to the formula To calculate the delay difference between them, here . Then, traverse the watermark sequence Each watermark bit in , according to the formula The delay of adjacent data packets is adjusted accordingly. If the watermark bit is "1", the next data packet The sending time of the data packet is delayed by several time units; if the watermark bit is "0", the next data packet The sending time is advanced by several time units. By resending the data packet, the watermark embedding process in the network traffic is completed; Indicates timestamp and timestamp the time interval between It is an integer, indicating the delay in several time units.
[0266] In this embodiment S2.1, for the watermark embedding method using the "hash algorithm", the data packet timestamp is hashed to generate the watermark sequence. This method has the advantages of concealment because the timestamp itself is common in network transmission, and the watermark generated by hashing the timestamp is not easy to detect. At the same time, the hash function has collision resistance, making the generated watermark robust and able to resist certain attacks and interference.
[0267] Furthermore, mapping the timestamp to a hash value of a preset fixed length ensures that the generated watermark sequence has a consistent length, which facilitates subsequent processing, storage, and transmission. At the same time, the fixed-length hash value simplifies the watermark embedding and extraction process, improving efficiency. Sequence partitioning allows hash values to be split into multiple subsequences or fragments to form a more complex watermark sequence. This refined control helps enhance the concealment and robustness of the watermark.
[0268] In addition, since the watermark is embedded by adjusting the delay, which is affected by many factors such as network conditions and device performance, the watermark is likely to be retained even in the event of network fluctuations or packet loss. This makes the watermark highly robust and able to resist network attacks and data tampering to a certain extent.
[0269] ②Time correlation mark:
[0270] Get the timestamp t of the first data packet of the first flow 01 , according to the first data packet timestamp t 01 and shared key A, generate watermark sequence through pseudo-random algorithm .
[0271] The watermark sequence W is used as the time fingerprint of the HS-RP circuit and embedded into all associated traffic by directionally adjusting the transmission delay of adjacent data packets. Associated traffic refers to the traffic generated when users access different services through the same HS-RP circuit. Different services include monitored websites and subsequent dark web services. That is:
[0272] In all associated traffic on the HS-RP circuit, the transmission delay of adjacent data packets is adjusted according to the watermark sequence W. For the i-th data packet, the delayed transmission time is w i , so that the interval between it and the previous data packet conforms to the watermark sequence W; let all traffic on the same HS-RP circuit (whether accessing monitoring websites or dark web services) carry the same time fingerprint W, and use the time fingerprint W as the unique time fingerprint of the circuit and bind it to the hidden service (HS) that created the circuit; among them, the HS-RP circuit is a hidden service communication circuit.
[0273] For the time-correlation tag in S2.1 of this embodiment, a watermark sequence is generated based on the timestamp of the first data packet, ensuring high synchronization between the watermark and the start time of the traffic, providing a precise time reference for subsequent traffic analysis and time-correlation tracking. The watermark is embedded by directionally adjusting the delay of adjacent data packets. This method is relatively covert and difficult to detect by ordinary users or malicious attackers, reducing the risk of detection and evasion. In addition, the watermark sequence is used as the time fingerprint of the HS-RP circuit and embedded in all associated traffic, enabling cross-service traffic tracking. Even if the user switches between different services during the access process, the watermark can still persist and function, ensuring the continuity of tracking.
[0274] In summary, the combination of watermarks and timestamps enables efficient tracking and correlation of traffic in subsequent processing, enabling accurate traffic identification and correlation analysis, whether within the same session or across sessions and network domains.
[0275] S2.2. Performing enhanced flow marking on the first flow to obtain enhanced flow marking information; wherein the session corresponding to the marked first flow is used by the user for all subsequent website access behaviors;
[0276] Among them, the enhanced traffic marking includes dynamic watermark marking composed of a composite watermark sequence and multi-protocol marking composed of encrypted watermark fragments. The composite watermark sequence is obtained by hashing the timestamp and packet size, and the encrypted watermark fragment is obtained by inserting a custom data field during the handshake process of the transport layer security protocol. The specific marking method is as follows:
[0277] Use a hash algorithm to perform a joint hash process on the timestamp, session ID, and packet size of the first flow to obtain a tamper-resistant composite watermark sequence;
[0278] The composite watermark sequence is divided into a first sequence, a second sequence and a third sequence;
[0279] The first segment of the sequence is embedded into the payload of the data packet, the second segment of the sequence is embedded into the option field of the IP data packet, and the third segment of the sequence is embedded into the unused bits of the TCP header, thereby obtaining a data packet payload with a watermark, an IP option field with a watermark, and a TCP header with a watermark, respectively. The data packet payload with a watermark, the IP option field with a watermark, and the TCP header with a watermark constitute dynamic watermark marking information.
[0280] In the TLS handshake phase, a custom encrypted watermark fragment is injected into the protocol to obtain the TLS handshake information with the encrypted watermark fragment;
[0281] The enhanced traffic marking information is composed of dynamic watermark marking information and TLS handshake information; among them, the TLS handshake phase is the key process in the Transport Layer Security Protocol (TLS) for establishing a secure connection between the client and the server; the session ID is a unique identifier used to identify the session between the user and the server; the IP data packet is the data packet of the network layer in the TCP / IP protocol; the TCP header is the header of the transport layer TCP protocol, which is used to control the transmission of information when transmitting data in the TCP connection; the IP option field is an optional field in the IP header, used to support some specific functions or testing purposes.
[0282] In this embodiment S2.2, due to the customization of multi-protocol tags and the irreversibility of hash operations, the enhanced traffic tag information is highly unique and hidden. Therefore, even in a complex network environment and huge traffic data, the hidden dark web accessed by the user can still be accurately identified.
[0283] S3. Within a preset time period, if the traffic watermark detector identifies user traffic containing a traffic watermark, the dark web website accessed by the user traffic is defined as the dark web visited by the user; wherein, the traffic watermark detector is deployed on a controlled guard node at the entrance and exit of the Tor network.
[0284] Step S3 of the embodiment of the present application includes S3.1 to S3.3, wherein S3.1 is the process of identifying the dark web and its address visited by the user, S3.2 is the process of updating the dark web fingerprint library in real time, and S3.3 is the process of identifying the hidden dark web visited by the user and classifying suspicious traffic to deepen the association identification method, specifically:
[0285] S3.1. If, within a preset time period, the traffic watermark detector on the controlled guard node (Guard) does not detect user traffic containing a traffic watermark (indicating that the current hidden service circuit has not selected a controlled guard node), it intercepts user traffic from the unselected controlled node, triggering a circuit anomaly and forcing the system to resend traffic to select a new guard node until the user traffic passes through the controlled guard node. Controlled guard nodes are deployed at the entrance and exit of the Tor network. The traffic watermark detector is responsible for detecting watermarks. It extracts possible watermark bits by analyzing the characteristics of the traffic packets of the selected carrier and compares the calculated characteristic function value with the preset watermark parameters to determine whether the watermark is present.
[0286] Within a preset time period, if the traffic watermark detector on the controlled guard node detects user traffic containing a traffic watermark (indicating that the current hidden service circuit has selected the controlled guard node), the Tor client is controlled to establish a three-hop circuit based on the guard node, intermediate node, and exit node of the Tor network at a default frequency of every 10 minutes; the three-hop circuit is used to encrypt and anonymize the communication between the user and the hidden service;
[0287] If the watermark detector detects that the exit node through which the three-hop circuit passes is a controlled relay node, the traffic watermark detector on the relay node determines the dark web website and the IP address of the dark web website visited by the user traffic, and uses the dark web website and its IP address as the dark web target information visited by the user; if the exit node through which the three-hop circuit passes is not a controlled relay node, the circuit needs to be destroyed immediately and the process repeated until the selected exit node is one of the controlled "honey relays".
[0288] In addition, the controlled guard nodes include controlled entry nodes and controlled exit nodes; the controlled entry nodes are deployed with traffic watermark generators, and the controlled exit nodes are deployed with traffic watermark detectors; based on the traffic watermark generator and the traffic watermark detector, the dark web websites visited by the same user can be identified.
[0289] Within the set maximum time range, if the traffic watermark detector is still unable to identify user traffic containing traffic watermarks, it means that the user who visited the monitored website did not visit the dark web.
[0290] The following is a detailed description of the detection method of the traffic watermark detector:
[0291] ① Detection method for watermark embedding (watermark embedding method based on "Begin cell"):
[0292] First, a detector is used to capture Begin cells in the network flow of the dark web fingerprint library and record their respective timestamps;
[0293] Then, the time interval (IPD) between each pair of adjacent Begin cells is calculated, that is, the time difference between each Begin cell and its previous cell;
[0294] Finally, the detector uses a similarity algorithm to carefully check the time interval of each Begin cell to determine whether it matches the Begin cell sequence pre-generated by the encoder. If the time interval sequence is highly similar or completely matches the Begin cell sequence generated by the encoder, the sequence is considered to be an embedded stream watermark. Based on this discovery, the detector can further analyze the information carried by the watermark, such as data source identifier, timestamp, user identifier, etc., to effectively track and accurately associate hidden services.
[0295] In addition, here is a description of the detection method of the watermark embedding method using the "hash algorithm":
[0296] (1) Using a network packet capture tool, obtain the second traffic that may contain the watermark from the Internet according to a preset period, and record the time when each data packet arrives;
[0297] Calculate the delay between adjacent data packets in the second flow to obtain the second delay sequence ;in, Indicates the delay between adjacent data packets containing watermark information during the watermark extraction period ;
[0298] Calculate the preset third time delay sequence and second time delay sequence according to the Pearson correlation coefficient formula The correlation between several time delay sequences in the , and the correlation set is obtained; wherein the third time delay sequence is the time delay sequence corresponding to the known embedded watermark sequence , among which Indicates the delay between adjacent data packets containing watermark information during the watermark embedding period ;
[0299] The correlation set is greater than the preset threshold T (i.e. >T) is determined as the first flow after the watermark sequence is embedded.
[0300] Among them, the Pearson correlation coefficient formula is:
[0301]
[0302] Here, x and y are the two variables to be analyzed (such as "time" in traffic data), and n is the number of observations.
[0303] In this embodiment of the watermark embedding detection method using a hash algorithm, correlation calculation can quantify the degree of similarity between the third delay sequence and the first delay sequence. By setting a preset threshold, traffic containing the watermark can be accurately screened. This method reduces the possibility of false positives and false negatives, thereby improving detection accuracy.
[0304] ② Detection method for time-related markers:
[0305] Monitor the traffic of the deployed monitoring website and extract the timestamp t of the first data packet of the visiting user 02 ;
[0306] Using the shared key A and the first packet timestamp t 02 Regenerate the watermark sequence W';
[0307] Detect the user's subsequent traffic to the dark web service through the same HS-RP circuit, extract the adjacent data packet interval sequence, and verify whether it matches the watermark sequence W'. If the match is successful, it is determined that the dark web service access traffic and the monitored website access traffic belong to the same user.
[0308] In addition, the specific comparative identification methods include active linkage and passive linkage user behavior identification methods, specifically:
[0309] ① Active linkage: This system actively controls a certain number of routing nodes and waits for them to be selected by the hidden service as its guard nodes (i.e., entry nodes). Once selected, these nodes can communicate directly with the hidden service, revealing its true IP address. Furthermore, even if a user subsequently accesses other dark web services through the same set of controlled nodes, the system can accurately determine that the traffic originates from the same user by detecting the same timing patterns in the communications. Specifically, when a user first visits a monitored website, their HS-RP circuit is marked with a specific timing signature as described above. If the user subsequently accesses another dark web service through the same HS-RP circuit and the circuit is again marked with the same timing signature, it can be concluded that the two visits belong to the same user, enabling cross-site correlation and tracing. For example, if a user first visits a monitored website (and is marked), and then later visits a dark web forum through the same HS-RP circuit, the system can correlate the two by comparing the timing patterns, thereby determining the user's identity or true IP address, as well as the dark web sites and dark web addresses visited.
[0310] ② Passive linkage: When a hidden service matching a signature in the dark web fingerprint database doesn't select a system-controlled routing node as its guard node, the system adopts a passive strategy, frequently creating links in an attempt to be selected by the hidden service as the second hop. This way, even if the system can't directly obtain the hidden service's true IP address, it can indirectly locate the hidden service by tracing its binding relationship with the guard node. This passive linkage strategy provides the system with a fallback in the event that active linkage fails, enhancing the flexibility and reliability of tracking and associating hidden services.
[0311] To apply this application example, please refer to Figure 3 , Figure 3 It is an association traceability diagram provided by an embodiment of the present application, which shows the process of adding traffic watermarks to traffic according to the watermark embedder, detecting traffic watermarks according to the watermark detector, and establishing a three-hop circuit for association traceability on this basis.
[0312] In this embodiment, S3.1, when no watermark is detected, intercepts and reroutes traffic, ensuring that all target traffic passes through the controlled guard node, thereby improving monitoring coverage and accuracy. When the exit node of a three-hop circuit is a controlled relay node, the traffic watermark detector on the relay node can be used to accurately locate the dark web website visited by the user and its IP address. This refined monitoring capability provides valuable intelligence support for network security and law enforcement agencies.
[0313] Furthermore, due to the unique nature of the Tor network, user access behavior often spans multiple network domains and different access paths. By deploying appropriate watermark processing equipment at the entry and exit nodes, cross-domain traffic correlation can be supported. Even if a user switches between different network paths or uses different hidden services during access, as long as the traffic carries the same watermark, it can be accurately correlated and identified by the system.
[0314] S3.2. Update the dark web fingerprint database in real time based on the dark web and its real-time access information.
[0315] The Dark Web Fingerprint Library is a database dedicated to collecting, organizing, and analyzing the characteristics of dark web services. It aims to obtain dark web data resources through legal and compliant channels, such as collaborating with law enforcement agencies, using public search tools, and participating in industry research projects. An exemplary construction method for the Dark Web Fingerprint Library is as follows:
[0316] Targeting and data collection: Identify the types of services and activity areas that the dark web fingerprint database should cover, focusing on specific areas such as illegal transactions and the circulation of cyberattack tools. Obtain dark web data resources through legal and compliant channels, including collaborating with law enforcement agencies, using public search tools, and participating in industry research projects. Focus on collecting web content, transaction records, and user interaction information from dark web forums, trading platforms, and instant messaging spaces.
[0317] Data content analysis: Deeply analyze the collected network traffic to extract statistical features such as packet size, transmission interval, and traffic rate. Analyze the structure of dark web pages and extract key identifiers such as page elements, dynamic scripts, and image hash values.
[0318] Time and space dimension tagging and classification system: Combine traffic transmission timestamps and circuit path information to tag dark web traffic in time and space dimensions. Establish a classification system based on the nature of the service, set business tags such as illegal drug transactions and hacking tool transactions, and achieve refined classification of service types;
[0319] Database Design and Data Storage: Design a multidimensional storage architecture to optimize table structure, field type matching, and retrieval efficiency. Fully annotated dark web data will be systematically stored according to the classification system to form a scalable dark web service feature library. This database supports functions such as time series analysis, path tracing, and behavioral pattern recognition, providing underlying data support for subsequent security monitoring.
[0320] Among them, the update method of the dark web fingerprint library is:
[0321] Using Tor network monitoring tools or dark web traffic analysis systems, extract traffic data of users accessing the dark web according to high-frequency time periods to obtain comprehensive traffic data. The collected data is then cleaned to remove invalid data such as outliers and duplicates. The high-frequency time period is the time window when the number of users accessing the dark web exceeds the preset number of dark web visits. In addition, by introducing a dynamic adjustment mechanism, the high-frequency time period is recalculated regularly (such as daily or weekly) to adapt to changes in user access behavior.
[0322] In the comprehensive traffic data, the target features of users accessing the dark web are extracted to obtain the target feature set; the user access behavior pattern is identified on the comprehensive traffic data based on the clustering algorithm to obtain the behavior pattern recognition result;
[0323] The fingerprint feature set is composed of the target feature set and the behavior pattern recognition result;
[0324] The fingerprint feature set is stored in the dark web fingerprint database according to the preset update mechanism to obtain an updated dark web fingerprint database; wherein the update mechanism specifically includes the following steps:
[0325] ① Add new features to the database: When a new dark web service or an unknown variant of its existing service is detected, these new features (such as API call patterns) are included in the dark web fingerprint library, and a unique ID is assigned to each new feature; where "ID" is the abbreviation of "Identifier" and is used to uniquely identify or distinguish each new feature.
[0326] ② Dynamic weight adjustment: In the event that old fingerprints become invalid due to data updates or other reasons, the weight of the features will be adjusted according to their activity. Specifically, the weight of highly active features will be increased so that they are detected first; at the same time, the weight of low-activity features or features with high false positive rates will be reduced and moved to the historical archive.
[0327] ③ Feature merging and de-redundancy: Use clustering algorithms to identify and merge similar feature groups to generate composite fingerprints to reduce redundant information.
[0328] ④ Eliminate obsolete features: For features that exceed the preset time limit and have no matching records, they will be marked as "invalid" and migrated to the archive database for storage.
[0329] It should be noted that the association identification capability of this first embodiment relies on a dark web fingerprint library. This fingerprint library, serving as the core data source for the feature comparison engine, supports multi-dimensional feature matching of target traffic (such as timing watermarks and protocol interaction fingerprints), thereby enabling precise mapping of HS-RP circuits and dark web services. The following will analyze the implementation of the association identification process in conjunction with the architectural design of the fingerprint library.
[0330] In this embodiment S3.2, the real-time update method can ensure that the information in the dark web fingerprint library is always consistent with the current network environment, avoiding misjudgment or missed judgment due to information lag, thereby improving the accuracy of monitoring and identification.
[0331] By extracting the most frequent periods of dark web access, we can precisely locate user activity windows, avoiding wasting resources on inactive periods during data collection. Clustering algorithms automatically group similar traffic data together, allowing for rapid identification of distinct user behavior patterns and improving identification efficiency.
[0332] S3.3. After a preset time interval, obtain the current dark web fingerprint database;
[0333] According to the enhanced index, the first dark web traffic having the same enhanced traffic marking information as the first traffic is obtained from the current dark web fingerprint library, and the website corresponding to the first dark web traffic is defined as the hidden dark web visited by the user; wherein, the enhanced index is obtained by establishing an index containing the enhanced traffic marking information for each dark web traffic record in the dark web fingerprint library; and, the hidden dark web referred to here specifically refers to dark web services in which users use specific means to hide session information, thereby making it difficult for conventional queries to track user access details.
[0334] Detect traffic containing a preset tag information set in the current dark web fingerprint database at a preset period (e.g., every 5 minutes), and define the traffic that successfully matches the set as the traffic to be tested. Traffic containing a preset tag information set refers to traffic that carries any (or all) of the following: tag information (data information obtained through the aforementioned "Begin cell"-based watermark embedding method, the "hash algorithm" watermark embedding method, or time-correlation tagging) or enhanced traffic tag information;
[0335] Score the traffic to be tested according to multi-dimensional association rules to obtain the user's suspiciousness score;
[0336] Traffic with a user suspicion score above a preset threshold in the tested traffic is defined as suspicious traffic. Traffic with the same preset tag information in the suspicious traffic is classified to obtain dark web traffic classification results for the user's suspicious access behavior. The dark web traffic classification results can not only be used to deeply explore dark web sites that the current user may have accessed but that previous methods have not fully identified, but also help network security teams quickly target traffic and users involved in illegal or high-risk activities. In addition, by analyzing these classification results, organizations can gain insight into the types and frequency of users' dark web access, thereby accurately optimizing and adjusting their security policies.
[0337] Among them, multi-dimensional association rules are established based on the number of occurrences of preset tags in dark web traffic, the category of dark web websites, dark web access time, and historical behavior deviation thresholds; the specific dimensions are:
[0338] ① Preset Tag Occurrences: This dimension focuses on the frequency of occurrence of specific preset tags in traffic. For example, if a preset tag appears in a user's traffic more than a certain threshold (e.g., ≥5 times / week) over a period of time (e.g., weekly), this may indicate that the user frequently accesses dark web services or content associated with the tag, increasing the suspiciousness of their behavior.
[0339] ② Darknet website category: This dimension assesses traffic suspicion based on the category of the darknet website. Darknet websites may engage in a variety of illegal or high-risk activities. The system assigns different suspicion weights to traffic based on the website's category. For example, visiting a darknet platform that trades illegal financial instruments may receive a higher suspicion score than visiting a data leak website, as illegal financial instrument trading is often associated with financial crime and fraud.
[0340] ③Dark web access time and historical behavior deviation threshold: This dimension analyzes the time patterns of users' dark web access and compares them with their historical behavior. The system records the time periods when users typically access the dark web and sets a deviation threshold. If a user's visits to the dark web at unusual times (such as late at night) exceed this threshold, this may indicate a significant change in their behavior, increasing their suspicion. For example, if a user typically visits the dark web during the day but suddenly begins frequenting late at night, this may trigger an increase in the system's suspicion score.
[0341] To apply this application example, please refer to Figure 4 , Figure 4 This is a data processing flow chart provided by an embodiment of the present application, which shows the process of correlating and tracing the source based on real-time traffic data and non-real-time traffic data. There are two cases, specifically:
[0342] The first is the real-time processing mode, which is deployed on the traffic collection server and performs the following four operations on the real-time traffic:
[0343] ① Separation: After pre-processing the network traffic, onion service traffic can be separated through three-layer filters or other methods.
[0344] ② Encoding: For the separated traffic, the wavelet leader multifractal form (WLMF) is used to extract multifractal features. Combined with the surface features, the dimensionality is reduced using the t-SNE technology, and then input into an encoder with a multi-head attention layer to obtain traffic embedding.
[0345] ③Classification: Map the traffic embedding to the embedding space, use cosine similarity to calculate the distance to the centroid of the monitored website, and combine the difference between the second closest distance and the closest distance to decide whether to accept the prediction result, thereby distinguishing whether the visit is to a monitored website or other websites.
[0346] ④ Tracing the source: For the traffic accessing the monitoring website, identify the user's other dark web access traffic through the five-tuple and statistical characteristics, use the traffic watermark method to track the user's traffic in and out of the anonymous network, encode the traffic watermark based on the timing, load and other characteristics of the packet, and use the detector to determine the traffic path to achieve the associated backtracing of the hidden service.
[0347] The second is offline detection mode, which sends the traffic to be tested to a pre-deployed server for filtering, encoding, and classification. After discovering traffic to the monitored website, it uses the five-tuple and statistical features to identify the user's other dark web traffic, intercepts it, and sends it to the server for correlation and tracing.
[0348] In this embodiment, S3.3, feature matching is performed through a dark web fingerprint library, which can achieve efficient dark web identification without adding excessive system burden. Compared with the method of in-depth analysis of all traffic, this feature matching method is more lightweight and suitable for deployment and application in large-scale network environments. In addition, the enhanced index is established based on the enhanced traffic tag information, which contains highly unique dynamic watermark tags and multi-protocol tags. This means that each dark web traffic record has a corresponding, unique enhanced traffic tag. Therefore, when it is necessary to retrieve dark web traffic with the same tag information as specific traffic, the enhanced index can provide accurate matching capabilities to ensure the accuracy of the retrieval results.
[0349] Furthermore, by detecting traffic containing preset tag information sets in the dark web fingerprint library, traffic data potentially associated with suspicious user behavior can be accurately located, avoiding the tedious process of filtering large amounts of data required by traditional methods. Multidimensional association rules not only consider the number of occurrences of preset tags in dark web traffic, but also incorporate multiple dimensions such as the category of dark web sites, dark web access time, and historical behavioral deviations, providing a more comprehensive and accurate basis for user suspicion scoring. Existing classification methods may rely more on a single or limited feature dimension for classification. By classifying suspicious traffic with the same preset tag information, we can obtain a classification result for the user's suspicious access behavior, helping network security personnel gain a deeper understanding of the user's suspicious behavior patterns and locate the dark web sites currently visited by the user.
[0350] To apply this application example, please refer to Figure 5 , Figure 5 This is the overall framework diagram provided by the embodiment of this application, showing the process of processing anonymous traffic in this embodiment to trace the dark web; it mainly includes: 1. Compressed sensing of anonymous traffic in a large-scale complex network environment, 2. Multi-granular anonymous user traffic identification, 3. Fine-grained anonymous user behavior identification based on cross-level feature fusion, 4. Dark web tracing based on multi-scale traffic confirmation attack method. These steps are closely corresponding to the above solution, so the specific implementation details are not elaborated in detail. The general framework is summarized as follows:
[0351] "1. Anonymous traffic compression sensing in large-scale complex network environments": Traffic (compressed traffic) is collected from Internet traffic;
[0352] "2. Multi-granularity anonymous user traffic identification": process the traffic collected in step 1 to obtain onion service traffic. It should be noted that Figure 5 The corresponding method shown in is only one possible preset method for extracting onion service traffic from the collected traffic. In addition, various other methods can be used to achieve this goal based on specific application scenarios and requirements. These methods may include different traffic analysis techniques, data processing algorithms, or specific rule sets to more effectively identify and extract onion service traffic to meet the needs of different situations.
[0353] 3. Fine-grained anonymous user behavior recognition based on cross-layer feature fusion: Process the onion service traffic output from step 2 to obtain user traffic (first traffic).
[0354] "4. Dark web tracing based on multi-scale traffic confirmation attack method": Process the user traffic output in step 3 and finally obtain the correlation tracing result; and the image between "Load extraction of traffic watermark encoding" and "Cross-correlation of light and dark web based on flow watermark" in this section is Figure 3 .
[0355] This embodiment S3 is based on the active and passive linkage dark web service association tracking method and stream watermark technology, which realizes the reconstruction of anonymous network links and the tracking and positioning of dark web sites, further improving the accuracy and efficiency of cross-domain anonymous user association.
[0356] It should be noted that in network communications, user network activity is typically conducted in the form of sessions. During a session, a user's network requests and responses form a collection of related traffic. Therefore, even though each request generates new traffic, this traffic still belongs to the same session. When a user begins a new session (for example, visiting a monitored clearnet website), the initial traffic of the session (i.e., the "first traffic") is captured and marked. This mark (e.g., the watermark embedded and time-correlatedly marked in the "first traffic" in this embodiment to obtain the traffic watermark) is unique and represents the specific attributes of the session. Importantly, this mark is applied to all subsequent traffic in the session until the session ends. On the backend, a traffic correlation algorithm tracks these marks. When a user continues to visit other websites (whether on the clearnet or darknet) within the same session, the traffic generated will carry the same mark. Therefore, even if these traffic flows are physically independent, the correlation algorithm in this embodiment can identify them as part of the same session through the mark, thereby achieving cross-domain anonymous user correlation.
[0357] Overall, this application has the following beneficial effects:
[0358] This application performs user behavior identification on the first network traffic on the Internet, and uses a classification prediction method based on a set confidence level to effectively distinguish the traffic of users visiting monitored websites. This method can improve the accuracy and robustness of behavior identification through strict confidence judgment. Embedding a watermark into the first traffic is equivalent to labeling the traffic with a unique label. This watermark technology is concealed and not easily tampered with or removed, providing a reliable basis for subsequent tracking and identification. Traffic is associated with a specific session through time-related tagging, which means that even if the user switches websites or uses different network paths in subsequent visits, as long as the session remains active, the traffic watermark can continue to work. Moreover, as a cross-domain identifier, watermarks can maintain consistency across different network domains and access paths, thereby supporting cross-domain access behavior association; time correlation tags further help determine the temporal sequence and logical relationship of these associated behaviors. This combined tagging method can enhance recognition accuracy and improve tracking continuity; therefore, once the watermark detector identifies that user traffic containing traffic watermarks has visited a dark web website, the information in the watermark can be used to effectively associate the dark web access behavior with previously monitored website access behavior. This method breaks down data silos and enables cross-domain tracking. In addition, the traffic watermark detector is deployed on a controlled guard node at the entrance and exit of the Tor network. The Tor network is one of the main channels for accessing the dark web, so this deployment location is of strategic significance. The watermark detector can cover a large amount of potential dark web access traffic, improving the efficiency of monitoring and identification.
[0359] In summary, this application proposes a traffic compression sensing algorithm for high-speed network environments, which efficiently and collaboratively processes traffic in multiple autonomous domain backbone networks, solves the scalability problem of flow analysis, and adapts to the challenge of large data volumes. For anonymous user traffic, fine-grained flow analysis technology is used to process encrypted and obfuscated traffic in real time and distinguish between dark web and open web access. By constructing multi-level and multi-dimensional traffic fingerprint features, the accuracy of the traffic correlation algorithm is improved, overcoming problems such as missing attributes and signal attenuation. In terms of cross-domain anonymous user correlation, an correlation technology that adapts to large-scale anonymous networks is provided, which realizes anonymous network link reconstruction and dark web site tracking, and enhances network security analysis capabilities.
[0360] Example 2:
[0361] See also Figure 6 , an embodiment of the present application provides a multi-granularity device for identifying associations between user behaviors on the bright and dark webs, comprising a monitoring module 10, a marking module 20, and an identification module 30;
[0362] The monitoring module 10 is used to obtain network traffic from the Internet according to a preset period, classify and predict the network traffic according to a set confidence level, and obtain the first traffic of users visiting the monitored website;
[0363] The marking module 20 is configured to perform watermark embedding and time correlation marking on the first flow to obtain a flow watermark; wherein the session corresponding to the marked first flow is used for the user to perform all subsequent website access behaviors;
[0364] The identification module 30 is used to define the dark web website accessed by the user traffic as the dark web accessed by the user if the traffic watermark detector identifies the user traffic containing the traffic watermark within a preset time period; wherein the traffic watermark detector is deployed on the controlled guard node at the entrance and exit of the Tor network.
[0365] In one embodiment, the monitoring module 10 includes a screening subunit, a compression subunit, an onion subunit, a first feature subunit, a second feature subunit, a third feature subunit, and a prediction subunit; wherein the screening subunit and the compression subunit are processes for performing compressed sensing processing on data, the onion subunit and the first feature subunit are processes for performing wavelet analysis on data, the second feature subunit is a process for screening data based on the amount of leakage, the third feature subunit is a process for reducing the dimension of data based on the t-SNE technology, and the prediction subunit is a process for screening the traffic visiting the monitored website;
[0366] The screening subunit is used to legally obtain anonymous network traffic to be tested by cooperating with law enforcement agencies, using public data sets and platforms, and participating in legitimate cybersecurity research projects.
[0367] The screening subunit is also used to automatically extract features from anonymous network traffic using methods such as deep neural networks (DNNs) to obtain several feature vectors. These feature vectors are then screened using mainstream machine learning models and algorithms such as k-NN and SVM to identify feature vectors that can effectively characterize anonymous network traffic, thereby forming a feature vector set. k-NN is a non-parametric, supervised learning classifier, and SVM is a support vector machine.
[0368] The screening subunit is further configured to further screen out feature vectors that satisfy a k-sparse signal from the feature vector set based on a k-sparse constraint condition, where these feature vectors are sparse feature vectors; construct an anonymous traffic feature set based on the sparse feature vectors obtained by screening to characterize the anonymous network traffic; wherein the k-sparse constraint condition is established by analyzing the sparse distribution characteristics of historical anonymous traffic data;
[0369] The compression subunit is used to perform compressed sensing on the anonymous traffic feature set according to the sensing matrix to obtain compressed traffic;
[0370] The sensing matrix is obtained by mapping the generating matrix according to a preset method. The generating matrix is established according to the selected error correction code. This embodiment uses the classic error correction code - BCH code as an example to explain the specific generation and application process of the sensing matrix. The specific steps are as follows:
[0371] Step 1: Select an error-correcting code. First, a suitable error-correcting code must be selected. Given that BCH codes have a well-defined algebraic structure and meet the RIP (Restricted Isometric Property) requirement, this embodiment uses BCH codes as the basis for constructing the sensing matrix.
[0372] Step 2: Construct the perception matrix Φ. Next, take the BCH code as an example to construct the generator matrix, which will be used to generate the perception matrix Φ required for projection. The specific construction process is as follows: First, select the appropriate parameters of the BCH code, including the code length n, the number of information symbols k, and the error correction capability t; then, construct the generator matrix of the BCH code based on these parameters. , where each element is filled according to the coding rules of BCH code; finally, the bipolar mapping method is used to generate the matrix Convert to the perception matrix Φ, that is, map each 0 element in the matrix to " ", each 1 element in the matrix is mapped to " ";in, is the dimension of the perception matrix.
[0373] Step 3: Apply to anonymous traffic feature compression. Once the perception matrix Φ is constructed, it can be applied to anonymous traffic feature compression. Assume that given an N-dimensional sparse signal , through the perception matrix Map it to M-dimensional space to obtain a compressed M-dimensional feature vector This process can be described by the following mathematical expression:
[0374]
[0375] in, is the flow characteristic before compression, is the flow characteristic after compression, It is a perception matrix generated based on the error correction code.
[0376] It is worth noting that in addition to the BCH code-based method, there are other methods for generating perception matrices, such as Gaussian random projection, etc. These methods are similar to the BCH code method in principle, but differ in their specific implementation.
[0377] The screening subunit and compression subunit of this embodiment can effectively reduce the amount of network traffic data through sparse signal screening and compressed sensing processing. When identifying user behavior, the introduction of a confidence threshold can ensure the reliability and stability of the recognition results; only when the confidence level of the recognition result reaches or exceeds the set threshold will it be considered a valid user behavior, which helps to reduce false positives and false negatives.
[0378] Furthermore, the k-sparsity constraint requires that no more than k non-zero elements be selected from the dataset, which are considered the most representative. This selection process avoids considering all possible feature combinations, greatly improving the efficiency of feature extraction. Using compressed sensing technology combined with a specific perception matrix to compress anonymous traffic feature sets can significantly reduce the amount of data while retaining sufficient information for subsequent user behavior identification or traffic analysis. Furthermore, by reducing the amount of data and optimizing the feature extraction process, this process can save computing resources and increase processing speed, especially when processing large-scale network traffic data.
[0379] In addition, while previous studies have proposed various linear projection methods, this embodiment proposes a deterministic projection technique based on error-correcting codes. Based on this, it explores strategies for generating sensing matrix elements and studies the intrinsic relationship between the dimension of the compressed sensing matrix and signal sparsity and dimensionality. This effectively extracts useful information from traffic while minimizing the number of traffic samples, achieving efficient information extraction and traffic compression, thereby enabling efficient traffic analysis and processing in high-speed network environments. Furthermore, it ensures that the compressed data retains the key features of the original signal, improving the accuracy and scalability of traffic analysis.
[0380] The onion subunit is used to extract onion service traffic from the compressed traffic according to a preset method; the preset method may be based on port and protocol identification, deep packet inspection (DPI), traffic feature analysis, behavioral pattern recognition, and machine learning models;
[0381] The first feature subunit is used to perform wavelet analysis on the onion service traffic to obtain multifractal features. The specific processing method is as follows:
[0382] ① Basic function definition: Let is a basic function with compact support; the reconstruction error is minimized by L1 regularization to generate an adaptive wavelet basis function that matches the Tor traffic characteristics, satisfying the following conditions: and .
[0383] ② Traffic data preprocessing: Perform preprocessing on the onion service traffic, including formalization, segmentation and normalization, and assume that the traffic is a time series , The specific processing method is as follows: Convert the time series of traffic (expressed as a sequence of encrypted data packets) into a discrete time series to meet the requirements of wavelet analysis; Segment the traffic based on the time interval between packet arrivals or a set fixed window length (for example, every 1000 data packets as a group) to cope with the possible interference caused by multi-label mixed traffic; Use methods such as Z-score normalization to eliminate the scale differences between traffic data and prevent noise from having an adverse impact on the multifractal spectrum analysis. Among them, Z-score normalization, also known as standard score normalization or zero-mean normalization, is a commonly used data preprocessing technique.
[0384] ③ Discrete wavelet transform: Use an adaptive wavelet basis function to perform a discrete wavelet transform on the preprocessed traffic data to obtain wavelet transform coefficients ; Among them, , is the discrete input signal (traffic data) after preprocessing, usually expressed as a time series; is an integer, representing the dilation scale of the wavelet basis function, and the scale is inversely proportional to the frequency; is an integer, representing the translation position of the wavelet basis function on the time axis. and mentioned later all follow this definition.
[0385] ④ Interval and neighborhood definition: Define the binary interval (dyadic interval) as to lay the foundation for subsequent multifractal analysis; Among them, ;
[0386] Define the traditional three-neighborhood , and expand it into a cross-scale mixed neighborhood for the multi-label confusion characteristics of anonymous traffic: to capture the cross-label correlation features; Among them, .
[0387] ⑤ Feature extraction: At all finer scales, in the case, calculate the local supremum of the wavelet coefficients located within the spatial neighborhood; Among them, ; And there exists j′ < j, representing all wavelet coefficients at scale j′; is a variable used to index the wavelet coefficients for the k value at a given scale j′;
[0388] Structure function definition: Based on the local supremum of the wavelet coefficients, define the structure function to extract signal features; Among them, , q is a real number (positive or negative), which is used to adjust the sensitivity of the structure function to wavelet coefficients of different amplitudes. Its core function is to quantify the multi-scale statistical characteristics of the signal in different amplitude ranges;
[0389] Definition and calculation of scaling index: In view of the non-stationary characteristics of Tor traffic, such as the existence of extreme values, the scaling index Define and Truncation is performed according to quantiles (such as 95%) to avoid outliers dominating the scaling index estimate; .
[0390] ⑥ Multifractal spectrum calculation: The multifractal spectrum of signal X is obtained by Legendre transform of scaling exponent , which is a key step in multifractal analysis and is used to reveal the complexity and self-similarity of traffic data; h is used to measure the local smoothness or roughness of the signal and quantify local singularities. When h > 0, the signal is smooth (low-frequency trend); when h < 0, the signal has sharp mutations (such as sudden peaks).
[0391] According to the above formula and data processing method, wavelet analysis is performed on the onion service traffic to obtain multi-fractal features. These features can reflect the characteristics of Tor traffic that are difficult to directly discover but meaningful, providing strong support for subsequent traffic analysis.
[0392] In this embodiment, the onion sub-unit and the first feature sub-unit perform feature screening on the onion service traffic to obtain a traffic representation vector. This process can focus on the most critical feature information, which helps to make subsequent classification predictions more accurate. By combining advanced technologies such as wavelet analysis and attention mechanism, meaningful features can be automatically extracted from complex network traffic, solving the problem of low accuracy of traffic association algorithms and models caused by factors such as missing attributes, signal attenuation, and traffic jitter.
[0393] The second feature subunit is used to calculate the information leakage of all features in the multifractal feature using the information leakage as a quantitative indicator to obtain an information leakage set; wherein, "using information leakage as a quantitative indicator" means using information leakage to measure the amount of information that the system can learn from a certain feature of the website, and "information leakage" reflects the characteristics of the feature. What it can reveal about the status of the website The amount of information;
[0394] The second feature subunit is further used to sort the information leakage amount of the information leakage amount set according to the data size, and select the top ten features with the largest leakage amount to form the first feature set;
[0395] Among them, the amount of information leakage in the open world scenario , specifically:
[0396]
[0397] in, is a random variable The entropy of is the monitored website, is a characteristic of a particular representation, Represents additional information or context in the open world, is known And consider In the case of Conditional entropy.
[0398] The second feature subunit of this embodiment forms the first feature set by screening high-leakage features, which not only reduces the interference of redundant data on the classifier and encoder, but also ensures that the scheme focuses on the dimensions that are most likely to leak information. It can improve analysis efficiency while enhancing the robustness against attacks, thereby avoiding problems such as overfitting caused by an excessive number of features, increased model complexity, reduced generalization ability, and increased computational cost.
[0399] The third feature subunit is used to calculate the similarity (conditional probability) between each pair of data points in the first feature set in the high-dimensional space (current dimensional space) using the Gaussian kernel function based on the t-SNE technology. That is, it calculates the probability of a point selecting a point as a neighbor, and obtains the first conditional probability set (high-dimensional space similarity matrix). In the low-dimensional space, the probability of a point selecting a point as a neighbor is also expressed in the form of conditional probability, but this time the t distribution is used as the kernel function to obtain the second conditional probability set (low-dimensional space similarity matrix). Among them, the t-SNE technology is a nonlinear dimensionality reduction technology used for high-dimensional data visualization.
[0400] The third feature subunit is also used to minimize the Kullback-Leibler divergence (KL divergence) loss between the first conditional probability set and the second conditional probability set. It continuously adjusts the point positions of the first feature set in the low-dimensional space through optimization algorithms such as gradient descent, so that the distribution of data points in the low-dimensional space retains the relative proximity relationship in the high-dimensional space as much as possible, until convergence is achieved or a preset number of iterations is reached. Finally, the second feature set after dimensionality reduction is obtained;
[0401] The third feature subunit is further used to use the encoder to perform nonlinear transformation on the second feature set to obtain a traffic representation vector, and process the traffic representation vector into a specified format for classifier and encoder input.
[0402] Among them, t-SNE technology uses Gaussian kernel function to calculate each pair of data points and The conditional probability between , indicating a point Select Point The probability of being a neighbor can be expressed as:
[0403]
[0404] In low-dimensional space, the points are calculated according to the t-SNE technique. and The conditional probability between , which can be expressed as:
[0405]
[0406] The goal of the t-SNE technique is to minimize the Kullback-Leibler divergence (KL divergence) between high-dimensional and low-dimensional data. Its cost function is for:
[0407]
[0408] Among them, the data points Refers to the original sample (such as eigenvector) in the high-dimensional space to be analyzed, and calculates the point Refers to the mapping points in the low-dimensional space generated by the optimization algorithm, represents the kth sample point in the original high-dimensional data, It is the low-dimensional representation (usually 2D or 3D) of x subscript k after t-SNE mapping; It is the joint probability distribution of similarity between the original high-dimensional data points, calculated by the Gaussian kernel function; It is the joint probability distribution of similarity between low-dimensional mapping points (calculation point y), calculated by t distribution (heavy-tailed distribution); and It is the specific probability value when i and j are determined, and i and j are the indexes used to traverse all data point pairs.
[0409] The encoder is constructed by combining the embedding vector generation method of the multi-head attention mechanism, the residual connection method and the multi-layer perceptron, and the encoder has been fully trained on the training set;
[0410] The encoder includes an input layer, a multi-head attention layer, a subsequent processing layer, and a residual connection. The input layer is used to receive the preprocessed flow representation vector. The multi-head attention layer is used to process the input vector with a self-attention mechanism to extract internal correlation information. The subsequent processing layer includes a multi-layer perceptron (MLP) layer, a 2D convolution layer, a 2D maximum pooling layer, a 1D convolution layer, a 1D maximum pooling layer, and an adaptive maximum convolution pooling layer, which are used to further extract features and transform the output of the multi-head attention layer. The residual connection is used to alleviate the gradient vanishing problem in deep networks and ensure smooth information transmission in the network. Specifically:
[0411] ① Multi-head attention layer (MHSA):
[0412] Composition: This layer contains multiple independent attention heads, each of which can process the input data independently, thereby capturing different representations of the input data in parallel;
[0413] Formula Usage:
[0414] Formula 1: This formula is used to define the output of the multi-head attention layer. It is a learnable weight matrix used to linearly combine the outputs of multiple attention heads to obtain the final output vector; is the number of heads, which determines the number of attention heads processed in parallel; It is The output of an attention head, It means concatenating the outputs of all attention heads along the feature dimension to form a longer vector.
[0415] Second formula: For each attention head, this formula calculates the linear transformation of query (Q), key (K) and value (V). 、 and It is aimed at The learnable weight matrix of each head, 、 and They are query, key and value, Represents the attention mechanism.
[0416] The third formula: This formula is used to calculate the output of the attention mechanism. is the dimension of the key vector, used to scale the dot product to avoid excessive values; Represents the transpose operation of the matrix;
[0417] Among them, the first formula is:
[0418]
[0419] The second formula is:
[0420]
[0421] The third formula is:
[0422]
[0423] ②Multi-layer Perceptron (MLP) layer:
[0424] Composition: This layer is a fully connected neural network, usually containing multiple hidden layers, used to further process the output of the multi-head attention layer.
[0425] Purpose: Perform nonlinear transformation on the output of the multi-head attention layer to extract higher-level features.
[0426] ③Convolutional Layers:
[0427] Composition: The encoder consists of 2D convolutional layers and 1D convolutional layers. The convolutional layers use convolution kernels to slide over the input data and extract local features through convolution operations.
[0428] use:
[0429] 2D Convolutional Layer (2D Conv): processes data with a two-dimensional structure (such as images) and can capture the spatial characteristics of the data.
[0430] 1D Convolutional Layer (1D Conv): processes one-dimensional data (such as time series data) and extracts local patterns in the time series.
[0431] ④Pooling Layers:
[0432] Composition: The encoder contains a 2D max pooling layer and a 1D max pooling layer.
[0433] Purpose: Reduce the spatial size or temporal length of data by downsampling while retaining important features and reducing the amount of computation and the number of parameters.
[0434] 2D Max Pooling: Slide a window on the two-dimensional feature map and select the maximum value in the window as the output. It is often used to reduce the spatial dimension.
[0435] 1D Max Pooling: Slide a window on a one-dimensional feature sequence and select the maximum value in the window as the output. It is often used to reduce the time dimension.
[0436] ⑤Adaptive Max Pooling:
[0437] Composition: This is a special pooling layer whose output size is fixed.
[0438] Purpose: Dynamically adjust the size of the pooling window according to the size of the input data to ensure that the output has a fixed size for subsequent processing.
[0439] ⑥Residual connection
[0440] Composition: Directly add input data to the output of certain layers through skip connections.
[0441] Formula Purpose: The fourth formula illustrates the role of residual connections. Here, F(x) is the output after a series of nonlinear transformations, and x is the input. Residual connections add the input x directly to F(x), helping to alleviate the vanishing gradient problem in deep networks.
[0442] Purpose: Improve the training effect of deep networks, allowing information to flow more smoothly in the network, and improving the training efficiency and performance of the model.
[0443] Among them, the fourth formula is:
[0444]
[0445] To apply this application example, please refer to Figure 2 , Figure 2 This is a flow representation generation diagram provided in an embodiment of the present application, which shows the process of performing nonlinear transformation on the second feature set according to the encoder to generate a flow representation vector.
[0446] The prediction subunit is used to calculate the distance between each traffic embedding and the centroid of the monitored website set in the traffic representation vector using cosine similarity to obtain several distance sets;
[0447] The prediction subunit is further configured to: for the first traffic embedding in the distance set corresponding to the plurality of distance sets, if the nearest centroid distance is less than the radius of the website corresponding to the nearest centroid (and this condition can be defined as "condition 1", that is, the distance between the traffic embedding and the website centroid is less than the radius of the website), and the difference between the second nearest centroid distance and the nearest centroid distance is greater than a preset threshold (Also, this condition can be defined as "Condition 2", that is, the difference between the second closest distance and the closest distance exceeds the preset threshold ), the website corresponding to the nearest centroid is determined to be the first monitored website visited by the user, and the traffic of the user visiting the first monitored website is defined as the first traffic of the user visiting the monitored website; and, according to this detection method, all traffic embeddings in the traffic representation vector are traversed, so that the first traffic finally obtained includes traffic from the user visiting several first monitored websites;
[0448] Among them, the radius of the first website is calculated based on the distance between the data point and the data mean in the first website using the standard deviation; the nearest centroid distance is the distance between the first traffic embedding and the nearest centroid, and the second nearest centroid distance is the distance between the first traffic embedding and the second nearest centroid; and "monitored websites" refer to those websites whose access traffic has been collected and analyzed in advance; by extracting features from these traffic, the encoder is trained, and the centroids of these websites are determined based on the position of these feature vectors in the embedding space. In contrast, "unmonitored websites" refer to those websites that have not undergone similar traffic collection and training processes, that is, their traffic features have not been used for encoder training; and the classification method based on conditions 1 and 2 can be defined as a classifier.
[0449] For example, suppose there is a traffic sample whose distance from website C is 0.8 units and its distance from the next closest website D is 1.6 units. At the same time, the difference threshold is set to 0.3 units.
[0450] Judgment process: Condition 1: Distance from traffic to website C (0.8) < Radius of website C (1.0) → Satisfied; Here, the distance between the traffic sample and website C is less than the radius of website C, so condition 1 is satisfied. Condition 2: Distance from the next closest website D (1.6) - Distance from website C (0.8) = 0.8 > Threshold (0.3) → Satisfied; Here, the difference between the distance from the traffic sample to the next closest website D and the distance to the closest website C is greater than the preset threshold, so condition 2 is also satisfied.
[0451] Conclusion: Based on the above conditions, it is determined that the traffic belongs to the behavior of visiting website C.
[0452] Among them, assuming that the website The embedding representation of … , where each is a vector, then the centroid formula is:
[0453]
[0454] Use the standard deviation to calculate the distance between the data point and the mean to get the radius of the website , and its calculation formula is:
[0455]
[0456] Use cosine similarity to calculate each flow embedding and website The distance between the centroids is calculated as:
[0457]
[0458] Condition 1 - The distance between the traffic embedding and the centroid of the website is less than the radius, which can be expressed as:
[0459]
[0460] Condition 2 - The difference between the second closest distance and the closest distance exceeds the threshold, which can be expressed as:
[0461]
[0462] in, is the mean of the embeddings, It is data points; " indicates a vector and centroid The dot product of is a vector The Euclidean norm (length) of is the dimension of the embedding vector, is the index variable to be summed, is the feature vector of the data point at a specific j and k; is the Euclidean norm of the centroid; is the distance between the traffic embedding and the second closest website centroid, is the distance between the traffic embedding and the nearest website centroid, is the threshold, is the embedding vector dimension.
[0463] The prediction subunit in this embodiment calculates the distance between each traffic embedding and the centroid of the monitored website set, generating a distance set reflecting the similarity between traffic and each website. By combining absolute and relative distance constraints, a high-confidence classification framework is constructed within the embedding space. Predictions are only accepted when traffic clearly matches the characteristics of the target website, making this particularly suitable for scenarios requiring strict filtering of false positives.
[0464] The monitoring module 10 of this embodiment processes onion service traffic through wavelet analysis to obtain multifractal features. These features can capture the complexity and nonlinear characteristics of traffic data, providing a rich information basis for subsequent analysis. Using the amount of information leakage as a quantitative indicator to screen for high-leakage features in the multifractal features, this process ensures that the selected features are closely related to user privacy and information leakage risks, helping subsequent analysis to focus more on key features and improve the accuracy and efficiency of the analysis. By minimizing the KL divergence loss between the first conditional probability set and the second conditional probability set to perform dimensionality reduction processing, the dimensionality of the data can be significantly reduced while retaining key information.
[0465] In one embodiment, the marking module 20 includes a watermark unit and an enhancement unit. The watermark unit is a process of embedding a watermark and marking time correlation on the first traffic, and the enhancement unit is a process of enhancing traffic marking on the first traffic, specifically:
[0466] The watermark unit is configured to embed a watermark and perform time correlation marking on the first flow to obtain a flow watermark; wherein the session corresponding to the marked first flow is used for the user to perform all subsequent website access behaviors;
[0467] The watermark embedding refers to embedding the watermark into the first traffic according to the pseudo-random number generator and the shared key, specifically:
[0468] Generate a binary sequence according to a pseudo-random number generator and a shared key, and convert the binary sequence into a time interval sequence of Begin cells;
[0469] Begin cells are inserted into the transmission process of the first traffic according to a time interval sequence.
[0470] The time correlation marking refers to performing time correlation marking on the first traffic according to the timestamp, specifically:
[0471] Obtain the timestamp of the first data packet of the first flow, and generate a watermark sequence according to the timestamp of the first data packet;
[0472] The watermark sequence is used as the time fingerprint of the HS-RP circuit and is embedded into all associated traffic by directionally adjusting the transmission delay of adjacent data packets. Associated traffic refers to the traffic generated when users access different services through the same HS-RP circuit. Different services include monitored websites and subsequent dark web services.
[0473] The following is a detailed description of the marking method:
[0474] ① Watermark embedding. Its core strategy is this: during an access request, if the client knows the address of the hidden service, the system uses watermarking technology to secretly transmit the client's known identity (i.e., the hidden service address) to the controlled assembly point. When the hidden service receives a packet containing a Relay-Drop, it will directly ignore and discard it, ensuring that this operation does not cause any interference to upper-layer applications. "Relay-Drop" packets refer to packets that carry watermark signals during the network relay process and are designed to be directly discarded.
[0475] This example uses the encoding of communication cells using Relay-Drop packets (as the transmission medium for watermark signals) as an example to explain the watermark embedding method in detail. Specifically, the Tor protocol's Relay-Drop communication cells are used to covertly encode dark web service addresses, achieving highly concealed watermark embedding and detection. The detailed steps are as follows:
[0476] Preparation phase: A key is shared between the encoder and detector for subsequent watermark generation and detection.
[0477] Watermark Embedding: 1. Begin Cell Sequence Generation: Using a pseudo-random number generator (PRNG) and a shared key, a binary sequence B is generated based on the extended Gilbert model. In sequence B, "1" indicates that a Begin cell needs to be generated, and "0" indicates that it should not be generated. The Gilbert model is a statistical model used to describe data transmission errors in digital communication systems, and the Begin cell is a service response mechanism in the Tor protocol, used to identify and coordinate data transmission during network communication.
[0478] 2. Time Interval Calculation: Convert the binary sequence B into a sequence D of Begin cell time intervals. Specifically, iterate through the binary sequence B. Whenever a "1" is encountered, calculate the corresponding Begin cell transmission time based on the preset time interval parameter Δt. This yields a sequence D of time intervals containing the Begin cell transmission times.
[0479] 3. Watermark encoding: The dark web service address information is encoded in a covert manner into the time interval sequence D of the Begin cell without changing the other contents of the Relay-Drop packet. This ensures that the hidden service will directly discard the Relay-Drop packet upon receiving it, without affecting the upper-layer application.
[0480] Begin cells are inserted into the transmission process of the first flow according to the time interval sequence D. In addition to Begin cells, other service response mechanisms in the Tor protocol can also be combined to enhance the concealment and robustness of the watermark.
[0481] The watermark embedded in the watermark unit of this embodiment relies on a shared key to generate the watermark, ensuring that only legitimate recipients can interpret the time pattern of the Begin cell. Even if an attacker intercepts the traffic, they cannot recover the watermark information. Furthermore, the sequence generated by the pseudo-random number generator is time-sensitive, making it impossible for attackers to reproduce historical sequences, thus preventing watermark forgery through traffic replay.
[0482] In addition to the above-mentioned watermark embedding method based on the "Begin cell", this embodiment also provides a watermark embedding method using the "hash algorithm" as a supplement, as follows:
[0483] Record each data packet in the first flow Time of arrival at the network node , get the time when the first data packet in the traffic arrives , this time can be accurate to milliseconds or even higher precision to ensure its uniqueness;
[0484] Calculating delays between adjacent data packets in the first flow to obtain a first delay sequence;
[0485] Based on the watermark embedder, according to the watermark sequence The watermark bit in the first delay sequence is adjusted to the corresponding delay, and the data packet in the network traffic is sent according to the adjusted delay to complete the watermark embedding;
[0486] Among them, the watermark sequence It is obtained by capturing the arrival time of the first data packet in the network traffic and obtaining the timestamp; then mapping the timestamp to a hash value of a preset fixed length, performing base conversion and sequence division on the hash value, specifically: selecting a suitable hash function (such as SHA-256 hash function) as the basis for generating the watermark sequence; the arrival time of the first data packet in the first flow is used as the input of the hash function to calculate the initial hash value ,in, , It's time, is a hash function, Is the hash value; the initial hash value Convert to a binary sequence (every 4 bits correspond to 1 hexadecimal bit) to get a binary sequence ; Divide the binary sequence into groups of 8 bits , if the decimal value of a set of binary numbers is greater than 127, then let Otherwise, ; This way, a unique watermark sequence can be generated based on the time of the first data packet ;
[0487] In summary, the watermark embedding process is specifically described as follows: for each pair of adjacent data packets and , according to the formula To calculate the delay difference between them, here . Then, traverse the watermark sequence Each watermark bit in , according to the formula The delay of adjacent data packets is adjusted accordingly. If the watermark bit is "1", the next data packet The sending time of the data packet is delayed by several time units; if the watermark bit is "0", the next data packet The sending time is advanced by several time units. By resending the data packet, the watermark embedding process in the network traffic is completed; Indicates timestamp and timestamp the time interval between It is an integer, indicating the delay in several time units.
[0488] In the watermark unit of this embodiment, for the watermark embedding method using the "hash algorithm", the data packet timestamp is used for hash processing to generate the watermark sequence. This method has the advantages of concealment, because the timestamp itself is common in network transmission, and the watermark generated by hashing the timestamp is not easy to detect. At the same time, the hash function has the anti-collision property, which makes the generated watermark robust and able to resist certain attacks and interferences.
[0489] Furthermore, mapping the timestamp to a hash value of a preset fixed length ensures that the generated watermark sequence has a consistent length, which facilitates subsequent processing, storage, and transmission. At the same time, the fixed-length hash value simplifies the watermark embedding and extraction process, improving efficiency. Sequence partitioning allows hash values to be split into multiple subsequences or fragments to form a more complex watermark sequence. This refined control helps enhance the concealment and robustness of the watermark.
[0490] In addition, since the watermark is embedded by adjusting the delay, which is affected by many factors such as network conditions and device performance, the watermark is likely to be retained even in the event of network fluctuations or packet loss. This makes the watermark highly robust and able to resist network attacks and data tampering to a certain extent.
[0491] ②Time correlation mark:
[0492] Get the timestamp t of the first data packet of the first flow 01 , according to the first data packet timestamp t 01 and shared key A, generate watermark sequence through pseudo-random algorithm .
[0493] The watermark sequence W is used as the time fingerprint of the HS-RP circuit and embedded into all associated traffic by directionally adjusting the transmission delay of adjacent data packets. Associated traffic refers to the traffic generated when users access different services through the same HS-RP circuit. Different services include monitored websites and subsequent dark web services. That is:
[0494] In all associated traffic on the HS-RP circuit, the transmission delay of adjacent data packets is adjusted according to the watermark sequence W. For the i-th data packet, the delayed transmission time is w i , so that the interval between it and the previous data packet conforms to the watermark sequence W; let all traffic on the same HS-RP circuit (whether accessing monitoring websites or dark web services) carry the same time fingerprint W, and use the time fingerprint W as the unique time fingerprint of the circuit and bind it to the hidden service (HS) that created the circuit; among them, the HS-RP circuit is a hidden service communication circuit.
[0495] For the time-correlation marker in the watermark unit of this embodiment, a watermark sequence is generated based on the timestamp of the first data packet, ensuring high synchronization between the watermark and the start time of the traffic flow, providing a precise time reference for subsequent traffic analysis and time-correlation tracking. The watermark is embedded by directionally adjusting the transmission delay of adjacent data packets. This method is relatively covert and difficult to detect by ordinary users or malicious attackers, reducing the risk of detection and evasion. In addition, the watermark sequence is used as the time fingerprint of the HS-RP circuit and embedded in all associated traffic, enabling cross-service traffic tracking. Even if the user switches between different services during the access process, the watermark can still exist and function, ensuring the continuity of tracking.
[0496] In summary, the combination of watermarks and timestamps enables efficient tracking and correlation of traffic in subsequent processing, enabling accurate traffic identification and correlation analysis, whether within the same session or across sessions and network domains.
[0497] The enhancement unit is configured to perform an enhancement flow mark on the first flow to obtain enhancement flow mark information; wherein the session corresponding to the marked first flow is used for the user to perform all subsequent website access behaviors;
[0498] Among them, the enhanced traffic marking includes dynamic watermark marking composed of a composite watermark sequence and multi-protocol marking composed of encrypted watermark fragments. The composite watermark sequence is obtained by hashing the timestamp and packet size, and the encrypted watermark fragment is obtained by inserting a custom data field during the handshake process of the transport layer security protocol. The specific marking method is as follows:
[0499] Use a hash algorithm to perform a joint hash process on the timestamp, session ID, and packet size of the first flow to obtain a tamper-resistant composite watermark sequence;
[0500] The composite watermark sequence is divided into a first sequence, a second sequence and a third sequence;
[0501] The first segment of the sequence is embedded into the payload of the data packet, the second segment of the sequence is embedded into the option field of the IP data packet, and the third segment of the sequence is embedded into the unused bits of the TCP header, thereby obtaining a data packet payload with a watermark, an IP option field with a watermark, and a TCP header with a watermark, respectively. The data packet payload with a watermark, the IP option field with a watermark, and the TCP header with a watermark constitute dynamic watermark marking information.
[0502] In the TLS handshake phase, a custom encrypted watermark fragment is injected into the protocol to obtain the TLS handshake information with the encrypted watermark fragment;
[0503] The enhanced traffic marking information is composed of dynamic watermark marking information and TLS handshake information; among them, the TLS handshake phase is the key process in the Transport Layer Security Protocol (TLS) for establishing a secure connection between the client and the server; the session ID is a unique identifier used to identify the session between the user and the server; the IP data packet is the data packet of the network layer in the TCP / IP protocol; the TCP header is the header of the transport layer TCP protocol, which is used to control the transmission of information when transmitting data in the TCP connection; the IP option field is an optional field in the IP header, used to support some specific functions or testing purposes.
[0504] In the enhancement unit of this embodiment, due to the customization of multi-protocol tags and the irreversibility of hash operations, the enhanced traffic tag information is highly unique and concealed. Therefore, even in a complex network environment and huge traffic data, it is still possible to accurately identify the hidden dark web visited by the user.
[0505] In one embodiment, the identification module 30 includes a truncation unit, a circuit unit, an association unit, a fingerprint library unit, and a deepening unit. The truncation unit, the circuit unit, and the association unit are processes for identifying the dark web accessed by the user and its address. The fingerprint library unit is a process for updating the dark web fingerprint library in real time. The deepening unit is a process for identifying the hidden dark web accessed by the user and classifying suspicious traffic to deepen the association identification method. Specifically,
[0506] The truncation unit is configured to, if the traffic watermark detector on the controlled guard node (Guard) does not detect user traffic containing a traffic watermark within a preset time period (indicating that the current hidden service circuit has not selected a controlled guard node), then truncate the user traffic from the unselected controlled node, triggering a circuit anomaly and forcing the system to resend traffic to select a new guard node until the user traffic passes through the controlled guard node. The controlled guard node is deployed at the entrance and exit of the Tor network. The traffic watermark detector is responsible for detecting watermarks. It extracts possible watermark bits by analyzing the characteristics of the traffic data packets of the selected carrier and compares the calculated characteristic function value with the preset watermark parameters to determine whether the data contains a watermark.
[0507] The circuit unit is configured to control the Tor client to establish a three-hop circuit based on the guard nodes, intermediate nodes, and exit nodes of the Tor network at a default frequency of every 10 minutes if the traffic watermark detector on the controlled guard node detects user traffic containing the traffic watermark (indicating that the controlled guard node has been selected for the current hidden service circuit). The three-hop circuit is used to encrypt and anonymize communications between the user and the hidden service.
[0508] The association unit is used to determine the dark web website and the IP address of the dark web website visited by the user traffic based on the traffic watermark detector on the relay node if the watermark detector detects that the exit node passed by the three-hop circuit is a controlled relay node, and use the dark web website and its IP address as the dark web target information visited by the user; if the exit node passed by the three-hop circuit is not a controlled relay node, the circuit needs to be destroyed immediately and this process is repeated until the selected exit node is one of the controlled "honey relays".
[0509] In addition, the controlled guard nodes include controlled entry nodes and controlled exit nodes; the controlled entry nodes are deployed with traffic watermark generators, and the controlled exit nodes are deployed with traffic watermark detectors; based on the traffic watermark generator and the traffic watermark detector, the dark web websites visited by the same user can be identified.
[0510] Within the set maximum time range, if the traffic watermark detector is still unable to identify user traffic containing traffic watermarks, it means that the user who visited the monitored website did not visit the dark web.
[0511] The following is a detailed description of the detection method of the traffic watermark detector:
[0512] ① Detection method for watermark embedding (watermark embedding method based on "Begin cell"):
[0513] First, a detector is used to capture Begin cells in the network flow of the dark web fingerprint library and record their respective timestamps;
[0514] Then, the time interval (IPD) between each pair of adjacent Begin cells is calculated, that is, the time difference between each Begin cell and its previous cell;
[0515] Finally, the detector uses a similarity algorithm to carefully check the time interval of each Begin cell to determine whether it matches the Begin cell sequence pre-generated by the encoder. If the time interval sequence is highly similar or completely matches the Begin cell sequence generated by the encoder, the sequence is considered to be an embedded stream watermark. Based on this discovery, the detector can further analyze the information carried by the watermark, such as data source identifier, timestamp, user identifier, etc., to effectively track and accurately associate hidden services.
[0516] In addition, here is a description of the detection method of the watermark embedding method using the "hash algorithm":
[0517] (1) Using a network packet capture tool, obtain the second traffic that may contain the watermark from the Internet according to a preset period, and record the time when each data packet arrives;
[0518] Calculate the delay between adjacent data packets in the second flow to obtain the second delay sequence ;in, Indicates the delay between adjacent data packets containing watermark information during the watermark extraction period ;
[0519] Calculate the preset third time delay sequence and second time delay sequence according to the Pearson correlation coefficient formula The correlation between several time delay sequences in the , and the correlation set is obtained; wherein the third time delay sequence is the time delay sequence corresponding to the known embedded watermark sequence , among which Indicates the delay between adjacent data packets containing watermark information during the watermark embedding period ;
[0520] The correlation set is greater than the preset threshold T (i.e. >T) is determined as the first flow after the watermark sequence is embedded.
[0521] Among them, the Pearson correlation coefficient formula is:
[0522]
[0523] Here, x and y are the two variables to be analyzed (such as "time" in traffic data), and n is the number of observations.
[0524] In this embodiment of the watermark embedding detection method using a hash algorithm, correlation calculation can quantify the degree of similarity between the third delay sequence and the first delay sequence. By setting a preset threshold, traffic containing the watermark can be accurately screened. This method reduces the possibility of false positives and false negatives, thereby improving detection accuracy.
[0525] ② Detection method for time-related markers:
[0526] Monitor the traffic of the deployed monitoring website and extract the timestamp t of the first data packet of the visiting user 02 ;
[0527] Using the shared key A and the first packet timestamp t 02 Regenerate the watermark sequence W';
[0528] Detect the user's subsequent traffic to the dark web service through the same HS-RP circuit, extract the adjacent data packet interval sequence, and verify whether it matches the watermark sequence W'. If the match is successful, it is determined that the dark web service access traffic and the monitored website access traffic belong to the same user.
[0529] In addition, the specific comparative identification methods include active linkage and passive linkage user behavior identification methods, specifically:
[0530] ① Active linkage: This system actively controls a certain number of routing nodes and waits for them to be selected by the hidden service as its guard nodes (i.e., entry nodes). Once selected, these nodes can communicate directly with the hidden service, revealing its true IP address. Furthermore, even if a user subsequently accesses other dark web services through the same set of controlled nodes, the system can accurately determine that the traffic originates from the same user by detecting the same timing patterns in the communications. Specifically, when a user first visits a monitored website, their HS-RP circuit is marked with a specific timing signature as described above. If the user subsequently accesses another dark web service through the same HS-RP circuit and the circuit is again marked with the same timing signature, it can be concluded that the two visits belong to the same user, enabling cross-site correlation and tracing. For example, if a user first visits a monitored website (and is marked), and then later visits a dark web forum through the same HS-RP circuit, the system can correlate the two by comparing the timing patterns, thereby determining the user's identity or true IP address, as well as the dark web sites and dark web addresses visited.
[0531] ② Passive linkage: When a hidden service matching a signature in the dark web fingerprint database doesn't select a system-controlled routing node as its guard node, the system adopts a passive strategy, frequently creating links in an attempt to be selected by the hidden service as the second hop. This way, even if the system can't directly obtain the hidden service's true IP address, it can indirectly locate the hidden service by tracing its binding relationship with the guard node. This passive linkage strategy provides the system with a fallback in the event that active linkage fails, enhancing the flexibility and reliability of tracking and associating hidden services.
[0532] To apply this application example, please refer to Figure 3 , Figure 3 It is an association traceability diagram provided by an embodiment of the present application, which shows the process of adding traffic watermarks to traffic according to the watermark embedder, detecting traffic watermarks according to the watermark detector, and establishing a three-hop circuit for association traceability on this basis.
[0533] In this embodiment, the interception unit, circuit unit, and association unit intercept and reroute traffic when no watermark is detected, ensuring that all target traffic passes through the controlled guard node, thereby improving monitoring coverage and accuracy. When the exit node of the three-hop circuit is a controlled relay node, the traffic watermark detector on the relay node can be used to accurately locate the dark web website visited by the user and its IP address. This refined monitoring capability provides valuable intelligence support for network security and law enforcement agencies.
[0534] Furthermore, due to the unique nature of the Tor network, user access behavior often spans multiple network domains and different access paths. By deploying appropriate watermark processing equipment at the entry and exit nodes, cross-domain traffic correlation can be supported. Even if a user switches between different network paths or uses different hidden services during access, as long as the traffic carries the same watermark, it can be accurately correlated and identified by the system.
[0535] The fingerprint library unit is used to update the dark web fingerprint library in real time based on the dark web and its real-time access information.
[0536] The Dark Web Fingerprint Library is a database dedicated to collecting, organizing, and analyzing the characteristics of dark web services. It aims to obtain dark web data resources through legal and compliant channels, such as collaborating with law enforcement agencies, using public search tools, and participating in industry research projects. An exemplary construction method for the Dark Web Fingerprint Library is as follows:
[0537] Targeting and data collection: Identify the types of services and activity areas that the dark web fingerprint database should cover, focusing on specific areas such as illegal transactions and the circulation of cyberattack tools. Obtain dark web data resources through legal and compliant channels, including collaborating with law enforcement agencies, using public search tools, and participating in industry research projects. Focus on collecting web content, transaction records, and user interaction information from dark web forums, trading platforms, and instant messaging spaces.
[0538] Data content analysis: Deeply analyze the collected network traffic to extract statistical features such as packet size, transmission interval, and traffic rate. Analyze the structure of dark web pages and extract key identifiers such as page elements, dynamic scripts, and image hash values.
[0539] Time and space dimension tagging and classification system: Combine traffic transmission timestamps and circuit path information to tag dark web traffic in time and space dimensions. Establish a classification system based on the nature of the service, set business tags such as illegal drug transactions and hacking tool transactions, and achieve refined classification of service types;
[0540] Database Design and Data Storage: Design a multidimensional storage architecture to optimize table structure, field type matching, and retrieval efficiency. Fully annotated dark web data will be systematically stored according to the classification system to form a scalable dark web service feature library. This database supports functions such as time series analysis, path tracing, and behavioral pattern recognition, providing underlying data support for subsequent security monitoring.
[0541] Among them, the update method of the dark web fingerprint library is:
[0542] Using Tor network monitoring tools or dark web traffic analysis systems, extract traffic data of users accessing the dark web according to high-frequency time periods to obtain comprehensive traffic data. The collected data is then cleaned to remove invalid data such as outliers and duplicates. The high-frequency time period is the time window when the number of users accessing the dark web exceeds the preset number of dark web visits. In addition, by introducing a dynamic adjustment mechanism, the high-frequency time period is recalculated regularly (such as daily or weekly) to adapt to changes in user access behavior.
[0543] In the comprehensive traffic data, the target features of users accessing the dark web are extracted to obtain the target feature set; the user access behavior pattern is identified on the comprehensive traffic data based on the clustering algorithm to obtain the behavior pattern recognition result;
[0544] The fingerprint feature set is composed of the target feature set and the behavior pattern recognition result;
[0545] The fingerprint feature set is stored in the dark web fingerprint database according to the preset update mechanism to obtain an updated dark web fingerprint database; wherein the update mechanism specifically includes the following steps:
[0546] ① Add new features to the database: When a new dark web service or an unknown variant of its existing service is detected, these new features (such as API call patterns) are included in the dark web fingerprint library, and a unique ID is assigned to each new feature; where "ID" is the abbreviation of "Identifier" and is used to uniquely identify or distinguish each new feature.
[0547] ② Dynamic weight adjustment: In the event that old fingerprints become invalid due to data updates or other reasons, the weight of the features will be adjusted according to their activity. Specifically, the weight of highly active features will be increased so that they are detected first; at the same time, the weight of low-activity features or features with high false positive rates will be reduced and moved to the historical archive.
[0548] ③ Feature merging and de-redundancy: Use clustering algorithms to identify and merge similar feature groups to generate composite fingerprints to reduce redundant information.
[0549] ④ Eliminate obsolete features: For features that exceed the preset time limit and have no matching records, they will be marked as "invalid" and migrated to the archive database for storage.
[0550] It should be noted that the association identification capability of this second embodiment relies on a dark web fingerprint library. This fingerprint library, serving as the core data source for the feature matching engine, supports multi-dimensional feature matching of target traffic (such as timing watermarks and protocol interaction fingerprints), thereby enabling precise mapping of HS-RP circuits and dark web services. The following will analyze the further implementation of the association identification process, combining the architectural design of the fingerprint library.
[0551] In the fingerprint library unit of this embodiment, the real-time update method can ensure that the information in the dark web fingerprint library is always consistent with the current network environment, avoiding misjudgment or missed judgment due to information lag, thereby improving the accuracy of monitoring and identification.
[0552] By extracting the most frequent periods of dark web access, we can precisely locate user activity windows, avoiding wasting resources on inactive periods during data collection. Clustering algorithms automatically group similar traffic data together, allowing for rapid identification of distinct user behavior patterns and improving identification efficiency.
[0553] A deepening unit, used to obtain the current dark web fingerprint database after a preset time interval;
[0554] The deepening unit is further used to obtain the first dark web traffic having the same enhanced traffic marking information as the first traffic from the current dark web fingerprint library according to the enhanced index, and define the website corresponding to the first dark web traffic as the hidden dark web visited by the user; wherein, the enhanced index is obtained by establishing an index containing the enhanced traffic marking information for each dark web traffic record in the dark web fingerprint library; and, the hidden dark web referred to here specifically refers to dark web services in which users use specific means to hide session information, thereby making it difficult for conventional queries to track user access details.
[0555] The deepening unit is further configured to detect traffic containing a preset tag information set in the current dark web fingerprint database at a preset period (e.g., every 5 minutes), and define traffic that successfully matches as traffic to be tested; wherein, traffic containing a preset tag information set refers to traffic carrying any one (or all) of the tag information (data information obtained by the aforementioned watermark embedding method based on the "Begin cell", the watermark embedding method using the "hash algorithm", and time-correlation tagging) or enhanced traffic tag information;
[0556] The deepening unit is also used to score the traffic to be tested according to multi-dimensional association rules to obtain the user suspicion score;
[0557] The deepening unit is also used to define traffic with a user suspicion score above a preset threshold as suspicious traffic, classify suspicious traffic with the same preset tag information, and obtain dark web traffic classification results for suspicious user access behavior. The dark web traffic classification results can not only be used to deeply explore dark web sites that the current user may access but that previous methods have not fully identified, but also help network security teams quickly target traffic and users involved in illegal or high-risk activities. In addition, by analyzing these classification results, organizations can gain insight into the types and frequency of users' access to the dark web, thereby accurately optimizing and adjusting their security policies.
[0558] Among them, multi-dimensional association rules are established based on the number of occurrences of preset tags in dark web traffic, the category of dark web websites, dark web access time, and historical behavior deviation thresholds; the specific dimensions are:
[0559] ① Preset Tag Occurrences: This dimension focuses on the frequency of occurrence of specific preset tags in traffic. For example, if a preset tag appears in a user's traffic more than a certain threshold (e.g., ≥5 times / week) over a period of time (e.g., weekly), this may indicate that the user frequently accesses dark web services or content associated with the tag, increasing the suspiciousness of their behavior.
[0560] ② Darknet website category: This dimension assesses traffic suspicion based on the category of the darknet website. Darknet websites may engage in a variety of illegal or high-risk activities. The system assigns different suspicion weights to traffic based on the website's category. For example, visiting a darknet platform that trades illegal financial instruments may receive a higher suspicion score than visiting a data leak website, as illegal financial instrument trading is often associated with financial crime and fraud.
[0561] ③Dark web access time and historical behavior deviation threshold: This dimension analyzes the time patterns of users' dark web access and compares them with their historical behavior. The system records the time periods when users typically access the dark web and sets a deviation threshold. If a user's visits to the dark web at unusual times (such as late at night) exceed this threshold, this may indicate a significant change in their behavior, increasing their suspicion. For example, if a user typically visits the dark web during the day but suddenly begins frequenting late at night, this may trigger an increase in the system's suspicion score.
[0562] To apply this application example, please refer to Figure 4 , Figure 4 This is a data processing flow chart provided by an embodiment of the present application, which shows the process of correlating and tracing the source based on real-time traffic data and non-real-time traffic data. There are two cases, specifically:
[0563] The first is the real-time processing mode, which is deployed on the traffic collection server and performs the following four operations on the real-time traffic:
[0564] ① Separation: After pre-processing the network traffic, onion service traffic can be separated through three-layer filters or other methods.
[0565] ② Encoding: For the separated traffic, the wavelet leader multifractal form (WLMF) is used to extract multifractal features. Combined with the surface features, the dimensionality is reduced using the t-SNE technology, and then input into an encoder with a multi-head attention layer to obtain traffic embedding.
[0566] ③Classification: Map the traffic embedding to the embedding space, use cosine similarity to calculate the distance to the centroid of the monitored website, and combine the difference between the second closest distance and the closest distance to decide whether to accept the prediction result, thereby distinguishing whether the visit is to a monitored website or other websites.
[0567] ④ Tracing the source: For the traffic accessing the monitoring website, identify the user's other dark web access traffic through the five-tuple and statistical characteristics, use the traffic watermark method to track the user's traffic in and out of the anonymous network, encode the traffic watermark based on the timing, load and other characteristics of the packet, and use the detector to determine the traffic path to achieve the associated backtracing of the hidden service.
[0568] The second is offline detection mode, which sends the traffic to be tested to a pre-deployed server for filtering, encoding, and classification. After discovering traffic to the monitored website, it uses the five-tuple and statistical features to identify the user's other dark web traffic, intercepts it, and sends it to the server for correlation and tracing.
[0569] The deepening unit of this embodiment performs feature matching through the dark web fingerprint library, which can achieve efficient dark web identification without adding excessive system burden. Compared with the method of performing in-depth analysis of all traffic, this feature matching method is more lightweight and suitable for deployment and application in large-scale network environments. In addition, the enhanced index is established based on the enhanced traffic tag information, which contains highly unique dynamic watermark tags and multi-protocol tags. This means that each dark web traffic record has a corresponding, unique enhanced traffic tag. Therefore, when it is necessary to retrieve dark web traffic with the same tag information as specific traffic, the enhanced index can provide accurate matching capabilities to ensure the accuracy of the retrieval results.
[0570] Furthermore, by detecting traffic containing preset tag information sets in the dark web fingerprint library, traffic data potentially associated with suspicious user behavior can be accurately located, avoiding the tedious process of filtering large amounts of data required by traditional methods. Multidimensional association rules not only consider the number of occurrences of preset tags in dark web traffic, but also incorporate multiple dimensions such as the category of dark web sites, dark web access time, and historical behavioral deviations, providing a more comprehensive and accurate basis for user suspicion scoring. Existing classification methods may rely more on a single or limited feature dimension for classification. By classifying suspicious traffic with the same preset tag information, we can obtain a classification result for the user's suspicious access behavior, helping network security personnel gain a deeper understanding of the user's suspicious behavior patterns and locate the dark web sites currently visited by the user.
[0571] To apply this application example, please refer to Figure 5 , Figure 5 This is the overall framework diagram provided by the embodiment of this application, showing the process of processing anonymous traffic in this embodiment to trace the dark web; it mainly includes: 1. Compressed sensing of anonymous traffic in a large-scale complex network environment, 2. Multi-granular anonymous user traffic identification, 3. Fine-grained anonymous user behavior identification based on cross-level feature fusion, 4. Dark web tracing based on multi-scale traffic confirmation attack method. These steps are closely corresponding to the above solution, so the specific implementation details are not elaborated in detail. The general framework is summarized as follows:
[0572] "1. Anonymous traffic compression sensing in large-scale complex network environments": Traffic (compressed traffic) is collected from Internet traffic;
[0573] "2. Multi-granularity anonymous user traffic identification": process the traffic collected in step 1 to obtain onion service traffic. It should be noted that Figure 5 The corresponding method shown in is only one possible preset method for extracting onion service traffic from the collected traffic. In addition, various other methods can be used to achieve this goal based on specific application scenarios and requirements. These methods may include different traffic analysis techniques, data processing algorithms, or specific rule sets to more effectively identify and extract onion service traffic to meet the needs of different situations.
[0574] 3. Fine-grained anonymous user behavior recognition based on cross-layer feature fusion: Process the onion service traffic output from step 2 to obtain user traffic (first traffic).
[0575] "4. Dark web tracing based on multi-scale traffic confirmation attack method": Process the user traffic output in step 3 and finally obtain the correlation tracing result; and the image between "Load extraction of traffic watermark encoding" and "Cross-correlation of light and dark web based on flow watermark" in this section is Figure 3 .
[0576] The identification module 30 of this embodiment is based on the active and passive linkage dark web service association tracking method and stream watermark technology, which realizes the reconstruction of anonymous network links and the tracking and positioning of dark web sites, further improving the accuracy and efficiency of cross-domain anonymous user association.
[0577] It should be noted that in network communications, user network activity is typically conducted in the form of sessions. During a session, a user's network requests and responses form a collection of related traffic. Therefore, even though each request generates new traffic, this traffic still belongs to the same session. When a user begins a new session (for example, visiting a monitored clearnet website), the initial traffic of the session (i.e., the "first traffic") is captured and marked. This mark (e.g., the watermark embedded and time-correlatedly marked in the "first traffic" in this embodiment to obtain the traffic watermark) is unique and represents the specific attributes of the session. Importantly, this mark is applied to all subsequent traffic in the session until the session ends. On the backend, a traffic correlation algorithm tracks these marks. When a user continues to visit other websites (whether on the clearnet or darknet) within the same session, the traffic generated will carry the same mark. Therefore, even if these traffic flows are physically independent, the correlation algorithm in this embodiment can identify them as part of the same session through the mark, thereby achieving cross-domain anonymous user correlation.
[0578] Overall, this application has the following beneficial effects:
[0579] This application performs user behavior identification on the first network traffic on the Internet, and uses a classification prediction method based on a set confidence level to effectively distinguish the traffic of users visiting monitored websites. This method can improve the accuracy and robustness of behavior identification through strict confidence judgment. Embedding a watermark into the first traffic is equivalent to labeling the traffic with a unique label. This watermark technology is concealed and not easily tampered with or removed, providing a reliable basis for subsequent tracking and identification. Traffic is associated with a specific session through time-related tagging, which means that even if the user switches websites or uses different network paths in subsequent visits, as long as the session remains active, the traffic watermark can continue to work. Moreover, as a cross-domain identifier, watermarks can maintain consistency across different network domains and access paths, thereby supporting cross-domain access behavior association; time correlation tags further help determine the temporal sequence and logical relationship of these associated behaviors. This combined tagging method can enhance recognition accuracy and improve tracking continuity; therefore, once the watermark detector identifies that user traffic containing traffic watermarks has visited a dark web website, the information in the watermark can be used to effectively associate the dark web access behavior with previously monitored website access behavior. This method breaks down data silos and enables cross-domain tracking. In addition, the traffic watermark detector is deployed on a controlled guard node at the entrance and exit of the Tor network. The Tor network is one of the main channels for accessing the dark web, so this deployment location is of strategic significance. The watermark detector can cover a large amount of potential dark web access traffic, improving the efficiency of monitoring and identification.
[0580] In summary, this application proposes a traffic compression sensing algorithm for high-speed network environments, which efficiently and collaboratively processes traffic in multiple autonomous domain backbone networks, solves the scalability problem of flow analysis, and adapts to the challenge of large data volumes. For anonymous user traffic, fine-grained flow analysis technology is used to process encrypted and obfuscated traffic in real time and distinguish between dark web and open web access. By constructing multi-level and multi-dimensional traffic fingerprint features, the accuracy of the traffic correlation algorithm is improved, overcoming problems such as missing attributes and signal attenuation. In terms of cross-domain anonymous user correlation, an correlation technology that adapts to large-scale anonymous networks is provided, which realizes anonymous network link reconstruction and dark web site tracking, and enhances network security analysis capabilities.
[0581] Example 3:
[0582] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the multi-granularity method for identifying associations between user behaviors on the light and dark webs;
[0583] The multi-granularity method for identifying associations between user behavior on the light and dark webs, if implemented as a software functional unit and used as a standalone product, can be stored in a computer-readable storage medium. Based on this understanding, the present invention can also implement all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunications signal, and software distribution medium.
[0584] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A multi-granularity method for identifying associations between user behaviors on the bright and dark webs, characterized by: include: Obtaining network traffic from the Internet according to a preset period, classifying and predicting the network traffic according to a set confidence level, and obtaining a first traffic of users accessing the monitored website; Performing watermark embedding and time-correlation marking on the first traffic to obtain a traffic watermark includes: inserting Begin cells during the transmission of the first traffic according to a time interval sequence, and using the watermark sequence as a time fingerprint of the HS-RP circuit, and embedding the watermark sequence into all associated traffic by directionally adjusting the transmission delay of adjacent data packets; wherein the session corresponding to the marked first traffic is used for the user to perform all subsequent website access behaviors; Within a preset time period, if the traffic watermark detector identifies user traffic containing the traffic watermark, the dark web website accessed by the user traffic is defined as the dark web visited by the user, specifically: If the traffic watermark detector on the controlled guard node does not detect user traffic containing the traffic watermark within a preset time period, the user traffic of the unselected controlled nodes is intercepted, so that the user traffic passes through the controlled guard node; wherein the controlled guard node is deployed at the entrance and exit of the Tor network; If, within the preset time period, the traffic watermark detector on the controlled guard node detects user traffic containing the traffic watermark, a three-hop circuit is established based on the guard node, the intermediate node, and the exit node of the Tor network; wherein the three-hop circuit is used to encrypt and anonymize communications between the user and the hidden service; If the exit node through which the three-hop circuit passes is a controlled relay node, the dark web website and the IP address of the dark web website visited by the user traffic are determined based on the traffic watermark detector on the relay node, and the dark web website visited by the user traffic is defined as the dark web visited by the user; wherein the traffic watermark detector is deployed on a controlled guard node at the entrance and exit of the Tor network.
2. A multi-granularity method for identifying associations between user behaviors on the bright and dark webs as claimed in claim 1, characterized in that: The first traffic is watermarked and time-correlatedly marked to obtain a traffic watermark, specifically: The watermark is embedded into the first traffic according to a pseudo-random number generator and a shared key, and the first traffic is time-correlatedly marked according to a timestamp to obtain the traffic watermark.
3. The multi-granularity method for identifying associations between user behaviors on the bright and dark webs as described in claim 2 is characterized in that: The watermark is embedded into the first traffic according to the pseudo-random number generator and the shared key, specifically: Generate a binary sequence according to a pseudo-random number generator and a shared key, and convert the binary sequence into a time interval sequence of Begin cells; Begin cells are inserted into the transmission process of the first traffic according to the time interval sequence.
4. The multi-granularity method for identifying associations between user behaviors on the bright and dark webs as described in claim 2, characterized in that: The first traffic is marked with time correlation according to the timestamp, specifically: Obtaining a timestamp of a first data packet of the first traffic, and generating a watermark sequence according to the timestamp of the first data packet; The watermark sequence is used as the time fingerprint of the HS-RP circuit and is embedded into all associated traffic by directionally adjusting the transmission delay of adjacent data packets. The associated traffic refers to the traffic generated when users access different services through the same HS-RP circuit. The different services include monitored websites and subsequent dark web services.
5. The multi-granularity method for identifying associations between user behaviors on the bright and dark webs according to claim 1, characterized in that: The controlled guard nodes include a controlled entry node and a controlled exit node; The controlled ingress node is deployed with a traffic watermark generator, and the controlled egress node is deployed with a traffic watermark detector; The dark web websites visited by the same user are identified based on the traffic watermark generator and the traffic watermark detector.
6. The multi-granularity method for identifying associations between user behaviors on the bright and dark webs as claimed in claim 1, characterized in that: The network traffic is classified and predicted according to the set confidence level to obtain the first traffic of users visiting the monitored website, specifically: Performing sparse signal screening and compressed sensing processing on the network traffic to obtain compressed traffic; The compressed traffic is classified and predicted according to a set confidence level to obtain the first traffic of the user visiting the monitored website.
7. The multi-granularity method for identifying associations between user behaviors on the bright and dark webs according to claim 6, characterized in that: The network traffic is subjected to sparse signal screening and compressed sensing processing to obtain compressed traffic, specifically: According to a k-sparsity constraint, sparse feature vectors are screened from the network traffic to construct an anonymous traffic feature set; wherein the k-sparsity constraint is established by performing sparse distribution statistics on historical anonymous traffic data; Compressed sensing is performed on the anonymous traffic feature set according to a perception matrix to obtain the compressed traffic; wherein the perception matrix is obtained by mapping a generation matrix in a preset manner, and the generation matrix is established according to a selected error correction code.
8. The multi-granularity method for identifying associations between user behaviors on the bright and dark webs as claimed in claim 6, characterized in that: The compressed traffic is classified and predicted according to a set confidence level to obtain the first traffic of the user visiting the monitored website, specifically: Extracting onion service traffic from the compressed traffic in a predetermined manner; Perform feature screening on the onion service traffic to obtain a traffic representation vector; The traffic representation vector is classified and predicted according to a set confidence level to obtain the first traffic of the user visiting the monitored website.
9. The multi-granularity method for identifying associations between user behaviors on the bright and dark webs according to claim 8, characterized in that: Perform feature screening on the onion service traffic to obtain a traffic representation vector, specifically: Performing wavelet analysis on the onion service traffic to obtain multifractal features; Calculating information leakage amounts of several features in the multifractal features according to conditional entropy to obtain an information leakage amount set; taking features corresponding to the N largest information leakage amounts in the information leakage amount set from the multifractal features to form a first feature set; In the current dimensional space and the preset low-dimensional space, calculating the conditional probability between each pair of data points in the first feature set, to obtain a first conditional probability set and a second conditional probability set respectively; With the goal of minimizing the KL divergence loss between the first conditional probability set and the second conditional probability set, mapping the first feature set from the current dimensional space to the low-dimensional space to obtain a second feature set; The second feature set is subjected to a nonlinear transformation according to an encoder to obtain the flow representation vector; wherein the encoder is established based on an attention mechanism and a multi-layer perceptron.
10. The multi-granularity method for identifying associations between user behaviors on the bright and dark webs according to claim 8, characterized in that: The traffic representation vector is classified and predicted according to a set confidence level to obtain the first traffic of the user visiting the monitored website, specifically: In the traffic representation vector, the distance between each traffic embedding and the centroid of the monitored website set is calculated to obtain a number of distance sets; For a distance set corresponding to a first flow embedded in the plurality of distance sets, if the nearest centroid distance is less than the radius of the website corresponding to the nearest centroid, and the difference between the second nearest centroid distance and the nearest centroid distance is greater than a preset threshold, then it is determined that the website corresponding to the nearest centroid is the first monitored website visited by the user, and the flow of the user visiting the first monitored website is defined as the first flow of the user visiting the monitored website; The closest centroid distance is the distance between the first flow embedding and the closest centroid, and the second closest centroid distance is the distance between the first flow embedding and the second closest centroid.
11. A multi-granularity method for identifying associations between user behaviors on the bright and dark webs according to any one of claims 1 to 10, characterized in that: After obtaining the first flow of users visiting the monitored website, the method further includes: By matching the first traffic with known dark web service features in a dark web fingerprint library, the hidden dark web accessed by the user is identified; wherein the dark web fingerprint library is updated in real time based on the dark web and its real-time access information.
12. A multi-granularity method for identifying associations between user behaviors on the bright and dark webs according to claim 11, characterized in that: The dark web fingerprint database is updated in real time based on the dark web and its real-time access information, including: Extract the traffic data of users accessing the dark web according to the high-frequency time period to obtain comprehensive traffic data; wherein the high-frequency time period is the time window in which the number of users accessing the dark web is higher than the preset number of dark web accesses; Extracting target features of users accessing the dark web from the comprehensive traffic data to obtain a target feature set; performing user access behavior pattern recognition on the comprehensive traffic data according to a clustering algorithm to obtain a behavior pattern recognition result; The target feature set and the behavior pattern recognition result form a fingerprint feature set; According to the preset update mechanism, the fingerprint feature set is stored in the dark web fingerprint library to obtain the updated dark web fingerprint library.
13. The multi-granularity method for identifying associations between user behaviors on the bright and dark webs according to claim 11, characterized in that: After defining the dark web website accessed by the user traffic as the dark web accessed by the user, the method further includes: After a preset time interval, obtain the current dark web fingerprint database; Detecting traffic containing a preset tag information set in the current dark web fingerprint library, and defining traffic that successfully matches as traffic to be tested; Scoring the traffic to be tested according to multi-dimensional association rules to obtain a user suspicion score; wherein the multi-dimensional association rules are established based on the number of occurrences of preset tags in dark web traffic, the category of dark web websites, dark web access time, and historical behavior deviation thresholds; The traffic to be tested whose user suspicion score is higher than a preset threshold is defined as suspicious traffic, and the traffic with the same preset tag information in the suspicious traffic is classified to obtain the dark web traffic classification result of the user's suspicious access behavior.
14. The multi-granularity method for identifying associations between user behaviors on the bright and dark webs according to claim 11, characterized in that: By matching the first traffic with known dark web service features in a dark web fingerprint library, the hidden dark web accessed by the user is identified, specifically: Performing enhanced flow marking on the first flow to obtain enhanced flow marking information; wherein the enhanced flow marking includes a dynamic watermark mark composed of a composite watermark sequence and a multi-protocol mark composed of an encrypted watermark segment, the composite watermark sequence is obtained by performing a hash operation on a timestamp and a data packet size, and the encrypted watermark segment is obtained by inserting a custom data field during a handshake process of a transport layer security protocol; After a preset time interval, the current dark web fingerprint library is obtained, and the first dark web traffic having the same enhanced traffic marking information as the first traffic is obtained from the dark web fingerprint library according to the enhanced index, and the website corresponding to the first dark web traffic is defined as the hidden dark web visited by the user; wherein the enhanced index is obtained by establishing an index containing the enhanced traffic marking information for each dark web traffic record in the dark web fingerprint library.
15. A multi-granularity device for identifying the association between user behavior on the bright and dark webs, characterized in that: Including monitoring module, marking module and identification module; The monitoring module is configured to obtain network traffic from the Internet according to a preset period, classify and predict the network traffic according to a set confidence level, and obtain a first traffic volume of users accessing the monitored website; The marking module is configured to perform watermark embedding and time-correlation marking on the first traffic to obtain a traffic watermark, including: inserting Begin cells during the transmission of the first traffic according to a time interval sequence, and using the watermark sequence as a time fingerprint of the HS-RP circuit, and embedding the watermark sequence into all associated traffic by directionally adjusting the transmission delay of adjacent data packets; wherein the session corresponding to the marked first traffic is used by the user for all subsequent website access behaviors; The identification module is configured to define the dark web website accessed by the user traffic as the dark web accessed by the user if the traffic watermark detector identifies the user traffic containing the traffic watermark within a preset time period, specifically: If the traffic watermark detector on the controlled guard node does not detect user traffic containing the traffic watermark within a preset time period, the user traffic of the unselected controlled nodes is intercepted, so that the user traffic passes through the controlled guard node; wherein the controlled guard node is deployed at the entrance and exit of the Tor network; If, within the preset time period, the traffic watermark detector on the controlled guard node detects user traffic containing the traffic watermark, a three-hop circuit is established based on the guard node, the intermediate node, and the exit node of the Tor network; wherein the three-hop circuit is used to encrypt and anonymize communications between the user and the hidden service; If the exit node through which the three-hop circuit passes is a controlled relay node, the dark web website and the IP address of the dark web website visited by the user traffic are determined based on the traffic watermark detector on the relay node, and the dark web website visited by the user traffic is defined as the dark web visited by the user; wherein the traffic watermark detector is deployed on a controlled guard node at the entrance and exit of the Tor network.
16. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 15, characterized in that: The marking module is specifically: The watermark is embedded into the first traffic according to a pseudo-random number generator and a shared key, and the first traffic is time-correlatedly marked according to a timestamp to obtain the traffic watermark.
17. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 16, characterized in that: The watermark is embedded into the first traffic according to the pseudo-random number generator and the shared key, specifically: Generate a binary sequence according to a pseudo-random number generator and a shared key, and convert the binary sequence into a time interval sequence of Begin cells; Begin cells are inserted into the transmission process of the first traffic according to the time interval sequence.
18. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 16, characterized in that: The first traffic is marked with time correlation according to the timestamp, specifically: Obtaining a timestamp of a first data packet of the first traffic, and generating a watermark sequence according to the timestamp of the first data packet; The watermark sequence is used as the time fingerprint of the HS-RP circuit and is embedded into all associated traffic by directionally adjusting the transmission delay of adjacent data packets. The associated traffic refers to the traffic generated when users access different services through the same HS-RP circuit. The different services include monitored websites and subsequent dark web services.
19. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 15, characterized in that: The controlled guard nodes include a controlled entry node and a controlled exit node; The controlled ingress node is deployed with a traffic watermark generator, and the controlled egress node is deployed with a traffic watermark detector; The dark web websites visited by the same user are identified based on the traffic watermark generator and the traffic watermark detector.
20. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 15, characterized in that: The monitoring module includes a compression unit and an identification unit; The compression unit is configured to perform sparse signal screening and compressed sensing processing on the network traffic to obtain compressed traffic; The identification unit is used to classify and predict the compressed traffic according to a set confidence level to obtain the first traffic of the user visiting the monitored website.
21. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 20, characterized in that: The compression unit includes a screening subunit and a compression subunit; The screening subunit is configured to screen out sparse feature vectors from the network traffic according to a k-sparsity constraint condition to construct an anonymous traffic feature set; wherein the k-sparsity constraint condition is established by performing sparse distribution statistics on historical anonymous traffic data; The compression subunit is used to perform compressed sensing on the anonymous traffic feature set according to a sensing matrix to obtain the compressed traffic; wherein the sensing matrix is obtained by mapping the generation matrix in a preset manner, and the generation matrix is established according to a selected error correction code.
22. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 20, characterized in that: The recognition unit includes an onion subunit, a feature subunit and a prediction subunit; The onion sub-unit is used to extract onion service traffic from the compressed traffic in a preset manner; The feature subunit is used to perform feature screening on the onion service traffic to obtain a traffic representation vector; The prediction subunit is used to perform classification prediction on the traffic representation vector according to a set confidence level to obtain the first traffic of the user visiting the monitored website.
23. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 22, characterized in that: The characteristic subunit is specifically: Performing wavelet analysis on the onion service traffic to obtain multifractal features; Calculating information leakage amounts of several features in the multifractal features according to conditional entropy to obtain an information leakage amount set; taking features corresponding to the N largest information leakage amounts in the information leakage amount set from the multifractal features to form a first feature set; In the current dimensional space and the preset low-dimensional space, calculating the conditional probability between each pair of data points in the first feature set, to obtain a first conditional probability set and a second conditional probability set respectively; With the goal of minimizing the KL divergence loss between the first conditional probability set and the second conditional probability set, mapping the first feature set from the current dimensional space to the low-dimensional space to obtain a second feature set; The second feature set is subjected to a nonlinear transformation according to an encoder to obtain the flow representation vector; wherein the encoder is established based on an attention mechanism and a multi-layer perceptron.
24. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 22, characterized in that: The prediction subunit is specifically: In the traffic representation vector, the distance between each traffic embedding and the centroid of the monitored website set is calculated to obtain a number of distance sets; For a distance set corresponding to a first flow embedded in the plurality of distance sets, if the nearest centroid distance is less than the radius of the website corresponding to the nearest centroid, and the difference between the second nearest centroid distance and the nearest centroid distance is greater than a preset threshold, then it is determined that the website corresponding to the nearest centroid is the first monitored website visited by the user, and the flow of the user visiting the first monitored website is defined as the first flow of the user visiting the monitored website; The closest centroid distance is the distance between the first flow embedding and the closest centroid, and the second closest centroid distance is the distance between the first flow embedding and the second closest centroid.
25. A multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to any one of claims 15 to 24, characterized in that: After obtaining the first flow of users visiting the monitored website, the method further includes: By matching the first traffic with known dark web service features in a dark web fingerprint library, the hidden dark web accessed by the user is identified; wherein the dark web fingerprint library is updated in real time based on the dark web and its real-time access information.
26. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 25, characterized in that: The dark web fingerprint database is updated in real time based on the dark web and its real-time access information, including: Extract the traffic data of users accessing the dark web according to the high-frequency time period to obtain comprehensive traffic data; wherein the high-frequency time period is the time window in which the number of users accessing the dark web is higher than the preset number of dark web accesses; Extracting target features of users accessing the dark web from the comprehensive traffic data to obtain a target feature set; performing user access behavior pattern recognition on the comprehensive traffic data according to a clustering algorithm to obtain a behavior pattern recognition result; The target feature set and the behavior pattern recognition result form a fingerprint feature set; According to the preset update mechanism, the fingerprint feature set is stored in the dark web fingerprint library to obtain the updated dark web fingerprint library.
27. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 25, characterized in that: After defining the dark web website accessed by the user traffic as the dark web accessed by the user, the method further includes: After a preset time interval, obtain the current dark web fingerprint database; Detecting traffic containing a preset tag information set in the current dark web fingerprint library, and defining traffic that successfully matches as traffic to be tested; Scoring the traffic to be tested according to multi-dimensional association rules to obtain a user suspicion score; wherein the multi-dimensional association rules are established based on the number of occurrences of preset tags in dark web traffic, the category of dark web websites, dark web access time, and historical behavior deviation thresholds; The traffic to be tested whose user suspicion score is higher than a preset threshold is defined as suspicious traffic, and the traffic with the same preset tag information in the suspicious traffic is classified to obtain the dark web traffic classification result of the user's suspicious access behavior.
28. The multi-granularity device for identifying associations between user behaviors on the bright and dark webs according to claim 25, characterized in that: By matching the first traffic with known dark web service features in a dark web fingerprint library, the hidden dark web accessed by the user is identified, specifically: Performing enhanced flow marking on the first flow to obtain enhanced flow marking information; wherein the enhanced flow marking includes a dynamic watermark mark composed of a composite watermark sequence and a multi-protocol mark composed of an encrypted watermark segment, the composite watermark sequence is obtained by performing a hash operation on a timestamp and a data packet size, and the encrypted watermark segment is obtained by inserting a custom data field during a handshake process of a transport layer security protocol; After a preset time interval, the current dark web fingerprint library is obtained, and the first dark web traffic having the same enhanced traffic marking information as the first traffic is obtained from the dark web fingerprint library according to the enhanced index, and the website corresponding to the first dark web traffic is defined as the hidden dark web visited by the user; wherein the enhanced index is obtained by establishing an index containing the enhanced traffic marking information for each dark web traffic record in the dark web fingerprint library.
29. A storage medium, characterized in that The storage medium stores a computer program, which is called and executed by a computer to implement a multi-granularity method for identifying associations between light and dark web user behaviors as described in any one of claims 1 to 14 above.