Real-time flow regulation and control method for network fingerprint defense

By building anonymous sets and dynamic regulation matrix, the problems of insufficient anonymity and high bandwidth delay in the existing technology are solved, and efficient traffic regulation of anonymous communication systems such as Tor is realized, reducing the success rate of website fingerprint attacks.

CN120342718APending Publication Date: 2025-07-18HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510543352.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When resisting website fingerprint attacks, the existing technology has insufficient anonymity protection, high bandwidth and delay overhead, and poor adaptability to dynamic network traffic.

Method used

By building anonymous sets, using historical data to generate traffic matrix and regulation matrix, dynamic traffic regulation is carried out, and combining congestion mitigation and virtual packet filling, the success rate of website fingerprint attack recognition is reduced.

Benefits of technology

It effectively reduces the success rate of website fingerprint attack recognition, takes into account bandwidth and delay overhead, has good adaptability and scalability, and is suitable for anonymous communication systems such as Tor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342718A_ABST
    Figure CN120342718A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time flow regulation and control system and method for network fingerprint defense, a program, equipment and a storage medium, and mainly relates to the field of network security and user privacy protection. The method comprises the following steps: collecting a traffic sequence generated when a user accesses each site in real time, and constructing a historical traffic database; generating unified regulation and control information by using a traffic matrix based on a time slot, inter-site anonymous set clustering and anonymous in-set mode aggregation technology; and finally, according to the regulation matrix, performing dynamic intervention on the traffic through a real-time regulation module, and by adopting the means of congestion relief, virtual data packet filling and random noise introduction, unifying the traffic characteristics of each station, reducing the bandwidth and time delay overhead caused by virtual filling, and keeping the overall communication efficiency at the same time. According to the scheme of the invention, the method achieves the self-adaptive regulation and control of the dynamic network flow while improving the anonymity and safety of a user, has good system performance and deployability, and is suitable for Tor and other anonymous communication scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the field of network security and user privacy protection, and particularly relates to a real-time traffic regulation system, method, program, device and storage medium for network fingerprint defense. Background Art

[0002] With the continuous popularization of Internet technology, more and more people work and entertain through the network. It is estimated that the amount of data generated globally every day is about 491EB currently, which involves a large amount of personal privacy information and activity behaviors of users, containing inestimable value. Recently, the relatively frequent user privacy leakage incidents have promoted the social research and exploration on user network privacy protection, and a series of technical applications have been proposed. Among them, the onion routing Tor based on the MIX network stands out among a group of anonymous network communication systems with its powerful performance and anonymous security. However, the privacy activity information of users in using Tor may still be affected by website fingerprint attacks, thereby leaking the behavior information of users.

[0003] In website fingerprint attacks, local eavesdroppers passively monitor the data traffic between the target client and the first-hop node (usually a bridge) of the Tor network. Even if the data is encrypted, the attacker can still infer the website or service accessed by the target by analyzing and using the inherent characteristic pattern information such as timing, packet length, and traffic direction in network transmission, because these physical characteristics can expose the user's access behavior to a certain extent.

[0004] To resist website fingerprint attacks, researchers have extended two types of defense schemes based on traffic shaping and traffic obfuscation from three simple operations of padding, blocking, and sending data packets. However, there are many problems in the existing schemes. For example, although the overhead impact of traffic obfuscation is small, it is difficult to resist attacks based on deep learning in practice; while traffic shaping makes it impossible for attackers to distinguish the traffic information of the accessed site through normalization, but sacrifices a lot of bandwidth and delay, seriously affecting user experience. Therefore, most of the existing schemes are difficult to be deployed on a large scale.

[0005] In view of the problems existing in the existing schemes, the present invention provides a real-time traffic regulation system and method for network fingerprint defense. By combining technical means of traffic clustering anonymization and dynamic traffic regulation, and using historical data to assist in optimizing the construction of aggregation patterns, it can not only effectively reduce the success rate of website fingerprint attacks, but also control the bandwidth and delay overhead in real-time applications, thereby taking into account system performance and user experience while improving security, and having broad practical deployment value. Summary of the Invention

[0006] The object of the present invention is to propose a real-time traffic regulation method for network fingerprint defense to overcome several deficiencies existing in existing website fingerprint protection solutions, including but not limited to insufficient anonymity guarantee, high bandwidth and latency overheads, and poor adaptability to dynamic network traffic.

[0007] This method takes the concept of an anonymous set as the core. By classifying and aggregating the characteristics of user access behaviors, the behavior characteristics of individual users within the anonymous set to which they belong are made as consistent as possible with those of other members, thereby weakening the identifiability of traffic fingerprints. During the actual traffic transmission process, the network transmission mode within each time slot is dynamically confused to balance communication efficiency and system resource consumption, and externally it appears similar to that without defense processing. The finally realized integrated protection solution not only strengthens the privacy protection of user access behaviors, but also has good deployability, scalability, and adaptability to complex network environments, and is suitable for actual deployment in anonymous communication systems such as Tor.

[0008] The present invention provides the following technical solution to achieve the above object: This real-time traffic regulation method for network fingerprint defense mainly includes three steps: traffic collection, regulation information generation, and traffic regulation. Technical solution

[0009] 1. Traffic collection: Within the preset period T for generating regulation information, the data streams of the user's access behaviors to each site are collected as reference historical data to prepare for the generation of regulation information. This method uses a network packet capture tool to collect the data streams generated when the user accesses each site in real time. During the collection process, the system only extracts the time interval between the data packet and the first data packet of the access sequence, as well as the direction of the data packet flowing in and out of the current device, and then processes each data packet into the format of <time, transmission direction> to form a simplified traffic sequence, reducing the storage burden. For each site, the system collects the traffic sequences generated by multiple access instances as follows: where t i represents the time interval of the i-th data packet relative to the first data packet of this access, d i ∈{+1, -1} represents the transmission direction of the i-th data packet, +1 indicates that the data packet is flowing out (i.e., sent from the user side to the network), and -1 indicates that the data packet is flowing in (i.e., transmitted from the network to the user side), Z represents the total number of access instances collected for this site, and n k is the number of data packets in the k-th access. After the traffic sequences of each site are collected and processed, they will be updated or added to the traffic database to form a dynamically updated historical data set to provide reference for the subsequent generation of regulation information.

[0010] 2. Regulation Information Generation: Before reaching the preset period T, the traffic collection stores the historical data corresponding to each site in the traffic database. After reaching the period T, the regulation information generation module is triggered to comprehensively analyze the historical traffic characteristics of each site and extract the key information available for the real-time regulation module. This step specifically includes the following three sub-steps: 1) Intra-site Feature Integration: For each site, the system first retrieves the <time, transmission direction> traffic sequence collected during the period T from the traffic database. To capture the feature representation of its traffic transmission in time and space, this method uses a time-slot-based traffic matrix, that is, the entire traffic sequence of a single access is divided into a time-slot-based traffic matrix according to N fixed time slots, each time slot having a duration of Δt, as shown below: where a 1,i represents the number of outbound data packets in the i-th time slot, and a 2,i represents the number of inbound data packets in the i-th time slot. For multiple traffic sequences of the same site they are all converted into traffic matrices To cover all sequences within the site to show the same features externally, a representative matrix is constructed here. Each element is the maximum value of the corresponding elements of all traffic matrices of this site 2) Inter-site Anonymous Agglomerative Clustering: To achieve unified regulation of cross-site traffic characteristics, it is necessary to compare the similarity of traffic characteristics of different sites and cluster sites with similar characteristics into the same anonymous set, thereby obtaining the mapping between sites and anonymous sets. Let the representative feature matrices of site p and site q be and The Euclidean distance d is used as the similarity metric Based on this distance, sites with closer feature distances are grouped into the same anonymous set S through a clustering algorithm, and it is ensured that each anonymous set contains at least K sites, so that the set has enough elements to ensure anonymity. 3) Pattern Aggregation within Anonymous Set: To construct an aggregation matrix covering the traffic peaks of all sites within the anonymous set, for each anonymous set S, its included site set is {s1, s2, …, s k}, the maximum value of the number of packets in each time slot of all sites within the anonymous set is taken, that is, the aggregation matrix is defined as The aggregation matrix corresponding to each anonymized set ensures that regardless of which site within the anonymized set exhibits unique characteristics during a certain time slot, this unified pattern can cover such situations. Thus, during real-time regulation, by supplementing virtual data packets, the traffic can reach the unified pattern, eliminating the differences among sites. To further optimize the regulation strategy and reduce the additional overhead caused by virtual data packet filling, the present invention introduces the calculation of the probability mass function. Specifically for each anonymized set, a probability matrix P with the same shape as the aggregation matrix is constructed S , where the elements use the historical traffic matrix corresponding to the anonymized set to calculate the probability of data transmission and reception existing in each time slot, that is, each matrix element is calculated as follows where I represents whether there is data packet circulation within this time slot. If there is, it is recorded as 1; otherwise, it is 0. 4) Simulation regulation optimization: After aggregating and analyzing historical traffic data, information that can be used for regulation has been obtained. However, due to the diverse traffic sequences of different websites and their sub-pages, directly using the aggregation matrix M will cause huge bandwidth overhead. Therefore, simulation regulation will be performed on each anonymized set to train and optimize the aggregation matrix M into a regulation matrix Specifically, it is for to perform matrix optimization, where w and b are learnable parameters for training, with the same shape as the aggregation matrix M, and τ is a preset adjustment amplitude. This step aims to consider the overheads of both bandwidth and delay simultaneously. The regulation matrix is optimized and trained using simulated actual regulation. Matrix optimization uses the historical data set as the training set during training, and the aggregation matrix M is calculated with the learning parameters w and b to obtain the regulation matrix Subsequently, optimization is carried out with the assistance of the loss function and the optimizer. The design of the loss function includes two parts considering the delay and bandwidth overheads. To simulate actual regulation, first use the historical data set as the training set. For each anonymized set S corresponding to the traffic matrix A, perform the initialization regulation in the traffic regulation module, and then obtain the actual regulation matrix of this traffic sequence and repeat O times to balance the error caused by sampling. The first part considers the delay overhead and is used to measure the difference between the regulation matrix and the actual traffic matrix A during each time slot, and is calculated in the way of mean square error The second part considers the bandwidth overhead and calculates the difference between and A in the total bandwidth. The way to calculate the total bandwidth is as follows The loss function design of the bandwidth overhead is The final loss function is Among them, λ ∈ (0, 1) is used to regulate the optimization focus between delay and bandwidth overhead. The regulation matrix obtained through this training While minimizing the virtual packet filling, it maintains a certain consistency with the bandwidth characteristics of the real traffic, coordinates security and communication efficiency, and will be used to actually regulate traffic reception and transmission in the traffic regulation module.

[0011] 3. Traffic regulation: The traffic regulation of the regulation information is still from a macroscopic perspective. Considering that the content changes of the site, the link status of the sub-pages of the site and user communication may change, more refined real-time regulation is required. After receiving and updating the regulation information at the user end and the server end (such as the first-hop node of the Tor network), the traffic regulation module will allocate a state machine to intervene and adjust the real-time traffic generated by the user accessing the site. This step specifically includes the following sub-steps: 1) Start executing regulation: When it is detected that the user is about to or has already started accessing a certain site, the system will allocate a real-time regulation state machine for this access behavior, prepare to regulate the real-time traffic of the site and record the data therein. 2) Initialize regulation: According to the target site accessed by the user, its affiliated anonymity set S and the regulation matrix are matched through the mapping between the site and the anonymity set If the site is a new site that has not been included or has few accesses, an anonymity set can be randomly or default selected to ensure the alignment of traffic characteristics as much as possible. And the actual regulation matrix Then uses the scrambled probability matrix P S And the preset threshold a to determine the mask matrix mask of the regulation matrix through the sampling method of cumulative values: The mask matrix mask and the regulation matrix Jointly confirm the actual regulation matrix In addition, by randomly selecting (1, E) as e to obtain the congestion avoidance judgment for real-time regulation. 3) Real-time regulation: During the actual network transmission process of the regulation matrix, the real sequence data of <time, transmission direction> collected is continuously written into the buffer, and enters the real-time regulation module for processing after each time slot cycle. The specific real-time regulation processing flow is as follows: ① When entering a certain time slot, the state machine reads the number of packets received within the current time slot from the buffer; ② Compare this quantity with the total number of packets in the time slots from the current time slot to the next e time slots in the regulation matrix to determine whether there is a congestion risk; ③ If the actual number of packets exceeds the preset value, the state machine starts the congestion relief mechanism, immediately sends some data packets (until the preset congestion relief limit is reached or the buffer is cleared), adds 1 to the congestion relief counter, and then jumps to the next time slot; ④ If there is no congestion risk, the state machine further determines whether the current time slot allows data transmission; ⑤ For time slots that do not allow data to be sent, the system will temporarily block the data packets to be sent in the buffer and wait for the next time slot to process; ⑥ When the time slot allows transmission, the data packet is sent according to the preset value of the control matrix; for the insufficient part, the virtual data packet is automatically supplemented; when the number of virtual data packets exceeds the preset threshold, the system adds 1 to the overfill counter; ⑦ After each time slot is processed, the state machine automatically enters the control cycle of the next time slot until the site visit ends or exceeds the preset time limit. The state machine introduces noise ∈ on the basis of each execution of sending data packets, which is designed to obey the artificially set truncated normal distribution. This ensures that the noise is not too loud and affects the user experience, and at the same time achieves the purpose of obfuscation, so that the overall traffic characteristics appear to be unprotected, thereby reducing the attacker's attention to the access behavior. When the site visit is completed, the state machine will check the page address and the ratio of the "congestion relief counter" and the "overfill counter" to determine whether the current page needs to switch to an anonymous set control matrix with more or less traffic. To update the original regulation information to cope with the situation that the page content changes significantly during the cycle or the user communication link conditions change, so as to continuously ensure the coordinated optimization of anonymity and communication efficiency. Beneficial Effects

[0012] The present invention proposes a real-time traffic control system and method for network fingerprint defense. By introducing anonymous set construction and dynamic control mechanism, it can effectively unify the traffic characteristics between sites, significantly reduce the recognition success rate of website fingerprint attacks, and improve user anonymity. Compared with the existing technology, this method takes into account bandwidth and delay overhead while ensuring privacy. Through historical data-driven control matrix optimization and probability-guided filling strategy, efficient use of communication resources is achieved. The system has good adaptability and scalability, and has low update overhead. It can dynamically respond to changes in site content and fluctuations in the network environment. It has good deployability and scalability and is suitable for actual deployment in anonymous communication scenarios including Tor. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is a basic block diagram of the real-time traffic control method for network fingerprint defense of the present invention.

[0014] Figure 2 It is a processing flow chart for real-time regulation of the state machine. Specific implementation manner

[0015] The present invention will be further described below in conjunction with the accompanying drawings of the specification.

[0016] As Figure 1 shown, the real-time traffic regulation method for network fingerprint defense mainly includes three steps: traffic collection, regulation information generation, and traffic regulation.

[0017] In specific implementation, first, the traffic collection module uses a network packet capture tool to capture in real time the data stream generated when the user accesses each site, extracts the time interval between each data packet and the first packet of this access and the transmission direction of the data packet, so as to form a simplified traffic sequence, and updates the data of multiple accesses to the traffic database; secondly, the regulation information generation module extracts historical traffic data from the database, divides each access traffic sequence into a traffic matrix according to a fixed time slot, constructs a representative matrix for each site, then clusters similar sites into an anonymous set according to the Euclidean distance, and takes the maximum value of the data of each time slot in the anonymous set to generate a regulation matrix, and calculates the probability mass function at the same time to guide the optimization of the regulation matrix; finally, after receiving the regulation information at the user end or the server end, the traffic regulation module processes the real-time collected traffic data by allocating a real-time regulation state machine, including reading data from the buffer, comparing with the preset value of the regulation matrix, taking congestion mitigation or filling and sending (and introducing random noise) measures, ensuring that the traffic characteristics after regulation tend to be unified, and updating the regulation strategy according to the feedback information.

[0018] Embodiments of the present invention are as follows:

[0019] Step 1 Traffic collection:

[0020] Use a packet capture tool to continuously collect the traffic sequences of each site within a preset period of 3 days. For the collected data packet sequence, <time, transmission direction> can be extracted from each data packet, and it is converted into <interval time from the first data packet of this access, transmission direction>. Perform the same operation on the data packet sequence of a single access to the same site to obtain the traffic sequence f. After collecting sufficient data, then update the traffic sequence and its corresponding site information to the traffic database.

[0021] Step 2 Regulation information generation:

[0022] Step 2.1 Feature integration within the site; Using the historical traffic data of each site stored in the traffic database, the system triggers the regulation information generation module at the end of the preset period. This module first divides the traffic sequence of each site into a traffic matrix according to the setting of N fixed time slots, each time slot having a duration of Δt. For a single access instance, it is expressed as where a 1,i represents the number of outbound packets in the i-th time slot, and a 2,j represents the number of inbound packets in the j-th time slot. For multiple access instances of the same site, the system extracts the element-wise maximum values of the corresponding matrices to form the representative matrix of the site

[0023] Step 2.2 Anonymity-based clustering between sites; For the representative matrix of each site The system clusters sites with similar representative matrices into an anonymous set S based on the Euclidean distance, so that sites with similar traffic patterns are clustered, and multiple anonymous sets are formed. At the same time, the system will ensure that each anonymous set contains at least K sites to ensure that the elements of the anonymous set are sufficient to improve the ability to resist fingerprint attacks.

[0024] Step 2.3 Pattern aggregation within the anonymous set; For each anonymous set S, its set of sites is {s1, s2, …, s k}, for the representative matrices of all its sites Take the maximum value in each time slot to generate the aggregation matrix M S . At the same time, the probability mass function is used to count whether there are packet transmissions in each time slot in the historical traffic data, and a probability matrix P S with the same shape as the aggregation matrix is constructed to reflect the occurrence probability of data flow in each time slot, so as to provide data basis for the subsequent optimization of the regulation matrix.

[0025] Step 2.4 Simulation-based regulation optimization. Directly using the aggregation matrix M for real-time regulation will result in too much virtual packet filling, thus generating a large bandwidth overhead. Therefore, the actual traffic regulation process will be simulated to optimize the aggregation matrix M and train the optimized regulation matrix First, introduce learnable parameters w and b (with the same shape as the aggregation matrix M) and a preset adjustment amplitude τ, and use a shrinkage formula similar to linear regression for matrix optimization. During actual training, the probability matrix P S and the preset threshold a will be used to obtain the actual regulation matrix through cumulative random sampling, and O samplings will be performed to reduce the error caused by random sampling. The loss function consists of two parts considering delay overhead and considering bandwidth overhead and jointly. In addition, λ ∈ (0, 1) is introduced to regulate the optimization focus of delay and bandwidth overhead. The final loss function is Through training, the regulation matrices corresponding to each anonymous set are finally optimized

[0026] Step 3 Flow Regulation:

[0027] Step 3.1 Start Regulation Execution: When it is detected that the user is about to or has already started accessing a certain target site, the system immediately activates the real-time regulation state machine, identifies and records this access behavior, and establishes an initial context for subsequent regulation.

[0028] Step 3.2 Initialize Regulation: The system retrieves and matches the corresponding regulation matrix according to the mapping relationship between the target site and the anonymization set. If the target site is a new site with insufficient data, a default or randomly selected anonymization set can be used. In addition, the system uses a preset probability matrix and a sampling threshold a to determine the mask matrix mask through cumulative sampling, and preprocesses the regulation matrix to obtain the real-time regulation matrix.

[0029] Step 3.3 Real-Time Regulation: As Figure 2 shown, the traffic sequence data generated by the system in real time is continuously written into the buffer. The state machine triggers the regulation process within each fixed time slot period, making the traffic characteristics tend to be unified while ensuring the user experience, and improving the ability to resist website fingerprint attacks. In addition, random noise ∈ that follows a truncated normal distribution is introduced during each packet sending process to achieve the purpose of confusion at the same time, making the overall traffic characteristics externally appear approximately in an unprotected state, thereby reducing the attacker's attention to the access behavior. The specific process is as follows: ① After entering a certain time slot, the state machine reads the actual number of received packets within this time slot from the buffer. ② The system compares the read number of packets with the preset number of packets in the corresponding time slot in the actual regulation matrix to determine whether there is congestion or insufficient data. ③ When it is detected that the actual number of packets exceeds the total number of packets from the current time slot to the next e time slots, the state machine will start congestion mitigation measures, send enough packets according to the preset upper limit to reduce the buffer load, increment the congestion mitigation counter by 1, and skip the remaining data in the current time slot. ④ When it is detected that the number of packets is insufficient, the state machine supplements the necessary virtual packets according to the regulation matrix value to make the actual sent traffic close to the preset target. At the same time, if the filling exceeds the preset threshold, the overfilling counter is incremented by 1. ⑤ If it is determined that data sending is not allowed within the current time slot, the data in the buffer is temporarily blocked and processed in the next time slot. ⑥After each time slot is processed, the state machine automatically waits to move on to the next time slot for processing until the current site access ends or exceeds the preset control time limit.

[0030] Step 3.4 Access End: When the current site access ends, the state machine summarizes the access process, evaluates the current control performance based on the page address and the ratio of the congestion mitigation counter to the overfill counter, and feeds back to the control information generation module. Based on this, the system determines whether to switch to an anonymous set control matrix with more or less abundant traffic, so as to update the control strategy in the face of page content changes or user communication link condition changes, ensuring the continuous coordinated optimization of overall anonymity and communication efficiency.

Claims

1. A real-time traffic regulation method for network fingerprint defense, characterized in that It includes the following steps: Step 1, traffic collection: Use a network packet capture tool to capture the data stream generated when users access each site in real time, extract the time interval between the data packet and the first access packet, and the transmission direction of the data packet, generate a traffic sequence of <time, transmission direction>, and update and store the sequence in the traffic database; Step 2, generation of regulation information: Extract the historical traffic data of each site from the historical traffic database. By dividing the single access traffic sequence into several fixed time slots to form a time slot-based traffic matrix, construct a representative matrix for each site, and perform anonymous clustering on each site according to the Euclidean distance. Then, take the maximum value of the number of packets of all sites in each time slot within each anonymous set to construct a regulation matrix with a unified pattern. At the same time, combine the calculation of the probability mass function to obtain the corresponding probability matrix, and then use the two types of matrices to simulate traffic regulation to guide the optimization of the subsequent regulation matrix; Step 3, traffic regulation: After the updated regulation information is received at the user side and / or the server side, allocate a real-time regulation state machine for the user access behavior, and process the real-time collected <time, transmission direction> traffic data according to the regulation matrix, including reading the number of data packets in the buffer, comparing with the preset number of packets, and adopting congestion mitigation, filling virtual data packets, blocking data packets, and random noise injection measures to dynamically regulate the traffic according to the comparison results, so that the regulated traffic as a whole tends to be unified and externally appears approximately in an unprotected state. At the same time, record the congestion mitigation and overfilling situations, and update the regulation information according to the record after the site access ends.

2. The real-time traffic regulation method for network fingerprint defense according to claim 1, characterized in that In the step of generating the regulation information, it includes: 1) Intra-site feature integration: For the traffic sequence f of each site, which is divided into a traffic matrix A based on N fixed time slots, each with a duration of Δt, where each time slot includes the number of outbound and inbound data packets, construct a representative matrix composed of the maximum values of the corresponding elements of all traffic matrices of this site 2) Anonymous clustering between sites: Based on the Euclidean distance d between the representative matrices of the sites, cluster the sites with similar traffic characteristics to form an anonymous set S, that is: where p and q represent two different stations represents the element in the station representative matrix To ensure anonymity, each anonymous set needs to contain at least a preset number S of stations; for the station representative matrix within each anonymous set, the maximum value is taken in each time slot to generate an aggregation matrix M as part of the control information 3) Pattern aggregation within the anonymous set: For the representative matrices of the sites within each anonymous set, take the maximum value in each time slot to generate an aggregation matrix M, which is used as a part of the regulation information to cover the pattern characteristics of all traffic within the anonymous set; at the same time, use the historical traffic data to count whether there is data packet flow in each time slot, and construct a probability matrix P with the same shape as the aggregation matrix to reflect the probability of data flow occurrence in each time slot, and provide data basis for the optimization of the subsequent regulation matrix; 4) Simulation regulation optimization: Introduce learnable parameters w and b with the same shape as the aggregation matrix and a preset adjustment amplitude τ, and perform shrinkage optimization similar to linear regression on the aggregation matrix to obtain the regulation matrix. That is The design of the loss function aims to simulate the actual regulation state, comprehensively consider the delay loss and bandwidth loss, and achieve the balanced optimization between the two; simulating the actual regulation is to randomly sample the probability matrix using a preset sampling threshold to generate a mask matrix, and then calculate the actual regulation matrix with the regulation matrix. And reduce the error with the actual by averaging multiple samplings; for the delay loss part, the mean square error is used to measure the difference between the regulation matrix and the actual traffic matrix in each time slot, that is For the bandwidth loss part, it is used to calculate the absolute deviation between the optimized regulation matrix and the total bandwidth of the actual traffic, that is where B is the total bandwidth; the two are weighted and summed to obtain the final loss value, that is Among them, λ ∈ (0, 1) is used to regulate the optimization focus of the delay and bandwidth overhead.

3. A real-time traffic regulation method for network fingerprint defense according to claim 1, characterized in that After the system at the client or the server receives and updates the regulation information in the traffic regulation step, it will include the following sub-steps: 1) Start to execute regulation: After the system detects that the user is about to or has started to access a certain target site, allocate a real-time regulation state machine for this access and record the access information; 2) Initialize regulation: According to the mapping relationship between the target site and the anonymous set, retrieve and match the corresponding regulation matrix; for a new site, a default or randomly selected anonymous set can be used, and the mask matrix is determined by cumulative sampling using the disrupted probability matrix, so as to generate a real-time regulation matrix; 3) Real-time regulation processing: Within each fixed time slot period, the state machine reads the number of data packets from the buffer and compares it with the preset number of packets for the corresponding time slot in the real-time regulation matrix. If the actual number of packets exceeds the total number of data packets from the current time slot to the next e time slots, congestion mitigation is initiated, and data packets are immediately sent until the upper limit or the mitigation threshold is reached, and congestion events are recorded. If it is insufficient, virtual data packets are filled according to the regulation matrix. When the filled virtual data packets exceed the preset filling threshold, the overfilling time is recorded. If transmission is not allowed in the current time slot, the data to be sent is temporarily blocked. After each time slot ends, the state machine automatically enters the next time slot until the access is completed. In addition, random noise obeying a preset truncated normal distribution is additionally introduced to operations involving sending data packets to achieve an obfuscation effect, and the external manifestation is normal traffic that is approximately not traffic shaped. 4) After the access ends, the state machine updates the regulation information based on the page information and the feedback of congestion mitigation and overfilling counts, so as to dynamically adjust the regulation strategy in subsequent cycles.