Defense method for maximum gain attack and derived adaptive attack thereof
By identifying and correcting abnormal data, identifying attack target items and filtering fake users, the frequency estimation problem of OUE protocol under maximum gain attack and adaptive attack is solved, and effective recovery of frequency estimation results and reliability of data analysis is achieved.
Patent Information
- Application Number
- CN202510429525.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-08
AI Technical Summary
In data poisoning attacks, the OUE protocol is vulnerable to maximum gain attacks (MGA) and its derived adaptive attacks (MGA-A), resulting in the frequency estimation results being manipulated, affecting the reliability of data analysis.
A defense method is proposed to identify and correct abnormally perturbed data, identify target items of attack, and filter out fake users without the need for additional trusted parties, delete fake user data to restore frequency estimation results.
It effectively reduces the error caused by malicious attacks, enhances the robustness of the OUE protocol, improves the security and accuracy of the data collection system, and ensures the accuracy of frequency estimation.
Smart Images

Figure CN120165953A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information security technology, and particularly relates to a defense method against the maximum gain attack and its derived adaptive attacks. Background Art
[0002] How to ensure that personal sensitive information is not leaked while publishing and analyzing data, and how to ensure user privacy and data security have become major challenges that need to be solved urgently. As a privacy protection model, Centralized Differential Privacy (CDP) strictly defines the strength of privacy protection, making the statistical calculation results of the data set have little impact on the existence or non-existence of a single record, thus protecting the privacy of individual data.
[0003] Local Differential Privacy (LDP) is a variant of differential privacy, aiming to provide verifiable privacy protection for users while allowing untrusted servers to collect and aggregate statistical data from distributed users. Under the LDP framework, users perturb their privacy data locally and then send the perturbed data to an untrusted central server, which then extracts the required statistical information from these data. Under the LDP mechanism, the OUE (Optimal Unary Encoding) protocol is widely used in frequency estimation tasks due to its efficient encoding method, which allows the server to perform accurate statistical aggregation of data while protecting user privacy, and has become one of the most commonly used LDP protocols.
[0004] However, since OUE depends on distributed user data, it is vulnerable to the threat of poisoning attacks. Attackers can interfere with the statistical results of the protocol by injecting malicious users and submitting specifically constructed data.
[0005] In today's data-driven era, many key decisions rely on the analysis and learning of the collected data. For example, in financial risk control, banks and financial institutions use users' transaction records and credit data to evaluate loan risks; in medical diagnosis, hospitals rely on patient data to train AI to assist doctors in disease prediction and treatment plan recommendation; in autonomous driving systems, vehicles train models through massive environmental perception data to optimize decisions and improve driving safety.
[0006] However, if this data is subjected to a data poisoning attack (Poisoning Attack), an attacker can inject carefully designed malicious data to make the machine learning model make wrong decisions. For example, in the field of financial risk control, an attacker may forge a large number of false credit records, causing the system to misjudge high-risk users as low-risk users, resulting in incorrect lending decisions; in the field of medical AI, an attacker can manipulate the training data to make the model mispredict the probability of a certain disease, affecting the diagnosis results; in autonomous driving, malicious data may cause the system to misidentify traffic signals and even make dangerous driving decisions.
[0007] In the data poisoning attack against the OUE protocol, the Maximal Gain Attack (MGA) models the attack as an optimization problem, aiming to generate false user data by maximizing the objective function, so as to maximize the attack utility and manipulate the frequency estimation result.
[0008] In summary, the objective of the present invention is to propose a defense method in the frequency estimation scenario after being attacked by MGA, to minimize the impact of the attack on the OUE protocol, and to eliminate malicious interference and restore the true distribution of the data as much as possible without disclosing user privacy. Summary of the Invention
[0009] The present invention aims to provide a defense method against the Maximal Gain Attack and its derived adaptive attacks, to achieve frequency recovery under the condition that the OUE protocol is attacked by MGA and MGA-A, to identify true and false users while determining the attacker's target item, to minimize the missed judgment of false users and the misjudgment of real users, and to achieve a better defense effect in terms of ensuring the frequency recovery effect.
[0010] Technical solution: A defense method against the Maximal Gain Attack and its derived adaptive attacks. Each data item in the OUE protocol is encoded as a d-dimensional binary vector, where only the position of the target item is 1 and the rest are 0. This method considers that under the MGA attack, the attacker controls the value of each bit of the uploaded binary vector to increase the target frequency or decrease the value of the non-target frequency, so as to maximize the deviation of the frequency estimation. The frequency recovery steps of this defense method include:
[0011] S1. Calculate the expectation l of the number of 1s in the vector submitted by real users for known parameters;
[0012] S2. Traverse all users, count and judge whether the number of 1s in the vector they submit is equal to l, and divide all users who submit data into two sets U1 and U2. Set U1 contains all false users and some real users, and set U2 contains all the remaining real users;
[0013] S3. Compare the distribution differences of each item in set U1 and set U2, and identify the different items as the attack target items. Specifically:
[0014] Estimate the frequencies of the data in sets U1 and U2 respectively, calculate the frequency differences of each item in sets U1 and U2, and regard the top l items with the largest differences as the possible attacked targets. Denote the set as And make the following judgments:
[0015] If the OUE protocol is under MGA attack, then for all user data in sets U1 and U2 respectively Count the 0 / 1 arrangement patterns of the bits, calculate the percentages of each arrangement pattern in sets U1 and U2 and find the difference between them. Identify the positions with value 1 in the arrangement pattern with the largest difference as the attacked target items, denoted as T I ;
[0016] If the OUE protocol is under MGA-A attack, then for all user data in sets U1 and U2 respectively Count the 0 / 1 arrangement patterns of the bits, calculate the percentages of each arrangement pattern in sets U1 and U2 and find the difference between them. Analyze the top k items with the largest differences, find the positions that are all 0 in the first three arrangements, regard them as error items, and denote the number of 1s in each arrangement pattern as The remaining items are identified as the attacked target item set, denoted as T I-A , and denote the total number of items in the set as The top arrangement patterns with the largest differences are denoted as
[0017] S4. According to the target item set denoted as T I Find the false users in reverse and delete them to restore the frequency. Specifically:
[0018] If the OUE protocol is under MGA attack, then move the users in set U1 whose T I bits are not all 1 to set U2, and identify the users remaining in set U1 as false users and do not include them in the statistical scope;
[0019] If the OUE protocol is under MGA-A attack, then move the users in set U1 whose T I-A bits do not match each distribution in to set U2, and regard the users remaining in set U1 as false users and do not include them in the statistical scope.
[0020] Furthermore, in the OUE protocol suffering from the maximum gain attack and its derived adaptive attacks, by constructing fake users, the target position in the perturbation response sent by them is set to 1, and at the same time, l - 1 positions are randomly selected from the remaining positions and also set to 1, so as to form an l-hot vector with 1s in l positions, where l is the average expectation of the number of 1s in the vector submitted by real users, and hot represents the one-hot vector, specifically referring to that there are l positions with 1s in the vector submitted by the attacker and 0s in other positions.
[0021] Furthermore, in step S2 of the method, specifically: the set of all users used is denoted as U, and the defender will traverse all users and count the number of 1s u in the vector submitted by each user i , and compare the number of 1s u submitted by each user i with l. The users with equal numbers are added to the set U1, and the users with unequal numbers are added to the set U2.
[0022] The MGA attack refers to the maximum gain attack, and the MGA-A refers to the adaptive attack derived from the MGA attack.
[0023] Beneficial effects: Considering that in a decentralized data collection environment, there are some malicious users trying to manipulate the frequency estimation results by injecting forged data, affecting the reliability of data analysis. The present invention designs a defense mechanism to statistically analyze user data, identify and correct abnormally perturbed data, and ensure the accuracy of the final frequency estimation. After the server side collects and aggregates user data, applying the defense method of the present invention can effectively reduce the error caused by malicious attacks. The present invention can enhance the robustness of the OUE protocol without an additional trusted party, resist MGA and MGA-A attacks, and improve the security and accuracy of the data collection system. The present invention is applicable to large-scale distributed data collection environments and can maintain a high frequency estimation accuracy under different ratios of malicious users, target perturbation numbers, and privacy budgets. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is the overall flowchart of the present invention;
[0025] Figure 2 is the schematic diagram of the specific process on the defender side of the present invention;
[0026] Figure 3 is the experimental result (Fire dataset) of different ratios β of fake users for the OUE protocol suffering from MGA;
[0027] Figure 4 is the experimental result (Fire dataset) of different numbers r of target items for the OUE protocol suffering from MGA;
[0028] Figure 5Are the experimental results of the OUE protocol suffering from MGA with different privacy budgets ε (Fire dataset);
[0029] Figure 6 Are the experimental results of the OUE protocol suffering from MGA with different proportions β of fake users (IPUMS dataset);
[0030] Figure 7 Are the experimental results of the OUE protocol suffering from MGA with different numbers r of target items (IPUMS dataset);
[0031] Figure 8 Are the experimental results of the OUE protocol suffering from MGA with different privacy budgets ε (IPUMS dataset);
[0032] Figure 9 Are the experimental results of the OUE protocol suffering from MGA-A with different proportions β of fake users (Fire dataset);
[0033] Figure 10 Are the experimental results of the OUE protocol suffering from MGA-A with different numbers r of target items (Fire dataset);
[0034] Figure 11 .Are the experimental results of the OUE protocol suffering from MGA-A with different privacy budgets ε (Fire dataset);
[0035] Figure 12 Are the experimental results of the OUE protocol suffering from MGA-A with different proportions β of fake users (IPUMS dataset);
[0036] Figure 13 Are the experimental results of the OUE protocol suffering from MGA-A with different numbers r of target items (IPUMS dataset);
[0037] Figure 14 Are the experimental results of the OUE protocol suffering from MGA-A with different privacy budgets ε (IPUMS dataset). Specific implementation manners
[0038] The above scheme will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are for illustrating the present invention and not for limiting the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] To describe in detail the technical solution disclosed by the present invention, the following further elaborates in conjunction with the accompanying drawings of the specification and specific implementation steps.
[0040] As an advanced LDP protocol, the OUE protocol locally encodes the participating users into binary vectors, perturbs their private data, and sends the perturbed data to an untrusted central server. Subsequently, the server aggregates the statistical items it is interested in from these perturbed data.
[0041] However, due to its distributed user design, the OUE protocol is vulnerable to data poisoning attacks. Attackers inject malicious users to send carefully designed data to interfere with the statistical results of the OUE protocol. Among them, the Maximal Gain Attack (MGA) defines the attack as an optimization problem, constructing the data of fake users by maximizing the objective function to maximize the attack utility.
[0042] Specifically, attackers can inject some fake users into the LDP protocol. These fake users can send any data in the coding space to the central server. Specifically, assume there are n real users in the system, and the attacker injects m fake users into it. Therefore, the total number of users becomes n + m. Since the OUE protocol performs the encoding and perturbation steps locally at the user side, the attacker can access the implementation details of these steps. Therefore, the attacker knows various parameters of the LDP protocol, that is, the attacker knows the domain size d, the coding space D, and the support set S(y) of each perturbation value y ∈ D.
[0043] In the OUE protocol, each data item is encoded as a d-dimensional binary vector, where only the position of the target item is 1 and the rest are 0. Under the MGA attack, the attacker can control the value of each bit of the uploaded vector to increase the target frequency or decrease the value of the non-target frequency, thereby maximizing the deviation of the frequency estimate. By constructing fake users, making the target position (i.e., the position corresponding to the value whose frequency the attacker hopes to increase) in the perturbed response they send be 1, and randomly selecting l - 1 other positions to also be set to 1, thus forming an l-hot vector with 1s in l positions, where l is the average expected number of 1s in the vectors submitted by real users. The core reason behind this design is to enhance the concealment of the attack: if the fake user only sets 1 at the target position, its impact on the target frequency is strong but easy to detect; while when the target position is mixed among multiple positions with 1s, the attack behavior is not easily discovered, and the frequency estimate of the target value in the aggregation result can still be significantly increased. Therefore, the number of 1s is designed to be l to balance the attack effect and concealment. This manipulation misleads the final statistics by increasing the proportion of 1 bits in the perturbation probability, making the frequency of the attack target significantly higher or lower.
[0044] On this basis, the attacker further proposed an adaptive attack method MGA-A against the OUE protocol. By introducing an adaptive strategy, more refined grouping management is carried out on forged users, thereby improving the attack efficiency. In MGA-A against the OUE protocol, in order to better conceal the attack behavior, the attacker no longer attacks all target items simultaneously, but randomly selects a subset of its target item set to conduct the attack. Similarly, in order to better hide the behavior of fake users, the attacker still designs the number of 1s in its submitted vector to be l, that is, the average expected number of 1s in the vectors submitted by real users. In this way, the attacker can more effectively utilize the forged user resources, maximize the attack gain, and further improve the concealment at the same time.
[0045] In response to the above attacks, the present invention provides a defense method against the maximum gain attack and its derived adaptive attacks, which solves the frequency recovery problem of the OUE protocol under the maximum gain attack and its derived adaptive maximum gain attacks, and is used to identify the target items under attack and screen out fake users on this basis for subsequent processing. The implementation steps of the frequency recovery problem of the OUE protocol under the maximum gain attack and its derived adaptive maximum gain attacks are as follows:
[0046] S1. Considering that the number of 1s in the vectors submitted by the attacker is constant and known, in order to avoid detection by the server, the attacker will set the number of 1s in each binary vector submitted to the average expected number l of 1s among real users when the defense party is against the MGA attack suffered by OUE and its derived MGA-A attack.
[0047] S2. The defense party divides all users who submit data into two sets. One is a set U1 that contains all fake users and some real users, and the other is a set U2 that contains all the remaining real users. The specific steps for dividing users after OUE suffers from MGA and MGA-A are as follows:
[0048] S21. The defense party will traverse all users (denote the set of all users as U) and count the number of 1s u in the vectors submitted by each user. i 。
[0049] S22. Compare whether the number of 1s u submitted by each user i is equal to l. Add the users with equal numbers to set U1, and add the users with unequal numbers to set U2.
[0050] S3. Compare the proportion differences of each item in the set U1 of "true and false users" and the set U2 of "all true users", find out the items with obvious differences, and identify them as attack target items to form the target item set. The process by which the defense party identifies the target items under MGA or MGA-A attack is as follows:
[0051] S31. Estimate the frequencies of the data in sets U1 and U2 respectively, and calculate the frequency differences of each item in U1 and U2 item by item.
[0052] S32. Since the attacker can attack at most l targets, the top l items with the largest differences are regarded as the targets that may be attacked, and the set is denoted as (which includes the targets actually attacked by the attacker and some error terms).
[0053] S33. If the OUE protocol is under MGA attack, then for all user data in U1 and U2 respectively Count the 0 / 1 arrangement patterns of the bits, calculate the percentages of each arrangement pattern in U1 and U2 and find the difference between them. The positions with value 1 in the arrangement pattern with the largest difference are identified as the attacked target items, denoted as T I .
[0054] S34. If the OUE protocol is under MGA-A attack, then for all user data in U1 and U2 respectively Count the 0 / 1 arrangement patterns of the bits, calculate the percentages of each arrangement pattern in U1 and U2 and find the difference between them. Analyze the top k items with the most differences, find the positions that are all 0 in the first three arrangements, regard them as error terms, and denote the number of 1s in each arrangement pattern as The remaining items are identified as the attacked target item set, denoted as T I-A , and denote the total number of items in the set as Denote the top arrangement patterns with the largest differences as
[0055] S4. Reverse find the eligible "fake" users according to the identified attacked target item set and delete them to restore the frequency. The process for the defender to lock the "fake" users based on the attacked target item set is as follows:
[0056] S41. If the OUE protocol is under MGA attack, then move the users in U1 whose bits corresponding to T I are not all 1 to set U2, and identify the users remaining in U1 (i.e., the users whose bits corresponding to T I are all 1) as "fake" users and do not include them in the statistics.
[0057] Among them, the bits corresponding to T I mean that the set contains the corresponding items. For example, if T = {1, 3}, it corresponds to the first and third bits in the OUE vector, and a vector like 10100 will be considered as submitted by a fake user.
[0058] S42. If the OUE protocol is under MGA-A attack, then for T in U1I-A bitwise AND Users whose distributions do not match any of the above are moved to the set U2, and the users remaining in U1 are regarded as "false" users and not included in the statistics.
[0059] Embodiment 2: Combining Figure 2 In the frequency recovery method of the OUE protocol under the maximum gain attack and its derivative adaptive maximum gain attack, the operation process steps of the defender are as follows:
[0060] Step 1: Calculate the expected value l of the number of 1s in the vector submitted by the real user, where the parameters required for the calculation are known and public.
[0061] Step 2: Traverse all users, count and determine whether the number of 1s in the vector they submit is equal to l, and divide the user set into U1 and U2.
[0062] Step 3: Compare the proportion differences of each item in U1 and U2 to confirm the target item set T I .
[0063] Step 4: Reverse-screen "false users" according to the target item set T I .
[0064] Step 5: Estimate the recovered frequency.
[0065] The following are the experimental scenario settings and experimental results of the present invention. The simulation experiment uses two real datasets, Fire and IPUMS. Among them, the Fire dataset filters the data in its "Alarms" label, including 747,554 users and 315 classifications. The IPUMS dataset contains 1,608,982 users and 206 classifications. At the same time, it is assumed that in the MGA-A attack against the OUE protocol, the difference between the subset size selected by the attacker and the target item set size is always 2. Tables 1, 2, and 3 respectively list the specific values of the false user proportion β, the number of target items r, and the privacy budget ε in the test experimental groups. The false user proportion refers to the proportion of false users among all users, the number of target items refers to the number of targets in the target item set designed by the attacker, and the privacy budget refers to the privacy budget used in the OUE protocol.
[0066] Table 1. False user proportion β test experimental group
[0067]
[0068] Table 2. Number of target items r test experimental group
[0069]
[0070] Table 3. Privacy budget ε test experimental group
[0071]
[0072] To evaluate the effectiveness of the design scheme of the present invention, next, under the settings of different proportions of fake users β, the number of target items r, and the privacy budget ε, the frequency recovery of the OUE protocol suffering from MGA and MGA-A will be carried out. Then, the experimental results before and after defense for the same technology ratio with different proportions of fake users β, the same number of target items r, and the privacy budget ε, different numbers of target items r with the same proportion of fake users β and the privacy budget ε, and different privacy budgets ε with the same proportion of fake users β and the number of target items r will be compared and analyzed to obtain the estimation error MSE. To estimate the frequency, f v For the original data frequency, D and d are respectively the coding domain of the OUE protocol and its size, and v is any item in the coding domain.
[0073] Figures 3 to 8 In it, the vertical axis represents the estimation error, that is, MSE; the horizontal axis represents different experimental variables, that is, different settings of the proportion of fake users β, the number of target items r, and the privacy budget ε. Curves of different colors represent the changes of the error with the experimental variables before and after using the defense method of the present invention. Among them, the blue represents the change of the error of the OUE protocol suffering from the MGA attack, and the orange represents the change of the error of the OUE protocol suffering from the MGA attack after using the defense method of the present invention to recover the frequency.
[0074] Analyzing the experimental results of the two data sets, the following conclusions can be drawn: In experiments with different proportions of malicious users ( Figure 3 and Figure 6 ) and different numbers of target interferences ( Figure 4 and Figure 7 ), it can be clearly seen that the mean square error (MSE) without defense increases significantly with the increase of the proportion of fake users and the number of target items. Under the action of the defense mechanism, MSE always remains at a low level, proving the effectiveness of the present invention against malicious attacks. In experiments with different privacy budgets ( Figure 5 and Figure 8 ), MSE gradually decreases with the increase of the privacy budget, indicating that the increase of the privacy budget can improve the data accuracy. However, the MSE under the defense mechanism decreases faster, indicating that this defense strategy can more effectively improve the result accuracy.
[0075] In all experimental scenarios, the defense strategy can effectively reduce the estimation error. Especially in the case of a relatively high proportion of fake users and a relatively large number of target items, the effect of error reduction is more obvious. This indicates that this mechanism has stronger robustness in an environment with more serious malicious attacks.
[0076] Combined with Figures 9 to 14, the vertical axis represents the estimation error, i.e., MSE; the horizontal axis represents different experimental variables, i.e., different settings of the proportion of fake users β, the number of target items r, and the privacy budget ε. Curves of different colors represent the changes in error with the variation of experimental variables before and after using the defense method of the present invention. Among them, the blue color represents the error change of the OUE protocol suffering from the MGA-A attack, and the orange color represents the error change of the OUE protocol suffering from the MGA-A attack after restoring the frequency using the defense method of the present invention.
[0077] Analyzing the experimental results of the two datasets, the following conclusions can be drawn:
[0078] From Figure 9 and Figure 12 It can be seen that as the proportion of fake users β increases, the mean square error (MSE) without defense increases exponentially, while under the action of the defense mechanism proposed by the present invention, the MSE always remains at a low level. This indicates that the present invention can effectively resist the MGA-A attack against OUE and improve data accuracy, especially when the proportion of malicious users is relatively high, the defense effect is more significant.
[0079] Figure 10 and Figure 13 show that in the case of different numbers of target items r, the MSE of the non-defense scheme increases significantly with the increase of r, while the defense mechanism proposed by the present invention can effectively suppress the growth of errors and ensure the stability of data estimation. This shows that the present invention can adapt to more complex attack scenarios and still maintain a high level of accuracy when multiple targets are interfered.
[0080] Figure 11 and Figure 14 The experimental results show that as the privacy budget increases, the MSE gradually decreases, meaning that a larger privacy budget can improve data accuracy. In addition, the downward trend of the MSE under the defense mechanism is more obvious, indicating that the present invention can not only provide effective protection under low privacy budgets, but also further improve data quality in a higher privacy budget environment.
[0081] In summary, under all experimental settings, the defense mechanism significantly reduces the estimation error, especially in the cases of a relatively high proportion of malicious users, a large number of target items, or a low privacy budget, and the error reduction effect is particularly obvious. This indicates that the mechanism can effectively enhance the robustness of the OUE protocol and improve the reliability of data estimation in various attack scenarios.
[0082] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A defense method against maximum gain attack and its derived adaptive attack. In the OUE protocol, each data item is encoded as a d-dimensional binary vector, where only the position of the target item is 1 and the rest are 0. This method considers that under the MGA attack, the attacker increases the target frequency or decreases the value of the non-target frequency by controlling the value of each bit of the uploaded binary vector, thereby maximizing the deviation of the frequency estimation. The method is characterized by: The method includes the following steps for frequency recovery: S1. Calculate the expected number of 1s in the vector submitted by the real user for the known parameters in the OUE protocol. S2, traverse all users, count and determine whether the number of 1s in the vectors they submit is equal to l, and divide all users who submit data into two sets U1 and U2. Set U1 contains all fake users and some real users, and set U2 contains all the remaining real users; S3. Compare the distribution differences of each item in set U1 and set U2, and identify the difference items as attack target items, specifically: The frequency of the data in sets U1 and U2 is estimated respectively, and the frequency difference between them is calculated item by item. The first l items with the largest difference are regarded as possible targets of attack, and their set is recorded as And make the following judgment: If the OUE protocol is attacked by MGA, all user data in sets U1 and U2 will be The 0 / 1 arrangement patterns of the bits are counted, and the percentage of each arrangement pattern in sets U1 and U2 is calculated, and the difference between the two is calculated. The position with a value of 1 in the arrangement pattern with the largest difference is confirmed as the target item to be hit, which is recorded as T I ; If the OUE protocol is attacked by MGA-A, all user data in sets U1 and U2 will be The 0 / 1 arrangement patterns of the bits are counted, and the percentage of each arrangement pattern in sets U1 and U2 is calculated, and the difference between the two is calculated. The first k items with the largest difference are analyzed, and the positions that are all 0 in the first three arrangements are found, which are regarded as error terms. The number of 1s in each arrangement pattern is recorded as The remaining items are confirmed as the target item set to be hit, denoted as T I-A , and the total number of items in the collection is recorded as The largest difference The arrangement pattern is denoted as S4, according to the target item set is recorded as T I Reverse and find out the fake users and delete them to restore the frequency, specifically: If the OUE protocol is attacked by MGA, then the T I Users whose bits are not all 1 are moved to set U2, and users who remain in set U1 are identified as fake users and are not included in the statistics; If the OUE protocol is attacked by MGA-A, then the T I-A Position and Users that do not match any of the distributions in are moved to set U2, and users who remain in set U1 are considered false users and are not included in the statistics.
2. The defense method according to claim 1, characterized in that: In the maximum gain attack and its derived adaptive attack on the OUE protocol, a fake user is constructed so that the target position in the perturbation response sent by the fake user is 1, and l-1 positions are randomly selected from the remaining positions and set to 1, thereby forming an l-hot vector with l positions as 1, where l is the average expectation of the number of 1s in the vector submitted by the real user, and hot represents the hot encoding vector, which specifically means that l positions in the vector submitted by the attacker are 1 and the others are 0.
3. The defense method according to claim 1, characterized in that: Step S2 is as follows: the set of all users is recorded as U, and the defender will traverse all users and count the number of 1s in the vector submitted by each user u i , compare the number of 1s submitted by each user u i Whether it is equal to l, the equal users are added to set U1, and the unequal users are added to set U2.
4. The defense method according to claim 1, characterized in that: The MGA attack refers to a maximum gain attack, and the MGA-A refers to an adaptive attack derived from the MGA attack.
Citation Information
Patent Citations
Utility optimization set data protection method based on local differential privacy
CN115130119A
Method for dynamically measuring privacy protection effect of local differential privacy mechanism
CN117195281A
End-to-end learning method for communication interference integration
CN118041486A
Vehicle position privacy protection method and system based on local differential privacy
CN119545333A
Private joining, analysis and sharing of information located on a plurality of information stores
WO2022251399A1