High-efficiency localization differential privacy data processing method capable of resisting data poisoning attack

Through Pederson's commitment technology, the user-side data is disturbed and checked on server-side data, which solves the privacy protection and operation efficiency of existing protocols under data poisoning attacks, and achieves efficient data processing and security improvement.

CN120342735APending Publication Date: 2025-07-18NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510590429.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing local differential privacy protocol cannot effectively take into account privacy protection and operational efficiency when facing data poisoning attacks. Especially in big data processing scenarios, GRR's privacy protection effect is insufficient, and VGRR's operating speed bottleneck is obvious, making it difficult to meet the needs of real-time and security.

Method used

Pederson commitment technology is used to process the user side data, generate a commitment vector and upload it to the server side, and the server side checks and saves the perturbation results that meet the requirements, and calculates the mean through an optimized data aggregation algorithm to improve the operating efficiency and security of the protocol.

Benefits of technology

Without leaking user data, the balance between result accuracy and calculation overhead is achieved, saving promised opening time and data transmission time compared with VGRR, improving operational efficiency, especially in the case of large-number domains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342735A_ABST
    Figure CN120342735A_ABST
Patent Text Reader

Abstract

The invention discloses a localized differential privacy protocol aiming at a data poisoning attack and a data processing method, and belongs to the technical field of information security. The method mainly comprises the steps that a server side sets the number field size, the privacy budget and the threshold value of a mean value estimation item based on actual needs, calculates a length parameter and then publishes information to all user sides; the user side carries out privacy protection processing on single attribute data and uploads a result to the server side; and the server side checks the disturbance result, saves data meeting requirements, and finally performs statistical analysis on privacy protection processing results to calculate a mean value of all user privacy protection processing results. Through the optimized data aggregation and analysis algorithm, on the premise of guaranteeing data privacy, rapid data transmission and data information extraction are achieved, and the method is superior to a traditional local differential privacy protocol in the aspects of privacy protection intensity, information utilization rate and protocol operation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information security technology, and relates to a data privacy protection and processing method under a local differential privacy protocol, in particular to an efficient local differential privacy data processing method against data poisoning attacks. Background Art

[0002] With the rapid development of information technology, people's ability to acquire, store, and analyze data has been continuously enhanced, and the global data has grown explosively and the aggregation degree has become higher and higher. Therefore, big data computing technology has emerged as the times require. It is a series of methods designed specifically for processing extremely large or extremely complex data sets, aiming to extract knowledge from large-scale data using advanced data analysis techniques.

[0003] In a distributed computing environment, data needs to be processed on multiple nodes, which makes the system face the risk of being attacked and easily leads to serious security problems such as data leakage, tampering, or forgery. In addition, big data processing tasks usually involve a large amount of personal information, such as medical records, financial transaction records, etc. The leakage of this information may cause extremely serious losses to individuals, enterprises, and even countries in terms of economy, reputation, etc. As the computational problems caused by privacy leakage in big data statistical analysis are increasing. With the wide application of information technology in various fields, the data collection and analysis scenarios are becoming increasingly frequent, and data privacy protection has become a crucial issue.

[0004] As an effective means to protect user data privacy, when adding privacy protection technology in distributed computing, the resulting loss of result accuracy and computational overhead must be comprehensively considered. The local differential privacy technology perturbs the data at the data source, so that even if the data is maliciously obtained during sharing, it is difficult to infer the sensitive information of individuals. Among them, GRR (Generalized Randomized Response) and VGRR (Vectorial Generalized Randomized Response) are relatively classical local differential privacy protocols.

[0005] GRR protects data privacy to a certain extent by performing randomized response on the data. However, in complex scenarios, its privacy protection effect still has room for improvement. Especially when facing possible attacks, it is difficult to provide sufficient security guarantees. VGRR is designed for high-dimensional data, optimizes the data processing process, and protects the data privacy of data owners and the data security of data collectors at the same time. However, it has obvious bottlenecks in operating efficiency.

[0006] With the continuous expansion of data scale and the increasing requirements for real-time performance in application scenarios, the limitations of existing local differential privacy protocols have gradually become prominent. In practical applications, many scenarios not only require strict data privacy protection but also have high requirements for the running speed of the protocol. For example, in scenarios such as mobile device data collection and online questionnaire surveys, the amount of user data is large and rapid processing and analysis are required. If the privacy protection protocol runs too slowly, it will seriously affect the user experience and data processing efficiency; while in scenarios such as the collection of financial transaction data and medical and health data involving sensitive information, the requirements for privacy protection effects are extremely high, and existing protocols cannot fully meet their security needs.

[0007] Therefore, there is an urgent need to propose a new local differential privacy protocol that can overcome the deficiencies of the GRR privacy protection effect and make up for the defects of the VGRR running speed, while ensuring data privacy, improving the running efficiency of the protocol, and adapting to diverse data collection and analysis scenarios.

[0008] In summary, the objective of the present invention is to propose a secure and efficient local differential privacy protocol for scenarios that require both privacy protection and running efficiency, so as to improve the result accuracy and running efficiency. Summary of the Invention

[0009] Objective of the Invention: The present invention provides an efficient local differential privacy data processing method against data poisoning attacks, which can take into account both privacy protection and running efficiency.

[0010] Technical Solution: An efficient local differential privacy data processing method against data poisoning attacks, the processing steps of which for privacy protection include:

[0011] S1. The server side sets the number field size d, privacy budget ε, and threshold α of the mean estimation item based on actual needs, and calculates the length parameters l, l1, and l2, and then publishes this information to each user side participating in the mean estimation item;

[0012] The calculation method of the length parameter is: first, find the integer solutions of l1 and l2 that satisfy the relationship of αε ≤ ln(l1 / l2) < ε, and then calculate l according to l = l1 + (d - 1)l2;

[0013] S2. Each user side participating in the project perturbs the single-attribute data according to the received requirements and uploads the processing result to the server side;

[0014] The perturbation processing is as follows: First, a value k is randomly selected from the number field size d. Then, l1 - l2 vs, l2 values of k, and l - l1 null values are randomly combined into an original perturbation vector μ. Next, Pedersen commitment is used to commit to each non-null value item to obtain a commitment vector η. The user records the randomly selected value k by themselves and then sends the commitment vector η to the server; the null value is represented as -1.

[0015] S3. The server checks the privacy-processed data uploaded by each client and saves the perturbation results that pass the check.

[0016] S4. The server conducts statistical analysis based on the privacy protection processing results sent by the client, counts the values of the perturbation results of all clients, and calculates the estimated mean of the original data based on this.

[0017] Preferably, in step S2, the method of sending the commitment vector η to the server includes sending only the positions and corresponding values of non -1 (null values) to the server in the form of a matrix.

[0018] Furthermore, in the above method, the server's check and selection of the received commitment vector η are as follows:

[0019] S31. The server randomly selects a position j and returns j to the client.

[0020] S32. The client opens and returns the selected position j and l1 - 1 other positions to the server. The requirements for the returned opened commitments are: the first one must be the position j selected by the server, and among these returned values, there must be l1 - l2 null values and l2 opened identical non-null values. If the position j selected by the server corresponds to a null value, then the commitment corresponding to the recorded value k is opened.

[0021] S33. The server checks whether the opened commitments meet the requirements. If they do not meet the requirements, they are discarded.

[0022] S34. If the position j selected by the server is not a null value, the original value of the commitment opened at this position is used as the output; otherwise, a value is randomly selected from the number field size d other than the original value corresponding to the commitment opened this time as the output.

[0023] Furthermore, in the above method, the way the server calculates the estimated mean of the original data is: The server generates a data set based on all the received perturbation processing results, and then aggregates the data according to the GRR aggregation formula: (count - nq) / (p - q), where p = l1 / l, q = (1 - p) / (d - 1), n is the total number of data, and count is the number of times the corresponding item appears.

[0024] Beneficial effects: The present invention achieves a balance between result accuracy and computational overhead without disclosing the individual attribute data of each client. At the same time, compared with VGRR, it saves the time required for l1 - l2 times of opening commitments and part of the time required for transmitting vectors, improving the operating efficiency, which is particularly obvious when the number field size d and l2 are large. Brief Description of the Drawings

[0025] Figure 1 It is the privacy protection flowchart of the method of the present invention;

[0026] Figure 2 shows the OG results of different settings of Bangumi scores and TikTok Google Play scores in the embodiment;

[0027] Figure 3 shows the MSE results of different settings of Bangumi scores and TikTok Google Play scores in the embodiment;

[0028] Figure 4 is a comparison of the running times for Bangumi scores and TikTok Google Play scores based on the present invention. Detailed Embodiment

[0029] The following further illustrates the above - mentioned solution with specific embodiments. It should be understood that these embodiments are for illustrating the present invention and not for limiting the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0030] First, the present invention solves the problem that traditional GRR and VGRR cannot balance result accuracy and computational overhead due to security or efficiency issues. In addition, the invention can resist the maximum background knowledge attack of an attacker, prevent the attacker from inferring or intercepting the real data of the participating parties from the server side, so that the participating computing parties can avoid the risk of privacy leakage; it can also prevent the attacker from constructing malicious data to attack the final result. The present invention is a local differential privacy protection technology (protocol), and the client then performs privacy protection processing on the individual attribute data based on this technology and uploads it to the server side, ensuring that the data privacy security is higher than that of the GRR protocol and also ensuring that it is superior to the VGRR protocol in terms of overall computing efficiency.

[0031] Combined with the above - mentioned technical solution, combined with Figure 1 , the implementation process of the present invention is as follows:

[0032] (1) The operation process of the client is as follows:

[0033] (1) According to the received required parameters, perform privacy protection processing on the individual attribute data to generate the original vector μ, the randomly selected number k, and the commitment vector η;

[0034] (2) Upload the commitment vector η to the server.

[0035] (3) Open the commitment according to the server's requirements and submit it to the server.

[0036] (2) Server-side operation process is as follows:

[0037] (1) The server generates and calculates relevant parameters based on the actual situation and sends them to the client;

[0038] (2) After receiving the commitment vector, randomly select a position and send it to the client;

[0039] (3) Verify the commitment vector and the corresponding commitment value according to the requirements. If it passes, select the output value and record it in the data set;

[0040] (4) Aggregate the data.

[0041] Using this local differential privacy protocol for privacy protection processing, the data needs to ensure that the perturbation result cannot be modified, the running efficiency should be improved, and the data privacy must be protected. Therefore, the effectiveness of the present invention will be evaluated through theoretical explanations, mathematical derivations, and experiments below, and the scientific credibility of the present invention in the use process will be measured.

[0042] It should be noted theoretically that: since the user transmits a commitment vector, no valid information can be obtained before opening the commitment; since the vector is mixed with randomly selected values by the user, it is impossible to determine whether the opened commitment is the user's true value; since it is required to force the construction of the vector for harmless opening, poisoning attacks on the output data can be prevented.

[0043] The following is the mathematical proof of the present invention:

[0044] Theorem 1: For the following parameters l, l1, l2, if they satisfy l = l1 + (d - 1)l2 and l1 > l2, the new protocol satisfies ln(l1 / l2)-LDP.

[0045] Proof: From the given parameters and the construction process of the μ vector, it can be seen that the composition of the μ vector includes: (l1 - l2) real items V, l2 random items K, and l - l1 null values (-1). Then, in the process of random selection by the server, the probability of randomly selecting item V is:

[0046]

[0047] And the probability of randomly selecting other values is:

[0048]

[0049] Therefore, it satisfies:

[0050]

[0051] Therefore, the new protocol satisfies ε'-LDP, where ε' = ln(l1 / l2).

[0052] Theorem 2: As a verifiable LDP protocol, the new protocol is indistinguishable.

[0053] Proof: Since the Pedersen commitment guarantees the hiding of the original value, the value v of the original item cannot be inferred from the commitment scheme; moreover, since the user can arbitrarily select a random item K for authenticity checking, then K is not the user's item, so the server cannot know the user's true item.

[0054] Theorem 3: The new protocol is robust.

[0055] Proof: Since the protocol has a limit on the number of null values, the malicious user U can only construct a vector μ with l - l1 null values. Thus, when the malicious user U attempts an input poisoning attack on the target item, there can be at most l1 target items in each μ vector, and at this time, the attack effect is the same as that of the input poisoning attack.

[0056] Based on the above theorems, the privacy and security of the new protocol are established. From Theorem 1, it can be obtained that the new protocol satisfies ln(l1 / l2)-LDP, which is the most basic property of the new protocol. From Theorem 2, it can be seen that except for the perturbed data, the aggregator cannot obtain any information about the original items of the users; from Theorem 3, it can be known that the new protocol can defend against OPA well.

[0057] The following are the experimental scenario settings of the present invention and the experimental results obtained.

[0058] The simulation experiment uses two real-world datasets, namely: TikTok Google Play ratings, which contains 454,047 data entries, with ratings being integers from 0 to 5, and the attack target contains one data item; Bangumi ratings, which contains 126,908 data entries, with ratings being integers from 1 to 10, and the attack target contains two data items. During the experiment, the proportion of m malicious users among all n + m users is denoted as The default proportion is set to 0.05, i.e., β = 0.05; the privacy budget ε is set to 1.6, and the threshold α is set to 0.99. The parameters l1, l2, and l in VGRR are generated by solving αε ≤ ln(l1 / l2) < ε, with the method of generating length parameters in VGRR being used, where d depends on the specific dataset.

[0059] In this embodiment, the effects are compared and analyzed through two indicators: the overall gain (OG) and the mean squared error (MSE). The overall gain (OG) is a direct indicator used to reflect the intensity of the poisoning effect. OG measures the change in the estimated usage frequency of the target item after the attack compared to before the attack.

[0060] The mean squared error (MSE) is a way to indicate the accuracy of frequency estimation, expressed as: the average of the absolute error magnitudes between the estimated frequency after the attack and the actual frequency of all items.

[0061]

[0062] Among them, represents the estimated frequency of item i after the poisoning attack, represents the estimated frequency of item i before the poisoning attack, and f i represents the actual frequency of item i.

[0063] In the embodiment, all malicious users are grouped together as a whole. The total gain is the change in the estimation result caused by this whole. Here, it is defined as the total score, and its value is the sum of the scores of all target items.

[0064] Furthermore, we divide each dataset into two groups. Each group contains the input poisoning GRR (IPA - GRR), the output poisoning GRR (OPA - GRR), VGRR, and this invention (NEW). One group controls the privacy budget as a variable, and the other group controls the proportion of malicious users as a variable. Figure 2 shows the OG results for different settings. Figures 2(a) and 2(b) show how OG varies with the privacy budget ε. Figures 2(c) and 2(d) show how OG varies with the proportion of malicious users β.

[0065] Figure 3 shows the MSE results for different settings. Figures 3(a) and 3(b) show how MSE varies with the privacy budget ε. Figures 3(c) and 3(d) show how MSE varies with the proportion of malicious users β.

[0066] In Figures 2 and 3, the vertical axis represents the estimation error, and the horizontal axis represents the corresponding item change. It can be seen that the OG and MSE in VGRR, NEW, and the IPA baseline are almost the same. This means that VGRR has the same ability as NEW to defend against poisoning attacks.

[0067] In Figure 4, the vertical axis represents the running time in seconds, and the horizontal axis represents the privacy budget ε. Combining Figures 4(a) and 4(b), it can be seen that the running efficiency of NEW is higher than that of VGRR.

[0068] By analyzing the experimental results of two data sets, the following conclusions can be drawn: While conforming to the GRR protocol distribution, the present invention has the same defense performance as VGRR and higher operating efficiency than VGRR. In the case of a small privacy budget and a large data domain range, the present invention saves more time compared to VGRR.

[0069] In summary, the present invention has the same result accuracy as the VGRR protocol and is superior to the VGRR protocol in terms of computational overhead.

[0070] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An efficient local differential privacy data processing method against data poisoning attacks, characterized in that The processing steps for privacy protection of this method include: S1. The server sets the number field size d, privacy budget ε, and threshold α of the mean estimation item based on actual needs, calculates the length parameters l, l1, and l2, and then publishes this information to each client participating in the mean estimation project; The calculation method of the length parameter is: First, find the integer solutions of l1 and l2 that satisfy the relationship of αε ≤ ln(l1 / l2) < ε, and then calculate l according to l = l1 + (d - 1)l2; S2. Each client participating in the project perturbs the single-attribute data according to the received requirements and uploads the processing results to the server; The perturbation processing is: First, randomly select a value k from d, randomly form an original perturbation vector μ with l1 - l2 items v, l2 values k, and l - l1 null values, then use Pedersen commitment to commit to each non-null value item to obtain a commitment vector η. The user records the randomly selected value k by himself, and then sends the commitment vector η to the server; S3. The server checks the privacy processing data uploaded by each client and saves the perturbation results that pass the check; S4. The server performs statistical analysis based on the privacy protection processing results sent by the client, counts the values of the perturbation results of all clients, and calculates the estimated mean of the original data based on this; 2. The localization differential privacy protection and data processing method according to claim 1, characterized in that The way of sending the commitment vector η to the server in step S2 includes sending only the positions and corresponding values of non-null values to the server in the form of a matrix.

3. The localization differential privacy protection and data processing method according to claim 1, wherein The server's check and selection of the received commitment vector η are as follows: S31. The server randomly selects a position j and returns the position j to the client; S32. The client opens and returns the selected position j and other l1 - 1 positions to the server. The requirements for the returned opened commitments are: The first one must be the position j selected by the server, and there must be l1 - l2 null values and l2 opened identical non-null values among these returned values. If the position j selected by the server corresponds to a null value, then open the commitment corresponding to the recorded value k; S33. The server checks whether the opened commitments meet the requirements. If they do not meet the requirements, discard them; S34. If the position j selected by the server is not a null value, use the original value of the commitment opened at this position as the output; otherwise, randomly select one from the values in d other than the original value corresponding to the commitment opened this time as the output.

4. The localization differential privacy protection and data processing method according to claim 1, characterized in that The way for the server to calculate the estimated mean of the original data is: The server generates a data set based on all the received perturbation processing results, and then aggregates the data according to the aggregation formula of GRR: (count - nq) / (p - q), where p = l1 / l, q = (1 - p) / (d - 1), n is the total number of data, and count is the number of occurrences of the corresponding item.