Business Processing Method and Device Based on Privacy Protection

Through the two-stage business processing method, the server automatically distinguishes sensitive items from non-sensitive items, and the client performs differentiated high-low differential privacy processing, solving the balance between sensitive data privacy protection and result accuracy in the existing technology, and achieving more efficient privacy protection and accuracy.

CN114676457BActive Publication Date: 2025-07-25ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210299408.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2025-07-25
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

Existing differential privacy protocols use homogeneous processing when processing sensitive and non-sensitive data, making it difficult to adjust the balance between result accuracy and privacy protection, especially when the privacy protection needs of sensitive data are strong, conventional LDP protocols may lead to results bias and reduced availability.

Method used

Using a two-stage business processing method, first the server automatically distinguishes sensitive items from non-sensitive items based on the first disturbance data uploaded by the client, and then the client performs differentiated high and low differential privacy processing based on the detection results, adding different noises to sensitive data and non-sensitive data respectively.

Benefits of technology

It improves the accuracy of business processing and the rationality of privacy protection, avoids interference from human factors, and ensures efficient protection of sensitive data and the accuracy of non-sensitive data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114676457B_ABST
    Figure CN114676457B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a business processing procedure including two stages. A business processing method based on privacy protection is used to perform business processing related to target information, and multiple attribute items correspond to the target information. In the first stage, each client provides a single first perturbed data determined by the first differential privacy satisfied by the local attribute items of the target information to the server respectively. The server detects sensitive items and non-sensitive items in each attribute item of the target information according to each first perturbed data separately sent by each client, and feeds back the detection results to each client. In the second stage, each client samples the local attribute items of the target information to obtain corresponding second perturbed data that satisfies the second differential privacy based on the detection results and uploads them to the server; the server performs corresponding business processing on the target information based on each second perturbed data separately uploaded by each client.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of secure computing technology, and in particular, to a service processing method and apparatus based on privacy protection. Background Art

[0002] Differential Privacy is a mathematical technique that can add noise to data while calculating the degree of privacy, making the process of adding "noise" more rigorous. Generally, the stronger the privacy protection, the more noise is added, and the lower the accuracy and usability of the result. Conversely, the stronger the data usability, the less noise is added, and the higher the accuracy of the result and the weaker the privacy protection. Therefore, in the differential privacy process, a privacy budget can be given, and the amount of noise to be added can be determined based on the privacy budget, so as to balance accuracy and privacy.

[0003] Local Differential Privacy (LDP) can transfer the work of data privatization to each participant, and each participant processes and protects local data by itself, that is, desensitizes local data, further reducing the possibility of privacy leakage. However, in practice, even for the same type of data, the sensitivity may not be the same. For example, in the scenario of disease frequency statistics, a cold may be a disease that everyone has experienced, and the information about whether a person has had a cold can be provided to other parties, and its sensitivity is weak, while privacy diseases (such as cancer), which may be diseases that only a few people have, or patients do not want others to know that they have the disease, have strong sensitivity. If all data elements are homogenized, the accuracy and effectiveness of the data may be weakened. Summary of the Invention

[0004] One or more embodiments of this specification describe a service processing method and apparatus based on privacy protection to solve one or more problems mentioned in the background art.

[0005] According to a first aspect, a service processing method based on privacy protection is provided for performing service processing related to target information, where the target information corresponds to a plurality of attribute items; the method includes: each client provides respective first perturbed data about the target information to the server, and a single first perturbed data is determined by a single client based on first differential privacy satisfied by the attribute items of the target information locally; the server detects sensitive items and non-sensitive items in each attribute item of the target information according to the respective first perturbed data sent by each client, and feeds back the detection result to each client;

[0006] Each client, based on the detection result, samples according to whether the local attribute item of the target information is a sensitive item to obtain corresponding second perturbed data that satisfies differential privacy of order two, and uploads it to the server; the server performs corresponding service processing on the target information based on the second perturbed data uploaded by each client respectively.

[0007] In one embodiment, the multiple attribute items are identified by consecutive non - negative integers as the identifier of each attribute item, and the differential privacy of order one satisfied by the attribute items on a single client side is determined by perturbing the identifier of the local attribute item with a predetermined privacy budget.

[0008] In one embodiment, the differential privacy of order one adopts the Hadamard response mechanism. The local attribute item held by a single client is the first attribute item, and the first attribute item corresponds to the first identifier. The single client determines the first perturbed data corresponding to the first attribute item in the following way: according to the element values in the row corresponding to the first identifier in the pre - constructed first Hadamard matrix, determine the first sampling probability distribution of the first perturbed data corresponding to the first attribute item, where the first sampling probability distribution describes the probability that the first identifier is sampled as each column identifier of the first Hadamard matrix; sample according to the first sampling probability distribution among the column identifiers corresponding to the first Hadamard matrix to obtain the sampled column identifier as the corresponding first perturbed data.

[0009] In one embodiment, the number of attribute items of the target information is k, the order of the first Hadamard matrix is the smallest power of 2 greater than k, and the first identifier x corresponds to the (x + 2)-th row in the first Hadamard matrix.

[0010] In one embodiment, among the rows corresponding to the first identifier in the first Hadamard matrix, the column identifiers corresponding to the elements with a value of +1 form the first candidate set, and the column identifiers corresponding to the elements with a value of -1 form the second candidate set; the first sampling probability distribution includes: the probability that a value in the first candidate set is sampled is the first probability, and the probability that a value in the second candidate set is sampled is the second probability, where the first probability is greater than the second probability, and both the first probability and the second probability are determined based on a predetermined privacy budget.

[0011] In one embodiment, the server detects sensitive items and non - sensitive items in each attribute item of the target information according to the first perturbed data sent by each client respectively, including: using the first quantity of the first perturbed data sampled from the first candidate set to determine the first occurrence frequency of the first identifier in each first perturbed data; determining the preliminary distribution probability of the first identifier based on the first occurrence frequency; comparing the preliminary distribution probability with a preset probability threshold to identify the first attribute item as a sensitive item or a non - sensitive item.

[0012] In one embodiment, the second differential privacy satisfies that: for a sensitive item, it corresponds to a first privacy budget, and for a non-sensitive item, it corresponds to a second privacy budget, and the second privacy budget is greater than the first privacy budget.

[0013] In one embodiment, the second privacy budget tends to infinity.

[0014] In one embodiment, for a single client, the attribute item corresponding to the target information is a second attribute item, and the second attribute item corresponds to a second item identifier; for the second attribute item, the single client determines the corresponding second perturbed data in the following manner: matching the second item identifier with each item identifier of the sensitive items and non-sensitive items in the detection result; sampling the second perturbed data corresponding to the second attribute item based on the matching result.

[0015] In one embodiment, the matching result is that the second attribute item belongs to the sensitive items, and the single client samples the corresponding second perturbed result according to a second probability distribution, and the second probability distribution includes: for each column identifier in the first column identifier set, the sampling probability is a third probability; for each column identifier in the second column identifier set, the sampling probability is a fourth probability; where: the third probability is the first multiple of the fourth probability, and the first multiple is determined with the natural logarithm e as the base and the first privacy budget as the exponent; the first column identifier set includes each column identifier with the element value of +1 in the row corresponding to the second attribute item in the second Hadamard matrix, and the second column identifier set includes each column identifier with the value of -1 respectively in the row corresponding to the second attribute item in the second Hadamard matrix; the order S of the second Hadamard matrix is the smallest integer power of 2 greater than the number s of sensitive item items.

[0016] In one embodiment, the matching result is that the second attribute item belongs to the non-sensitive items, and the single client samples the corresponding second perturbed result according to a third probability distribution, and the third probability distribution includes: the fifth probability of sampling as each column identifier of the second Hadamard matrix is negatively correlated with the second privacy budget; the sixth probability of sampling as t consecutive data greater than the largest column identifier of the second Hadamard matrix is positively correlated with the second privacy budget; where: t is the number of non-sensitive item items; the order S of the second Hadamard matrix is the smallest integer power of 2 greater than the number s of sensitive item items.

[0017] In one embodiment, the service processing includes determining the frequency of each attribute item, and the service side performs corresponding service processing on the target information based on the second perturbed data respectively uploaded by each client, including: for a single sensitive item, based on the second perturbed data, the overall frequency of the sensitive item All attribute items are collected as the local frequency of the column identifiers with the element of +1 in the corresponding row of the second Hadamard matrix for this single sensitive item Determine the frequency of the corresponding attribute item; for a single non-sensitive item, determine the frequency of the corresponding attribute item based on the number of the corresponding sampling values in each of the second perturbation data, where the corresponding sampling values are within t consecutive data ranges greater than the maximum column identifier of the second Hadamard matrix.

[0018] According to a second aspect, there is provided a privacy protection-based service processing method for performing service processing related to target information, where the target information corresponds to multiple attribute items; the method is executed by a client and includes:

[0019] Provide the server with first perturbation data of the attribute items of the target information locally, where the first perturbation data is determined based on the first differential privacy satisfied by the attribute items of the target information locally;

[0020] Receive the detection results of the sensitive and non-sensitive items in each attribute item of the target information fed back by the server, where the detection results are determined based on the first perturbation data provided by each client;

[0021] Based on the detection results, perform sampling that satisfies the second differential privacy according to whether the local attribute items of the target information are sensitive items to obtain second perturbation data, and upload the second perturbation data to the server for the server to perform corresponding service processing on the target information based on the second perturbation data uploaded by each client respectively.

[0022] According to a third aspect, there is provided a privacy protection-based service processing method for performing service processing related to target information, where the target information corresponds to multiple attribute items; the method is executed by the server and includes:

[0023] Receive the first perturbation data of the attribute items of the target information locally provided by each client to the server respectively, where a single first perturbation data is determined by a single client based on the first differential privacy satisfied by the local attribute items;

[0024] Detect the sensitive and non-sensitive items in each attribute item of the target information according to the first perturbation data sent by each client respectively, and feedback the detection results to each client for each client to upload the second perturbation data obtained by performing sampling that satisfies the second differential privacy on the local attribute items according to whether they are sensitive items based on the detection results respectively;

[0025] Perform corresponding service processing on the target information based on the second perturbation data uploaded by each client respectively.

[0026] According to a fourth aspect, there is provided a privacy protection-based service processing device for performing service processing related to target information, where the target information corresponds to multiple attribute items; the device is disposed on the client and includes:

[0027] A first perturbation unit, configured to provide first perturbation data about the local attribute items of the target information to the server, where the first perturbation data is determined based on the first differential privacy satisfied by the local attribute items of the target information;

[0028] A receiving unit, configured to receive the detection results of the sensitive items and non-sensitive items in each attribute item of the target information fed back by the server, where the detection results are determined based on the first perturbation data provided by each client;

[0029] A second perturbation unit, configured to, based on the detection results, sample to obtain second perturbation data that satisfies the second differential privacy for whether the local attribute items of the target information are sensitive items, and upload the second perturbation data to the server for the server to perform corresponding service processing on the target information based on the second perturbation data respectively uploaded by each client.

[0030] According to a fifth aspect, a service processing device based on privacy protection is provided for performing service processing related to target information, where the target information corresponds to multiple attribute items; the device is disposed on the server and includes:

[0031] A receiving unit, configured to receive the first perturbation data about the local attribute items of the target information provided by each client to the server respectively, and a single first perturbation data is determined by a single client based on the first differential privacy satisfied by the corresponding local attribute items;

[0032] A detection unit, configured to detect the sensitive items and non-sensitive items in each attribute item of the target information according to the first perturbation data respectively sent by each client, and feedback the detection results to each client for each client to respectively upload the second perturbation data obtained by sampling according to whether the local attribute items are sensitive items and satisfying the second differential privacy;

[0033] A processing unit, configured to perform corresponding service processing on the target information based on the second perturbation data respectively uploaded by each client.

[0034] According to a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the second aspect or the third aspect.

[0035] According to a sixth aspect, a computing device is provided, including a memory and a processor, characterized in that an executable code is stored in the memory, and when the processor executes the executable code, the method of the second aspect or the third aspect is implemented.

[0036] Through the method and device provided by the embodiments of this specification, a business processing process including two stages is provided. In the first stage, in the manner of local differential privacy, each client uploads the first perturbed data of the local privacy data of the target information to the server, and the server distinguishes the sensitive items and non-sensitive items among the possible attribute items of the target information according to each first perturbed data uploaded by each client. Then, the server feeds back the distinction results of the sensitive items and non-sensitive items to each client. In the second stage, each client differentially performs local differential privacy according to the local privacy data corresponding to the target information, giving a higher privacy budget when the local data is a non-sensitive item and a lower privacy budget when the local data is a sensitive item, so as to obtain the second perturbed data of the local privacy data and upload it to the server. The server performs corresponding business processing based on each second perturbed data. Through the two-stage setting, in the first stage, the server automatically distinguishes sensitive items and non-sensitive items according to the perturbed data, avoiding the interference of human factors in the process of dividing sensitive items and non-sensitive items. In the second stage, the data accuracy of non-sensitive items is improved, so that the overall accuracy of business processing is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0038] Figure 1 A schematic diagram of a specific implementation architecture showing the technical concept of this specification;

[0039] Figure 2 A flowchart of a privacy protection-based business processing method under the interaction between a client and a server showing a specific example;

[0040] Figure 3 A schematic diagram showing the value range of the second perturbed data under the HL-LDP architecture according to an embodiment;

[0041] Figure 4 A flowchart of a privacy protection-based business processing method executed by a client according to an embodiment;

[0042] Figure 5 A flowchart of a privacy protection-based business processing method executed by a server according to an embodiment;

[0043] Figure 6 A schematic block diagram of a privacy protection-based business processing device provided at a client according to an embodiment;

[0044] Figure 7 A schematic block diagram of a privacy - protected service - side business processing device according to an embodiment is shown. Detailed implementation manners

[0045] The technical solutions provided in this specification will be described below with reference to the accompanying drawings.

[0046] With the development of big data, there are more and more scenarios for comprehensive processing of data from multiple data parties. For example, federated learning scenarios, multi - party secure computing scenarios, joint frequency estimation scenarios, and so on. With the continuous tightening of industry supervision and the continuous strengthening of the public's privacy awareness, privacy computing has become an important branch in the field of data analysis. For example, a frequency estimation scheme for privacy protection of reasonable user data that complies with laws and regulations supervision. For example, developers of a certain operating system or the service side of an instant messaging application use frequency estimation to calculate the usage frequencies of different emojis, developers of a certain operating system use frequency estimation to calculate the usage data of its own browser, the service side of a certain input method counts the word frequencies used by users, and developers of an operating system count the user volumes of various applications, and so on. These frequency estimations may all involve user privacy.

[0047] To meet the legal supervision requirements or the privacy protection requirements of users for their own data, the LDP (Local Differential Privacy) protocol is usually used to add noise to the local data of the client and then upload it when collecting data. Conventional LDP protocols are usually implemented through a local model of privacy. LDP can be regarded as a noisy channel, with the input being the original data and the output being the local differential privacy transformation (LDP view) of the data. Based on a certain LDP view, an attacker cannot determine whether the original data corresponding to this view is one of any two possible original data points.

[0048] Conventional LDP protocols usually homogenize all data elements. This may result in a significant loss of the accuracy of the processing results, making meaningful results rare or even completely unusable. For example, a medical platform wants to conduct a case survey. During the homogenized privacy processing, the privacy budgets and the ways of adding noise for "common cold" and "AIDS" are the same. Generally, "common cold" is a disease that almost every respondent has experienced, while only a small number of respondents have experienced "AIDS", and the leakage of this information will pose a serious potential privacy infringement to the respondents. Therefore, a larger privacy budget is needed to protect the privacy of "AIDS" patients. If the sensitivity levels of "common cold" and "AIDS" are regarded as the same during the privacy process, then the privacy of "common cold" will also be increased. In this way, the deviation of the results may also increase, which may cause excessive noise addition to the data, resulting in unreasonable survey results and reduced usability.

[0049] Recently, some papers (such as the paper Context-Aware Local Differential Privacy published by Jayadev Acharya, K.A. Bonawitz, Peter Kairouz, Daniel Ramage, Ziteng Sun, etc.) have proposed a differential privacy protection scheme. The concept of Context-aware LDP in it describes that the definition of local differential privacy can be related to the specific data elements to be protected. In particular, it proposes the concept of High-low LDP (HLLDP): allowing a part of non-sensitive data elements to be (almost) free from the noise interference of differential privacy, so as to obtain relatively accurate results, while for another part of sensitive data elements, corresponding privacy algorithms are used to protect them.

[0050] Although the HLLDP protocol provides a differential frequency estimation framework, the basic idea of this framework is to provide a set of sensitive elements A and a set of non-sensitive elements B. For the data in the set of sensitive elements A and the set of non-sensitive elements B, different privacy budgets are given respectively, and noise data is sampled through different probability distributions, so as to add different noises to sensitive data and non-sensitive data. However, this scheme needs to know which data elements are sensitive elements, which may not be available for specific problems, and thus requires manual determination and input of a division of non-sensitive data and sensitive data.

[0051] In view of this, this specification aims to provide a two-stage data desensitization scheme, so that in the first stage, the division results of sensitive / non-sensitive elements can be obtained in a privacy-protected and data-driven manner, and in the second stage, the data desensitization process of HLLDP privacy protection is implemented according to the division results of sensitive / non-sensitive elements.

[0052] Figure 1 Shows a schematic diagram of an implementation architecture of the technical concept of this specification. Figure 1 Shows a frequency estimation scenario, which is, for example, a survey scenario of the prevalence of diseases in a certain hospital. In this scenario, the current user as a patient can report disease data to the surveyor (such as the server of the hospital) through an online terminal. In the first stage, the traditional LDP method can be used to report LDP data under differential privacy. For example, a single user can upload the noisy data of the disease they have through the corresponding client. For example, if the current disease is a cold, then upload the perturbed data after perturbing the information based on the cold to meet differential privacy. Then, the server can detect the sensitivity of various diseases, that is, whether it is sensitive data, based on the perturbed data obtained from multiple clients. It can be understood that for disease data, the more frequently occurring diseases (such as colds) among multiple users may be more popular and the lower the sensitivity felt by users, and the less frequently occurring diseases (such as AIDS) may be less popular and the higher the sensitivity felt by users. Accordingly, the server can estimate the probability distribution of each attribute item, and compare the probability value according to the probability distribution with a predetermined threshold ( Figure 1 which is 0 in the example) to determine sensitive items and non-sensitive items as the sensitivity detection result. Provide the sensitivity detection result to each client. In the case where there are multiple pieces of privacy information, the server can also provide a set of sensitive data and non-sensitive data to each client.

[0053] Then, in the second stage, a single client performs HL-LDP (High-low LDP), that is, differential differential privacy, on the local privacy data according to the detection result of the data sensitivity received. Specifically, for sensitive data, add noise with a smaller privacy budget, and for non-sensitive data, add noise with a larger privacy budget. When the privacy budget is infinite, the added noise has almost no effect on the original data. Then, each client can provide the noisy data (such as denoted as HL-LDP data) obtained by local differential differential privacy to the server. The server can re-determine the frequency estimation result based on the HL-LDP data uploaded by each client. Among them, Figure 1 the frequency estimation method shown in is the Hadamard Response frequency estimation method.

[0054] In the above process, by using the LDP and HL-LDP methods in two stages respectively, the server automatically distinguishes sensitive data and non-sensitive data according to the LDP data and distributes them to the data providers (such as Figure 1The client in it), so that the data provider can perform differential local differential privacy based on the data sensitivity result, and add noise to the local private data. Since the data sensitivity is driven by the data itself and is more objective, it effectively improves the accuracy of the business result and the rationality of the local differential privacy.

[0055] Figure 2 FIG. shows a privacy protection-based service processing flow according to an embodiment of the present specification. This flow is used when the client provides data to the server. After desensitizing the local data, it is provided to the service provider for the service provider to use the desensitized data for service processing to obtain corresponding service processing results. For example, perform the frequency estimation service processing in the foregoing text to obtain a frequency estimation result. It can be understood that the data interaction between the client and the server is involved in this process. Therefore, Figure 2 In the following, the interaction between a single client and the server is taken as an example to describe the privacy protection-based service processing flow. In practice, the number of clients can be determined according to the actual situation, and each client interacts with the server in a similar manner. It should be noted that the following description of Figure 2 The embodiment shown is described by taking the frequency estimation scenario of diseases as an example. The implementation architecture provided in this specification can also be applied to other service scenarios, such as other frequency estimation scenarios (such as user consumption item statistics scenarios), centered federated learning scenarios, etc. In scenarios such as centered federated learning, each training member can be regarded as a client, and a trusted third party can be regarded as a server.

[0056] Such as Figure 2 As shown, the privacy protection-based service processing flow of an embodiment of the present specification may include the following steps:

[0057] Step 201, the client provides the server with first perturbation data that satisfies the first differential privacy for the local private data of the target information.

[0058] Here, the target information can be any information required for the service processing of the server. For example, in the frequency estimation service scenario of disease frequency, the first information can be the type of disease targeted by the user's current visit (such as cold, pneumonia, keratitis, etc.). A single client may have local private data corresponding to the target information. For example, the current client corresponds to the first user, and the first user corresponds to the disease targeted by the current diagnosis and treatment, such as a cold. Then, the disease of cold can be used as the local private data of the current client for the target information.

[0059] To protect local data privacy, the single client can perturb the privacy data while satisfying the differential privacy mechanism to form a single first perturbed data. Here, the first perturbed data can be understood as the perturbed data obtained by performing the first perturbation on the local data in the first stage before determining whether the information is sensitive data. That is to say, the "first" in "first perturbed data" corresponds to the "first stage", without other substantial limitations on the perturbed data. Before specifically describing the detailed process of adding perturbations below, the basic principle of differential privacy is first introduced briefly.

[0060] Differential Privacy (DP) is a means in cryptography, aiming to provide a way to maximize the accuracy of data queries while minimizing the chance of identifying its records when querying from a statistical database. Let there be a random algorithm M, and PM be the set of all possible outputs of M. For any two neighboring datasets x and x' (i.e., x and x' differ by only one data record) and any subset of PM If the random algorithm M satisfies:

[0061]

[0062] Then the algorithm M is said to provide ε-differential privacy protection, where the parameter ε is called the privacy protection budget, used to balance the degree of privacy protection and accuracy. ε can usually be preset in advance. The closer ε is to 0, e ε The closer it is to 1, the closer the processing results of the random algorithm for the two neighboring datasets x and x' are, and the stronger the privacy protection degree.

[0063] The implementation methods of differential privacy include the noise mechanism, the exponential mechanism, etc. In the case of the noise mechanism, generally, the amplitude of the added noise is determined according to the sensitivity of the query function. The above-mentioned sensitivity means the maximum difference in the query results of the query function when querying a pair of adjacent datasets x and x'. The noise mechanism includes, for example, the Gaussian noise mechanism, the Laplace noise mechanism, etc. Taking the Laplace mechanism as an example, given a dataset D, let there be a function f that maps the data in the dataset D to the dataset Rd, and its sensitivity is Δf. Then the random algorithm M(D) = f(D) + Y provides ε-differential privacy protection, where Y is a random noise that satisfies Lap(Δf / ε), that is, the noise follows the Laplace differential privacy with a scale parameter of Δf / ε. In the case where the data is frequency, the sensitivity is, for example, 1. For a given privacy budget ε and sensitivity, the corresponding noise can be sampled according to the Laplace distribution and added to the local privacy data.

[0064] Specifically in the scenario of disease frequency estimation, in order to perturb the disease types, various diseases can be identified using numerical values. For example, if there are k possible diseases, then various diseases can be identified by k consecutive integers. For instance, various diseases can be identified by numerical values 0, 1, 2,..., k - 1 respectively, with one numerical value corresponding to one optional element (disease). Thus, the local privacy data can include the numerical values corresponding to the respective diseases. For example, when the current disease is a cold, the corresponding local privacy data is 2, and so on. Since in the two-stage architecture of this specification, different differential privacy mechanisms need to be adopted in the two stages, the differential privacy in the first stage can be referred to as the first differential privacy. Under the first differential privacy, each client can adopt the same privacy budget ε.

[0065] In one embodiment, for the specific numerical value corresponding to the current privacy data, the current client can generate a random noise that satisfies Lap(Δf / ε) based on the predetermined privacy budget ε, and add this random noise to the corresponding numerical value to obtain the first perturbed data.

[0066] In the case where each attribute item of the target information is identified by consecutive numerical values starting from 0, according to a possible embodiment, this single client can also sample the corresponding noise data for the local privacy data based on the Hadamard matrix under the privacy budget ε according to the Hadamard Response (HR) mechanism. The method for generating a single first perturbed data under the HR mechanism is introduced below.

[0067] The Hadamard matrix is a square matrix, each element of which is +1 or -1, and each row is orthogonal to each other. According to this property, the Hadamard matrix can be constructed in a recursive manner of order 2 m The order of the Hadamard matrix is usually 1, 2, or a multiple of 4. Or rather, for non-negative integer m, any Hadamard matrix of order 2 m can be constructed. The Hadamard matrix is not unique. A relatively classic one is the Hadamard matrix given by Sylvester, which is more convenient to apply due to the following special properties: they are all symmetric matrices; the trace of the matrix is 0; the elements of the first row and the first column are all +1, and the elements of the other rows and columns are half +1 and half -1.

[0068] Under the HR mechanism, assuming that the constructed Hadamard matrix is the first Hadamard matrix, for an element x of the target information (such as when the current disease is a cold, x = 2), the corresponding sampling probability can be determined according to the following first sampling probability distribution:

[0069]

[0070] Among them, Q(z|x) represents the probability that the attribute item identifier x of the private data is sampled as the value z, z represents each column number in the first Hadamard matrix (for example, the serial number of the 3rd column is 3), C x represents the set of column numbers with the value +1 in the row elements corresponding to x in the Hadamard matrix, denoted as the first candidate set, Z represents the set of all element column numbers in the first Hadamard matrix, such as the column number set of the K-order Hadamard matrix includes K numerical values 1, 2, …… K, etc., Z\C x represents the set of other column numbers except C x in the row elements corresponding to x, that is, the set of column numbers with the element value of -1, denoted as the second candidate set, s represents the number of elements with the value +1 in the row elements corresponding to x in the first Hadamard matrix.

[0071] The order K of the first Hadamard matrix can be determined in advance according to the number of attribute items of the target information, so as to construct the first Hadamard matrix of the corresponding order. The first Hadamard matrix is, for example, the Hadamard matrix of the Sylvester mechanism. According to the constructed first Hadamard matrix, the number s of elements with the value +1 in each row and the positions of the elements with the value +1 in the row corresponding to any attribute item x can be determined. For a single client, the local private data corresponding to the target information is determined. For example, if it is a cold (such as x = 2), then the corresponding first perturbation data can be generated according to the row data in the Hadamard matrix corresponding to the cold (such as 2). Specifically, the first perturbation data corresponding to the local private data can be sampled according to the first probability distribution Q(z|x) in formula (1).

[0072] Specifically, when constructing the first Hadamard matrix, since the order of the Hadamard matrix is 2 to the power of m, and at the same time, the order of the first Hadamard matrix needs to be greater than the number of attribute items of the target information, so as to ensure that each attribute item corresponds to a different row in the first Hadamard matrix. Therefore, for the k possible values of the target information, the order This formula represents the smallest integer power of 2 greater than k, where K≥k + 1. Taking k = 3 as an example, k + 1 = 4, and the ceiling result of log24 is 2, then K = 2 2 = 4. In this way, in the K-order Hadamard matrix, the number of rows and the number of elements in a single row are both greater than the number of k attribute items. A 4-order Hadamard matrix is, for example:

[0073]

[0074] Thus, for any attribute item of the target information, it corresponds to an attribute item identifier x between 0 and k - 1, and can correspond to the (x + 2)-th row of the first Hadamard matrix. This is because under the Sylvester rule, the elements of the first row are all 1, which is of little significance for the sampling positions. The maximum value x = k - 1 corresponds to the (k + 1)-th row. Since the matrix order K ≥ k + 1, therefore, the numerical values of the attribute item identifiers in the target information can all find corresponding rows in the first Hadamard matrix. It can be understood that in the case where the first Hadamard matrix is constructed by other rules, the correspondence between the attribute item identifiers of the target information and the matrix rows can be determined by other means. The goal is to determine the sampling probability distribution rule according to the element values of +1 or -1, so as to distinguish the sampling values. For example, if the attribute items of the target information are identified from 1 to k, then the attribute item identifier x can correspond to the (x + 1)-th row. If the attribute items of the target information are identified from 2 to k + 1, then the attribute item identifier x can correspond to the x-th row, and so on. Starting from the second row of the Hadamard matrix under the Sylvester rule, half of the elements in each row are +1 and half are -1. Thus, the number s of elements with the value of +1 in the row corresponding to x in the Hadamard matrix is s = K / 2.

[0075] Further, for a specific local privacy data x of the target information (any value between 0 and k - 1, such as x = 2), sampling can be performed according to the sampling probability defined by Q(z|x). Taking the 4th-order Hadamard matrix H4 shown above as an example, for instance, if the current client's local privacy data is a cold, correspondingly x = 2, which corresponds to the 4th row in the Hadamard matrix. The element values at the 1st and 4th positions are +1, and the element values at the 2nd and 3rd positions are -1. Then C x can represent the column identifier set {1, 4}, corresponding to the first candidate set. Additionally, the second candidate set is other column identifier sets, such as {2, 3}. Sample the first perturbed data x' according to the sampling probability Q(z|x) described above. The probability that x = 2 is sampled as the column identifiers x' = 1, 4 in the first candidate set is both the first probability greater than the probability that it is sampled as the column identifiers x' = 2, 3 in the second candidate set, which is both the second probability For example, if the sampling result is 4, then 4 is taken as the first perturbed data after adding noise to the local privacy data of the target information (such as the disease suffered).

[0076] The HR method has a relatively low computational complexity and transmission volume (i.e., communication volume), and is an efficient LDP method. It can be understood that the Hadamard matrices used by multiple clients with the HP architecture can be consistent, and the corresponding Cs x can also be consistent. Therefore, when the possible values of the target information are determined, parameters such as the Hadamard matrix, C x K, Z, s, etc. can also be determined in advance.

[0077] In practice, a single client can also use other methods to perturb local privacy data to obtain corresponding perturbed data, which is not limited here. Similarly, each client can use a similar method to obtain corresponding first perturbed data for the local privacy data of the target information and provide it to the server. It can be understood that the first perturbed data can be regarded as the item identifier after perturbing the item identifier of the local attribute item, pointing to other attribute items. For example, for the cold attribute item x = 2, the first perturbed data obtained is 3, corresponding to the attribute item with the item identifier 3 such as rhinitis. For the cold attribute item x = 2, the first perturbed data obtained is 4, then it corresponds to the attribute item with the item identifier 4 such as pneumonia, and so on. Thus, the true attribute item of the corresponding client cannot be inferred from the first perturbed data.

[0078] Step 202, the server detects sensitive items and non-sensitive items in the target information according to the first perturbed data sent by multiple clients, and feeds back the detection results to each client.

[0079] In the frequency estimation scenario, the server can receive the first perturbed data about the target information from multiple clients and perform a preliminary evaluation on these first perturbed data to obtain the preliminary frequency estimation results of each possible attribute item of the target information, which are used to detect sensitive items and non-sensitive items in the target information. Taking the business scenario of disease frequency statistics as an example, the preliminary frequency estimation results of the target information can include, for example, the prevalence probabilities of various diseases, etc.

[0080] In an optional embodiment, the server can count the attribute items corresponding to each first perturbed data according to the quantity, and estimate the frequency of occurrence of each attribute item based on the total number of clients as the preliminary frequency estimation result. Assume that the total number of clients is n, and Z i is the first perturbed data of the i-th client for the first information, and the n first perturbed data provided by n clients for the target information are Z1, Z2,... Z n . Then for any value x between 0 and k - 1, corresponding to an attribute item of the target information, its occurrence frequency can be the ratio m / n of the number of clients m with Z i = x to the total number of clients n. For k values between 0 and k - 1, k occurrence frequencies can be obtained. Further, the client can compare these occurrence frequencies with a predetermined frequency threshold (such as 0.3). In the frequency estimation scenario of disease prevalence, the attribute items (disease items) with occurrence frequencies greater than the frequency threshold can be determined as non-sensitive items, and the attribute items with occurrence frequencies less than the frequency threshold can be determined as sensitive items.

[0081] In another optional embodiment, under the HP architecture, taking into account the sampling error, the server can determine the input distribution of the target information on each attribute item according to the value of each first disturbance data, and determine the sensitive items and non-sensitive items of the target information based on the input distribution.

[0082] Specifically, when the number of clients is n, the first disturbance data provided by each client for the target information are Z1, Z2, ..., Z n In the case of , according to the LDP principle under the HR architecture mentioned above, taking the n data as an example, all of which are values between 0 and k-1, for any item identifier (hereinafter referred to as the first item identifier) x, corresponding to an attribute item of the target information (such as the cold attribute item), its frequency of occurrence in the n first disturbance data (which can be regarded as an empirical probability estimate) can be estimated in the following way:

[0083]

[0084] in, It means that for n clients as a whole, the first disturbance data Z provided by the jth client is j Is it C x The elements in are represented by numerical values. Usually, Z j C x In the case of elements in Can be a non-zero value (such as 1), Z j Not for C x In the case of elements in Can be zero value (ie 0).

[0085] Furthermore, the initial probability distribution of the first identifier x is:

[0086]

[0087] Based on this probability distribution The probability of the attribute item corresponding to the first identifier x can be determined. In the frequency statistics scenario of the probability of disease, for the attribute item of cold, assuming that x is 2, it can be considered that in the current preliminary frequency estimation result, the frequency corresponding to the attribute item of cold is

[0088] Furthermore, k frequencies can be obtained for k attribute items. By comparing the k frequencies with the preset probability thresholds one by one, it is possible to detect whether they are sensitive items. For example, attribute items (disease items) whose corresponding probabilities are greater than the probability threshold are determined as non-sensitive items, and attribute items whose corresponding probabilities are less than the probability threshold are determined as sensitive items. At this time, a probability threshold can be pre-set, such as 0 or 0.1, etc. When the probability is less than the probability threshold, it is determined that the attribute item corresponding to the item identifier x (such as the AIDS attribute item) is a sensitive item. In When the probability is greater than the probability threshold, it is determined that the attribute item corresponding to the item identifier x (such as the cold attribute item) is a non-sensitive item.

[0089] In the service states obtained based on the LDP process above, various information is considered homogenously, or processed based on a unified privacy budget, without considering the differences between information. Especially in the HP-LDP process, in order to improve efficiency and save computational and communication amounts, the damage to data accuracy is relatively large. To distinguish the differences between information, the detection results of sensitive and non-sensitive items in the target information can be provided to the client, so that the client can re-perform differential LDP operations based on local data to improve the accuracy of the data sent.

[0090] In one embodiment, the sensitive and non-sensitive item information in the target information can be marked by sensitivity identifiers. For example, the first value (such as 1) is used to represent sensitive data, and the second value (such as 0) is used to represent non-sensitive data. For multiple attribute items of the target information, the sensitivity detection results of each attribute item can be described in the form of an array, vector, set, etc. For example, the vector combination M=(1, 0, 0) is used to represent that the first attribute item of the target information is a sensitive item, and the second and third attribute items are non-sensitive items.

[0091] In another embodiment, the server can also describe the sensitivity of each attribute item of the target information through set classification. For example, sets A and B are used to record sensitive and non-sensitive items respectively. Still taking the previous example as an example, assuming that the respective attribute items corresponding to the target information are identified by the numerical values in the following data set: {k}=0, 1... k-1, the server can determine two predetermined sets A and B, where set A records sensitive items and set B records non-sensitive items. In an optional implementation, x i is an element in the data set {k}, and set A records the sensitive items of the target information, such as A={x1, x2,... x s}, then set B can be used to record the non-sensitive items of the target information. For example, set B is denoted as {k}-A={x s+1 , x s+2 ,... x s+t}, indicating that the elements therein are the remaining elements of the data set {k} after removing the elements in set A, and t=k-s.

[0092] In other embodiments, the sensitive and non-sensitive items in the target information can also be described in other ways, which will not be elaborated here one by one. The server can provide the information of the sensitive and non-sensitive items in the target information to each client as the detection result. For example, providing set A and set B to each client, or providing vector M, etc.

[0093] Step 203: Based on the detection result, the client samples the local privacy data of the target information to obtain corresponding second perturbed data that satisfies differential privacy of the second level, and provides the desensitized data to the server as the second perturbed data.

[0094] Here, the second perturbed data can be understood as the perturbed data obtained by perturbing the local data a second time in the second stage after determining whether the relevant information is sensitive data. The "second" in the "second perturbed data" corresponds to the "second stage", and each client can obtain the second perturbed data of its corresponding local privacy data. The client can re-perturb the local privacy data according to the differential privacy budget based on the information in the detection result as to whether the attribute item corresponding to the local privacy data is a sensitive item, so as to obtain the corresponding second perturbed data.

[0095] Among them, the client can obtain the information as to whether the local privacy attribute item is a sensitive item from the detection result provided by the server. Depending on the form of the detection result, for example, the client can determine whether the local privacy attribute item is a sensitive item by the value of the element corresponding to the local attribute item (such as the cold item) in the vector describing the detection result, or can match the identifier corresponding to the local privacy attribute item with the elements in the set of sensitive items or non-sensitive items describing the detection result, so as to determine whether the local privacy attribute item is a sensitive item or a non-sensitive item.

[0096] A single client can determine different privacy budgets for sensitive items and non-sensitive items, and select the corresponding privacy budget to generate the second perturbed data according to the sensitive attributes of the local privacy data. For example, the sensitive items satisfy the first privacy budget, and the non-sensitive items satisfy the second privacy budget, and the second privacy budget is much larger than the first privacy budget. Specifically, in one embodiment, for non-sensitive items, the second privacy budget can tend to infinity, and for sensitive items, the first privacy budget is a predetermined ε (such as 0.1, 0.2, etc.). If the second perturbed data corresponding to the attribute item identifier x is denoted as x', and the detection result is represented by the set of sensitive items A and non-sensitive items B, the privacy budget is expressed, for example, as:

[0097]

[0098] Furthermore, a single client can generate second perturbed data that satisfies differential privacy for the local privacy data according to the privacy budget determined for the local privacy data.

[0099] In the implementation architecture of this specification, a variant of HR, High-low LDP (i.e., HL-LDP), can be adopted to generate the second perturbed data with a differential privacy protection strategy. HL-LDP allows non-sensitive items to use an infinite privacy budget, that is, there is no need to protect privacy. The privacy budget ε corresponding to the sensitive items in the above formula can be called the privacy budget of the HL-LDP protocol. The privacy budget ε can be preset as needed, such as 0.1, 0.2, etc.

[0100] Under the HL-LDP protocol, according to the relevant papers, the client can determine the second perturbed data in the following way:

[0101] Assume that the number of attribute items of the target information is k (such as the number of disease types to be frequency-estimated is k), and let s represent the number of sensitive items, then s = |A|. For example, the sensitive item set is denoted as A = {x1, x2, …… x s}, and S is the smallest power of 2 greater than s, that is As the order of the second Hadamard matrix for determining the construction of the second perturbed data. The second Hadamard matrix H s For example, it can be a Hadamard matrix of order S with a Sylvester structure. In addition, let the number of non-sensitive items be t = k - s. Then: S + t ≤ 2(s + t) = 2k. The non-sensitive item set is denoted as B = {x s+1 , x s+2 , …… x s+t}. Among them, x s+t is x k . Then, let the output range of the private data be the set {1, 2, …, S + t}, and the relationship between S, t, k, and s can be as Figure 3 shown. Thus:

[0102] When the local privacy attribute item is a sensitive item (such as x ∈ A), the sampling probability that the second perturbed data corresponding to the data x is sampled as the column number y is the second probability distribution:

[0103]

[0104] When the local privacy attribute item is a non-sensitive item (such as ), the sampling probability that the second perturbed data corresponding to the data x is sampled as the column number y is the third probability distribution:

[0105]

[0106] In the formula, [S] represents the set of S column numbers 1, 2, …, S, and “s.t.” means satisfying simultaneously. H S (s, y) represents the element value corresponding to the column number y in the row corresponding to x in the second Hadamard matrix.

[0107] In the second probability distribution of formula (2), among the rows corresponding to x in the second Hadamard matrix, the column identifiers corresponding to the element value +1 form the first column identifier set, and the column identifiers corresponding to the element value -1 form the second column identifier set. Then, since ε is greater than 0, the third probability that the attribute item identifier x is sampled as each column identifier in the first column identifier set is greater than the fourth probability that it is sampled as each column identifier in the second column identifier set. Assume that the third probability is the first multiple of the fourth probability, and the first multiple is determined with the natural logarithm e as the base and the first privacy budget as the exponent. And, when the second Hadamard matrix is a Hadamard matrix of order S with a Sylvester structure, the total probability that the attribute item identifier x is sampled as the first column identifier set and the second column identifier set is 1, that is, the second perturbed data corresponding to the sensitive item must be sampled as the data in the data set {1, 2, …… S}, that is, corresponding to Figure 3 the data in the first interval.

[0108] In the third probability distribution of formula (3), x is the item identifier of a non-sensitive item, and the fifth probability that it is sampled as each column identifier {1, 2 …… S} of the second Hadamard matrix (corresponding to Figure 3 the first interval) is less than the sixth probability that it is sampled as t consecutive data {S + 1, S + 2 …… S + t} greater than the maximum column identifier of the second Hadamard matrix (corresponding to Figure 3 the second interval). Among them, the fifth probability is negatively correlated with the second privacy budget, and the sixth probability is positively correlated with the second privacy budget. When the second privacy budget approaches infinity, the sixth probability that x is sampled as x + S - s approaches 1, and almost no privacy protection is performed.

[0109] Assume that the current client's local privacy attribute item is AIDS, corresponding to x = 10. If 10 ∈ A, then AIDS is a sensitive item. Further, the client can sample the AIDS item identifier 10 through the formula describing the sampling probability distribution of the sensitive item above to obtain the corresponding second perturbed data. According to the second probability distribution, among the set {1, 2,..., S + t}, S column numbers 1, 2 …… S may be sampled as the second perturbed data of the AIDS item identifier 10. Among them, assume that x = 10 corresponds to the (x + 2) = 12th row in the Hadamard matrix of order S, then the probability that the corresponding column identifier is sampled can be determined according to whether the element value of each column in the (x + 2)th row is +1 or -1. Among them, in this row, the probability that the column identifier with the element value +1 is sampled is the third probability the probability that the column number with the element value -1 is sampled is the fourth probability the probability that other values (such as S + 1 to S + t) are sampled is 0.

[0110] Assume that the local privacy attribute item is a cold, corresponding to x = 2, and If it is a non-sensitive item, the second perturbation item of the cold item identifier 2 is sampled through the third probability distribution in the above formula (3). In the set {1, 2,..., S + t}, the probabilities of sampling the S column numbers 1, 2,..., S are all the fifth probability. The probability of sampling y = x + S - s (since x is greater than s and less than k, so x + S - s is less than k + S - s = S + t and is a value in the set) is the sixth probability. The probability of sampling other values is 0. When ε approaches infinity, the probability of sampling y = x + S - s approaches 1, while the probabilities of sampling the S column numbers 1, 2,..., S approach 0. X is most likely to be sampled as x + S - s. The maximum value of x - s is t, which is equivalent to shifting x by S - s values to Figure 3 S+(x - s) in the second interval shown, so as to try to distinguish it from the sampling results of sensitive items.

[0111] In summary, under the privacy budget ε, the sensitive data x is usually sampled as a value between 1 and S, and is most likely to be sampled as the column number corresponding to the element value +1 in the row corresponding to it in the Hadamard matrix; under the privacy budget approaching infinity, the non-sensitive data x is most likely to be sampled as the value x + S - s in the interval from S + 1 to S + t. In this way, each client samples the corresponding second perturbation result according to the local actual data. The second perturbation results sampled by the n clients are respectively denoted as Y1, Y2,..., Y n . In the scenario of estimating the frequency of diseases suffered by patients, Y1, Y2,..., Y n respectively correspond to the perturbation values of the disease identifiers of n patients.

[0112] According to the description above, a part of the accuracy is sacrificed in the HP architecture of the first stage, and the server distinguishes sensitive items and non-sensitive items driven by data. In the HL-LDP architecture of the second stage, each client can process local data through a differential privacy protection scheme based on sensitive items and non-sensitive items according to local actual data, and can improve the accuracy.

[0113] Step 204, the server performs corresponding service processing based on the second perturbation data.

[0114] In the scenario of estimating the frequency of diseases suffered by patients, the service processing of each second perturbation data by the server is, for example, estimating the frequency of various diseases.

[0115] According to an optional implementation manner, the generation method of the second perturbation data is similar to that of the first perturbation data, so the same method as that for the preliminary frequency estimation result can be used to process each second perturbation data. Specifically, in an embodiment, for Y1, Y2,..., Yn Statistically analyze the occurrence frequencies of the various numerical values that appear, obtain the disease prevalence frequencies for various diseases, and use them as the service processing results. In another embodiment, each client determines the second perturbed data in the HP architecture, and the server can use the occurrence probability of the first perturbed data x and the probability distribution to determine the disease prevalence frequencies for various diseases and use them as the service processing results.

[0116] Specifically, according to a possible design, the second perturbed data is sampled and determined based on the HL-LDP architecture under the second probability distribution or the third probability distribution, then the server can consider the sensitive items and non-sensitive items separately. Specifically, the second perturbed data is within the range of 1 to S for the perturbed data of sensitive items, and within the range of S + 1 to S + t for the perturbed data of non-sensitive items.

[0117] s defines the value range [s] from 1 to s, and can correspond to each sensitive item, then the integer i can be used to take values within the value range [s], and the corresponding x i can represent the data in the sensitive item.

[0118] In the case where the data i ∈ [s], x i represents a sensitive item, and the value range of the corresponding second perturbed data is [S]. Let S i represent the set of column identifiers where the element values of the corresponding row of x in the first Hadamard matrix i are +1, that is, the corresponding first column identifier set of x i Due to the relatively small privacy budget and low data accuracy, the server can determine its distribution probability through the following method:

[0119]

[0120] i where the overall frequency of the attribute item x corresponding to i is determined through the following method:

[0121]

[0122] The local frequency of all the numerical values sampled into the dataset S i is determined by the number of numerical values in the dataset S in the second perturbed data, such as: i Here,

[0123]

[0124] represents n second perturbed data Y1, Y2... Y nwhere the numerical probability that the value remains within the range 1, 2,..., S defined by the data set S; describes n second perturbation data Y1, Y2,..., Y n in the data set S i the probability of occurrence of the numerical values in, and the sensitive item x i if the corresponding row of x in the Hadamard matrix is the (i + 1)-th row, the data set S i is the set represented by the columns with the value +1 in the (i + 1)-th row. For example, if x i has a corresponding row {1, -1, -1, 1} in the Hadamard matrix, then S i can be {1, 4}.

[0125] In the case of, for the non-sensitive item data x i , due to having a relatively large privacy budget and relatively high data accuracy, the server can directly determine its probability distribution through the following method:

[0126]

[0127] In this way, combining the sampling principle of the client in step 203, for the non-sensitive item x i , the sampling data is Y m = i + S - s with a probability of Therefore, the probability distribution is close to the ratio of the number of second perturbation data Y m = i + S - s to the number of clients n, and the loss of accuracy is extremely small.

[0128] To make the specific scheme shown in Figure 2 more explicit, the following further describes it with a more specific example. Suppose a scenario where a hospital samples the estimated frequency of current cases. Suppose the number of disease types k = 10, and the 10 diseases are respectively identified by 0, 1, 2, 3,..., 9. Under the HR method, a -order first Hadamard matrix H K = H 16 can be constructed. Taking an arbitrary client user as an example, suppose the disease he / she has is a cold, corresponding to an identifier 7 between 0 and k - 1. Then x = 7, and it can correspond to the 9th row in the first Hadamard matrix H 16 . There are 8 positions in the 9th row with the element value +1 and 8 positions with the element value -1. Suppose the set of positions / column identifier set with the element value +1 is R = C x = C7 = {1, 4, 6, 7, 9, 10, 13, 16}, and the set of positions / column identifier set with the element value +1 is T = {2, 3, 5, 8, 11, 12, 14, 15}, and s is the number of elements in the set R, s = 8. Then the client can select from the set R with a probability with probability in the set T The sum of the sampling probabilities of each element in the sets R and T is 1. Substitute s = 8 and K = 16 for sampling, and the sampling result is used as the first perturbation data of the disease suffered by the corresponding user. If the sampling result is 13, the client sends the disease identifier 13 to the server.

[0129] The server receives the disease identifiers Z1, Z2... Z sent by each client n , and can initially estimate the frequencies of various diseases. To estimate the disease frequencies, it is necessary to determine the probability distribution of the disease identifiers of various diseases in the received data. For a value x between 0 and k - 1 (in this specific example, k - 1 = 9), in each of the first perturbation data, the first perturbation data Z j has an occurrence probability of Furthermore, the probability distribution corresponding to the value x is determined by determined.

[0130] In this way, the server can determine the k initial frequency estimates of the k diseases corresponding to the k identifiers between 0 and k - 1 according to the n first perturbation data uploaded by n clients. Then, the server can distinguish the sensitivities of the k initial frequency estimates according to a pre-set frequency threshold. For example, the diseases corresponding to the initial frequency estimates greater than the frequency threshold are non-sensitive items, and the corresponding identifiers are included in the set B, and the diseases corresponding to the initial frequency estimates less than the frequency threshold are sensitive items, and the corresponding identifiers are included in the set A. X = 0, 1, 2... k - 1, x i is an element in the data set {k}, and the number of sensitive items is s, then it can be recorded as A = {x1, x2,... x s} and B = {k} - A. The server can feedback the sets A and B to n clients.

[0131] A single client receives the sets A and B, and can regard the sets A and B as a dictionary to query which set the local disease identifier such as x i = 7 belongs to. When x i = 7 belongs to the set A, it is a sensitive item, determine the privacy budget of differential privacy as ε, and sample the second perturbation value in the set [S] = {1, 2... S} (such as the integers in the first interval in Figure 3 ) according to the sampling probability distribution of formula (2), where is the order of the second Hadamard matrix. When x i = 7 belongs to the set B, it is a non-sensitive item, it can be determined that the privacy budget tends to infinity, and sample the second perturbation value according to the sampling probability distribution of formula (3), and the sampled value is in the set {S + 1, S + 2... S + t} (such as Figure 3The probability in the second interval of the integers) approaches 1. In this way, the sensitive item is sampled between 1 and S with high probability, and the non-sensitive value is sampled at S + t with high probability. Each client samples the local disease identifier as the second perturbed data and provides it to the server.

[0132] The data received by the server is n second perturbed data. For the sensitive items in the second perturbed data, when the current data is one of i = 1, 2... s, the corresponding disease identifiers are x1, x2... x s is a sensitive item, and the corresponding sampling interval is as in Figure 3 the first interval in. The set S of column numbers where the element value is +1 is sampled in the corresponding row (such as row i + 1) of the S-order Hadamard matrix i The local frequency of is: Disease item x i The empirical probability that may be a sensitive item Furthermore, the disease item x i The overall frequency that may be a sensitive item is If i belongs to the values between s + 1 and s + t = k, then the disease item x i is a non-sensitive item, and its frequency can be based on falling into Figure 3 Each second perturbed value shown in the second interval is determined: In this way, the server can obtain the frequency estimation results of each disease item based on the received second perturbed values and perform subsequent business processing. The subsequent business processing here is, for example, feature extraction, model training, epidemic prevention and control decision-making, etc.

[0133] Figure 2 The business processing process based on privacy protection is described in the way of interaction between a single client and the server. As Figure 4 、 Figure 5 shown, it is the business processing process based on privacy protection executed by a single client and the server in the corresponding embodiment.

[0134] As Figure 4 shown, the business processing process based on privacy protection executed by a single client includes:

[0135] Step 401, providing the server with the first perturbed data of the attribute items of the target information locally, where the first perturbed data is determined based on the first differential privacy satisfied by the attribute items of the target information locally;

[0136] Step 402, receiving the detection results of the sensitive items and non-sensitive items in each attribute item of the target information fed back by the server, where the detection results are determined based on the first perturbed data provided by each client respectively;

[0137] Step 403: Based on the detection result, sample to obtain second perturbed data that satisfies differential privacy of order two according to whether the local attribute items of the target information are sensitive items, and upload the data to the server for the server to perform corresponding business processing on the target information based on the second perturbed data uploaded by each client respectively.

[0138] As Figure 5 shown, the privacy protection-based business processing process executed by a single client includes:

[0139] Step 501: Receive the first perturbed data of each local attribute item of the target information provided by each client to the server respectively, where the first perturbed data is determined by a single client based on differential privacy of order one satisfied by the corresponding local attribute item;

[0140] Step 502: Detect the sensitive and non-sensitive items in each attribute item of the target information according to the first perturbed data sent by each client respectively, and feedback the detection result to each client for each client to upload the second perturbed data obtained by sampling according to whether the local attribute item is a sensitive item and satisfying differential privacy of order two based on the detection result;

[0141] Step 503: Perform corresponding business processing on the target information based on the second perturbed data uploaded by each client respectively.

[0142] It should be noted that Figure 4 the process shown in Figure 5 and Figure 2 the process shown in Figure 2 can be in accordance with Figure 4 the processes executed by the client and the server respectively in the interaction process shown in Figure 5 Therefore, the operations performed by the client and the server in the interaction process of

[0143] Reviewing the above process, it involves a business processing process with two stages. In the first stage, in the manner of local differential privacy, each client uploads the first perturbed data of the local privacy data of the target information to the server, and the server distinguishes the sensitive items and non-sensitive items among the possible attribute items of the target information according to the first perturbed data uploaded by each client respectively. Then, the server feeds back the distinction results of the sensitive items and non-sensitive items to each client. In the second stage, each client differentially performs local differential privacy based on the local privacy data corresponding to the target information. When the local data is a non-sensitive item, a higher privacy budget is given, and when the local data is a sensitive item, a lower privacy budget is given, so as to obtain the second perturbed data of the local privacy data and upload it to the server. The server performs corresponding business processing based on each second perturbed data. Through the two-stage setting, in the first stage, the server automatically distinguishes sensitive items and non-sensitive items according to the perturbed data, avoiding the interference of human factors in the process of dividing sensitive items and non-sensitive items. In the second stage, the data accuracy of non-sensitive items is improved, making the overall accuracy of business processing improved.

[0144] According to an embodiment of another aspect, there is also provided a privacy protection-based business processing device provided at a client. Figure 6 Figure 6 shows a privacy protection-based business processing device 600 according to an embodiment, which can be provided at a client and is used to perform a desensitization operation on data. As Figure 6 shown, the device 600 includes:

[0145] A first perturbation unit 601, configured to provide the server with the first perturbed data of the local attribute items of the target information, and the first perturbed data is determined based on the first differential privacy satisfied by the local attribute items of the target information;

[0146] A receiving unit 602, configured to receive the detection results of the sensitive items and non-sensitive items among the respective attribute items of the target information fed back by the server, and the detection results are determined based on the respective first perturbed data provided by each client;

[0147] A second perturbation unit 603, configured to, based on the detection results, perform sampling that satisfies the second differential privacy according to whether the local attribute items of the target information are sensitive items to obtain the second perturbed data, and upload it to the server for the server to perform corresponding business processing on the target information based on the respective second perturbed data uploaded by each client.

[0148] According to an embodiment of another aspect, there is also provided a privacy protection-based business processing device provided at a server. Figure 7 Figure 7 shows a privacy protection-based business processing device 700 according to an embodiment, which can be provided at a server and is used to perform a desensitization operation on data. As Figure 7As shown, the apparatus 700 includes:

[0149] A receiving unit 701, configured to receive first perturbation data provided by each client to the server respectively regarding the attribute items of the target information locally, where a single piece of first perturbation data is determined by a single client based on the first differential privacy satisfied by the corresponding local attribute item;

[0150] A detecting unit 702, configured to detect sensitive items and non-sensitive items in each attribute item of the target information according to the first perturbation data respectively sent by each client, and feedback the detection result to each client for each client to upload second perturbation data sampled based on whether the local attribute item is a sensitive item and satisfying the second differential privacy;

[0151] A processing unit 703, configured to perform corresponding service processing on the target information based on the second perturbation data respectively uploaded by each client.

[0152] It should be noted that Figure 6 、 Figure 7 The apparatuses 600 and 700 shown respectively correspond to Figure 4 、 Figure 5 the methods described, Figure 4 、 Figure 5 and the corresponding descriptions in the method embodiments of

[0153] also apply to the apparatuses 600 and 700, which will not be elaborated here. Figure 4 、 Figure 5 According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in combination with

[0154] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor. An executable code is stored in the memory, and when the processor executes the executable code, the method described in combination with Figure 4 、 Figure 5 is implemented.

[0155] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0156] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the technical concept of this specification. It should be understood that the above description is only the specific embodiments of the technical concept of this specification and is not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of this specification shall be included within the protection scope of the technical concept of this specification.

Claims

1. A business processing method based on privacy protection, which is used to perform business processing related to target information, and the target information corresponds to multiple attribute items; the method includes: Each client respectively provides the server with respective first perturbed data about the target information, and a single first perturbed data is determined based on the target information by a single client according to the first differential privacy satisfied by the local attribute items. The server detects sensitive items and non-sensitive items in each attribute item of the target information according to the respective first perturbed data sent by each client, and feeds back the detection result to each client. Each client respectively performs sampling that satisfies the second differential privacy to obtain corresponding second perturbed data according to whether the local attribute items of the target information are sensitive items based on the detection result, and uploads them to the server. The second differential privacy is different from the first differential privacy, and the second differential privacy satisfies: the sensitive item corresponds to the first privacy budget, the non-sensitive item corresponds to the second privacy budget, and the second privacy budget is greater than the first privacy budget. The server performs corresponding business processing on the target information based on the respective second perturbed data uploaded by each client.

2. The method according to claim 1, wherein The multiple attribute items use consecutive non-negative integers as the identification of each attribute item, and the first differential privacy satisfied by the local attribute items of a single client is determined by perturbing the attribute item identification of the local attribute items with a predetermined privacy budget.

3. The method according to claim 2, wherein, The first differential privacy adopts the Hadamard response mechanism. The local attribute item held by a single client is the first attribute item, and the first attribute item corresponds to the first item identification. The single client determines the first perturbed data corresponding to the first attribute item in the following manner: According to the element values in the row corresponding to the first item identification in the pre-constructed first Hadamard matrix, determine the first sampling probability distribution of the first perturbed data corresponding to the first attribute item. The first sampling probability distribution describes the probability that the first item identification is sampled as each column identification of the first Hadamard matrix. Perform sampling among the respective column identifications corresponding to the first Hadamard matrix according to the first sampling probability distribution, and obtain the sampled column identification as the corresponding first perturbed data.

4. The method according to claim 3, wherein, The number of attribute items of the target information is k, the order of the first Hadamard matrix is the smallest power of 2 greater than k, and the first item identification x corresponds to the (x + 2)-th row in the first Hadamard matrix.

5. The method according to claim 3, wherein, In the row corresponding to the first item identification in the first Hadamard matrix, the column identifications corresponding to the elements with a value of +1 form the first candidate set, and the column identifications corresponding to the elements with a value of -1 form the second candidate set; the first sampling probability distribution includes: The probability that a value in the first candidate set is sampled is the first probability, and the probability that a value in the second candidate set is sampled is the second probability. The first probability is greater than the second probability, and both the first probability and the second probability are determined based on a predetermined privacy budget.

6. The method according to claim 5, wherein, The server detects sensitive items and non-sensitive items in each attribute item of the target information according to the respective first perturbed data sent by each client, including: Determine the first occurrence frequency of the first item identifier in each of the first perturbation data by using the first quantity of the first perturbation data sampled from the first candidate set; Determine the preliminary distribution probability of the first item identifier based on the first occurrence frequency; According to the comparison between the preliminary distribution probability and a preset probability threshold, identify the first attribute item as a sensitive item or a non-sensitive item.

7. The method according to claim 1, wherein the second privacy budget tends to infinity.

8. The method according to claim 1, wherein for a single client, the attribute item corresponding to the target information is a second attribute item, and the second attribute item corresponds to a second item identifier; for the second attribute item, the single client determines the corresponding second perturbation data in the following manner: Match the second item identifier with each item identifier of the sensitive items and non-sensitive items in the detection result; Sample the second perturbation data corresponding to the second attribute item based on the matching result.

9. The method according to claim 8, wherein the matching result is that the second attribute item belongs to a sensitive item, and the single client samples the corresponding second perturbation result according to a second probability distribution, and the second probability distribution includes: For each column identifier in the first column identifier set, the sampling probability is a third probability; For each column identifier in the second column identifier set, the sampling probability is a fourth probability; Wherein: the third probability is the first multiple of the fourth probability, and the first multiple is determined with the natural logarithm e as the base and the first privacy budget as the exponent; the first column identifier set includes each column identifier with the element value of +1 in the row corresponding to the second attribute item in the second Hadamard matrix, and the second column identifier set includes each column identifier with the numerical value of -1 respectively in the row corresponding to the second attribute item in the second Hadamard matrix; the order S of the second Hadamard matrix is the smallest power of 2 greater than the number s of sensitive item items.

10. The method according to claim 8, wherein the matching result is that the second attribute item belongs to a non-sensitive item, and the single client samples the corresponding second perturbation result according to a third probability distribution, and the third probability distribution includes: The fifth probability of sampling as each column identifier of the second Hadamard matrix is negatively correlated with the second privacy budget; The sixth probability of sampling as t consecutive data greater than the largest column identifier of the second Hadamard matrix is positively correlated with the second privacy budget; Wherein: t is the number of non-sensitive item items; The order S of the second Hadamard matrix is the smallest power of 2 greater than the number s of sensitive item items.

11. The method according to any one of claims 8-10, wherein the service processing includes determining the frequency of each attribute item, and the service end performs corresponding service processing on the target information based on the second perturbation data respectively uploaded by each client, including: For a single sensitive item, based on each second perturbation data, the overall frequency of the sensitive item All attribute items are collected as the local frequencies of the column identifiers where the elements in the corresponding row of the second Hadamard matrix for the single sensitive item are +1 Determine the frequencies of the corresponding attribute items; For a single non-sensitive item, determine the frequency of the corresponding attribute item based on the quantity of the corresponding sampling value in each of the second perturbation data, wherein the corresponding sampling value is within the range of t consecutive data greater than the largest column identifier of the second Hadamard matrix.

12. A service processing method based on privacy protection is used for performing service processing related to target information, and multiple attribute items correspond to the target information; The method is executed by a client and includes: Provide first perturbation data regarding the local attribute items of the target information to the server, where the first perturbation data is determined based on the first differential privacy satisfied by the local attribute items of the target information; Receive the detection results of the sensitive and non-sensitive items among the respective attribute items of the target information feedback by the server, where the detection results are determined based on the respective first perturbation data provided by each client; Based on the detection results, perform sampling that satisfies the second differential privacy according to whether the local attribute items of the target information are sensitive items to obtain second perturbation data, and upload it to the server for the server to perform corresponding service processing on the target information based on the respective second perturbation data uploaded by each client. Among them, the second differential privacy is different from the first differential privacy, and the second differential privacy satisfies: the sensitive items correspond to the first privacy budget, the non-sensitive items correspond to the second privacy budget, and the second privacy budget is greater than the first privacy budget.

13. A service processing method based on privacy protection is used to perform service processing related to target information, and multiple attribute items correspond to the target information; The method is executed by the server and includes: Receive the respective first perturbation data regarding the local attribute items of the target information provided by each client to the server, and a single first perturbation data is determined by a single client based on the first differential privacy satisfied by the local attribute items; According to the respective first perturbation data sent by each client, detect the sensitive and non-sensitive items among the respective attribute items of the target information, and feedback the detection results to each client for each client to upload the second perturbation data obtained by performing sampling that satisfies the second differential privacy according to whether the local attribute items are sensitive items based on the detection results. Among them, the second differential privacy is different from the first differential privacy, and the second differential privacy satisfies: the sensitive items correspond to the first privacy budget, the non-sensitive items correspond to the second privacy budget, and the second privacy budget is greater than the first privacy budget; Perform corresponding service processing on the target information based on the respective second perturbation data uploaded by each client.

14. A service processing device based on privacy protection is used to perform service processing related to target information, and multiple attribute items correspond to the target information; The device is provided at the client and includes: A first perturbation unit configured to provide first perturbation data regarding the local attribute items of the target information to the server, where the first perturbation data is determined based on the first differential privacy satisfied by the local attribute items of the target information; A receiving unit configured to receive the detection results of the sensitive and non-sensitive items among the respective attribute items of the target information feedback by the server, where the detection results are determined based on the respective first perturbation data provided by each client; A second perturbation unit configured to, based on the detection results, perform sampling that satisfies the second differential privacy according to whether the local attribute items of the target information are sensitive items to obtain second perturbation data, and upload it to the server for the server to perform corresponding service processing on the target information based on the respective second perturbation data uploaded by each client. Among them, the second differential privacy is different from the first differential privacy, and the second differential privacy satisfies: the sensitive items correspond to the first privacy budget, the non-sensitive items correspond to the second privacy budget, and the second privacy budget is greater than the first privacy budget.

15. A service processing device based on privacy protection is used to perform service processing related to target information, and multiple attribute items correspond to the target information; The device is provided at the server and includes: A receiving unit, configured to receive first perturbation data provided by each client to the server respectively regarding the attribute items of the target information locally, where a single piece of first perturbation data is determined by a single client based on the first differential privacy satisfied by the corresponding local attribute item; A detection unit, configured to detect sensitive items and non-sensitive items in each attribute item of the target information according to the first perturbation data respectively sent by each client, and feedback the detection result to each client, so that each client uploads second perturbation data obtained by sampling the local attribute items according to whether they are sensitive items and satisfying the second differential privacy based on the detection result, where the second differential privacy is different from the first differential privacy, and the second differential privacy satisfies: the sensitive items correspond to a first privacy budget, the non-sensitive items correspond to a second privacy budget, and the second privacy budget is greater than the first privacy budget; A processing unit, configured to perform corresponding service processing on the target information based on the second perturbation data respectively uploaded by each client.

16. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method according to claim 12 or 13.

17. A computing device, comprising a memory and a processor, characterized in that, Executable code is stored in the memory, and when the processor executes the executable code, the method according to claim 12 or 13 is implemented.

Citation Information

Patent Citations

  • Cloud platform privacy protection method based on frequent item retrieval

    CN104123504A

  • A privacy protection method and system based on power system edge calculation

    CN109740346A