Differential privacy calculation method based on shuffling model
By constructing a personalized shuffling model and optimizing the randomization mechanism, the problem of limited application scope of shuffling model in non-statistical tasks is solved, and the effect of supporting personalized output and improving the intensity of privacy protection is achieved.
Patent Information
- Application Number
- CN202510240902.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-10
AI Technical Summary
When the shuffling model handles non-statistical tasks, its application scope is limited by statistical estimation and cannot support personalized output. The privacy amplification effect relies on anonymous shuffling messages, limiting the calculation type.
By constructing a personalized shuffling model, randomizing the data using a randomization mechanism optimized for non-statistical tasks, the application scope of shuffling model is expanded, so that it supports personalized output, and is independent of the randomizer.
It realizes that while protecting privacy, it supports a wider range of permutation isovariable computing, improves the intensity of privacy protection and the accuracy of personalized computing, and expands the application scope of shuffling models.
Smart Images

Figure CN120124088A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of data security and personal information protection. Specifically, it relates to a differential privacy calculation method based on a shuffle model. Background Art
[0002] Differential privacy is a technology for protecting personal data privacy. It guards against the leakage of personal information by adding noise to the data. Differential privacy ensures that the distribution of statistical analysis results is almost indistinguishable whether or not the data of a certain individual is included, thereby protecting the privacy of individuals. Differential privacy is further divided into centralized differential privacy and local differential privacy. In the centralized mode, users send data to a trusted centralized server, and the server processes the data and adds noise to protect privacy. This mode requires users to fully trust the server and is not applicable to decentralized environments. In the local mode, users perform randomization processing on the data locally and then send it to the server. This method eliminates the need to trust the centralized server but often leads to a significant reduction in utility because each user independently adds noise to the data.
[0003] The shuffle model is a new model between centralized differential privacy and local differential privacy. In the shuffle model, the data of the client is shuffled by a shuffler before being sent to the server, which makes it difficult to trace the source of the data, thereby enhancing privacy protection. This model provides better utility than local differential privacy while reducing the need to trust the centralized server.
[0004] That is to say, the shuffle model has an obvious limitation. The privacy amplification effect highly depends on anonymized shuffle messages, which greatly limits the types of computations that can be performed. So far, the only computational form achievable in the shuffle model is statistical estimation, that is, the server obtains the shuffled messages, aggregates them, and calculates a single output from them, such as a count, a sum, or a histogram. However, many real-world applications are inherently non-statistical. When multiple users gather their data together for joint computation, they expect that the personalized outputs of each user may be different. Solving this problem and expanding the application scope of the shuffle model to enable it to handle more types of tasks has become an urgent technical challenge. Summary of the Invention
[0005] In view of the problems in the related art, the present invention proposes a differential privacy calculation method based on a shuffle model to overcome the above-mentioned technical problems existing in the existing related technologies.
[0006] To this end, the specific technical solution adopted by the present invention is as follows:
[0007] A differential privacy calculation method based on a shuffle model, the method comprising:
[0008] S1. Construct a personalized shuffling model; optimize the data randomization performance of the personalized shuffling model using a randomization mechanism optimized for non-statistical tasks to obtain an improved personalized shuffling model;
[0009] S2. Use the improved personalized shuffling model to encrypt and decrypt the user's input data.
[0010] Further, constructing the personalized shuffling model includes:
[0011] Expand the application scope of the shuffling model and make the shuffling model support personalized output;
[0012] Make the shuffling model independent of the randomizer.
[0013] Further, optimizing the data randomization performance of the personalized shuffling model using a randomization mechanism optimized for non-statistical tasks to obtain an improved personalized shuffling model includes:
[0014] Configure the input domain and output domain of the randomization mechanism optimized for non-statistical tasks; use the randomization mechanism optimized for non-statistical tasks to select a first randomization value with a first probability and a second randomization value with a second probability;
[0015] Add the configured randomization mechanism optimized for non-statistical tasks to the personalized shuffling model to obtain an improved personalized shuffling model.
[0016] Further, configuring the input domain and output domain of the randomization mechanism optimized for non-statistical tasks includes:
[0017] Set the input domain and output domain to be spherical, and the radius of the output domain is larger than the radius of the input domain by a number of Minkowski distances.
[0018] Further, using the randomization mechanism optimized for non-statistical tasks to select a first randomization value with a first probability and a second randomization value with a second probability includes:
[0019] Determine the selection range according to the true value, and the first randomization value is within the selection range of the true value;
[0020] Select the second randomization value from outside the selection range of the true value;
[0021] Wherein, the first probability is greater than the second probability.
[0022] Further, using the improved personalized shuffling model to encrypt and decrypt the user's input data includes:
[0023] The user adds noise to the input data and encapsulates it into the public-key encrypted message of the computing server;
[0024] The improved personalized shuffling model is used to shuffle the public-key encrypted message, and the result of the shuffling process is sent to the computing server;
[0025] Based on the computing server, the public-key encrypted message after the shuffling process is decrypted, and equivariant computations such as permutation are performed;
[0026] Through the computing server, the result of the equivariant computation is output;
[0027] From the result of the equivariant computation, the user decrypts the entry related to their own public key.
[0028] Furthermore, the public-key encrypted message includes the user's input data, the noise added by the user, and the public key.
[0029] Furthermore, the public key is used to allow the computing server to encrypt the result of the equivariant computation and serve as an anonymous identifier for the key owner.
[0030] Furthermore, using the improved personalized shuffling model to shuffle the public-key encrypted message includes:
[0031] Using the improved personalized shuffling model to shuffle the data and the message of the public-key encrypted message, and encrypting the user's input data using the public key of the computing server.
[0032] Furthermore, outputting the result of the equivariant computation through the computing server includes:
[0033] Listing the result of the equivariant computation in the form of the user's public key and the encrypted computation result pair.
[0034] The beneficial effects of the present invention are:
[0035] (1) PIC model: Through an innovative encryption and anonymization scheme, it solves the pain points of privacy protection for non-statistical tasks and has broad application potential. Among them, the PIC model extends the application scope of the shuffling model to support personalized output, not limited to statistical estimation.
[0036] (2) Minkowski randomizer: Optimizes the utility-privacy trade-off problem of randomization, significantly improves the privacy protection strength and the accuracy of personalized computing, and performs excellently even in the case of a limited number of users. Among them, the Minkowski randomizer provides a randomized mechanism optimized for non-statistical tasks, achieving a balance between privacy protection and utility.
[0037] (3)Enhanced privacy protection: By introducing the shuffling technique, the randomness of the data is increased, making it difficult for attackers to infer individual information through the input-output relationship, thus effectively enhancing the data privacy protection ability.
[0038] (4)Improved computational efficiency: The proposed method optimizes the computational process while ensuring privacy protection, reduces the computational complexity, and improves the data processing efficiency, especially in large-scale data processing scenarios.
[0039] (5)Wide applicability: The present invention introduces a new paradigm called PIC, which extends the shuffling model of differential privacy to support a wider range of permutation-equivariant computations, not just statistical estimation. This is important because many real-world applications require personalized outputs for each user, rather than just a single aggregated output, which is the case for statistical estimation tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 is a flowchart of a differential privacy calculation method based on a shuffling model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To further illustrate the embodiments, the present invention provides drawings. These drawings are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are usually used to represent similar components.
[0043] According to an embodiment of the present invention, a differential privacy calculation method based on a shuffling model is provided.
[0044] Now, the present invention will be further described in conjunction with the drawings and specific embodiments. As Figure 1 shown, the differential privacy calculation method based on a shuffling model according to an embodiment of the present invention includes:
[0045] S1. Construct a personalized shuffling model; optimize the data randomization performance of the personalized shuffling model using a randomization mechanism optimized for non-statistical tasks to obtain an improved personalized shuffling model.
[0046] S2. Use the improved personalized shuffling model to encrypt and decrypt the user's input data.
[0047] In one embodiment, constructing the personalized shuffling model includes:
[0048] Expand the application scope of the shuffling model and make the shuffling model support personalized output; make the shuffling model independent of the randomizer.
[0049] In one embodiment, optimizing the data randomization performance of the personalized shuffling model by using a randomization mechanism optimized for non-statistical tasks to obtain an improved personalized shuffling model includes:
[0050] Configure the input domain and output domain of the randomization mechanism optimized for non-statistical tasks; use the randomization mechanism optimized for non-statistical tasks to select a first randomization value with a first probability and a second randomization value with a second probability.
[0051] Add the configured randomization mechanism optimized for non-statistical tasks to the personalized shuffling model to obtain an improved personalized shuffling model.
[0052] In one embodiment, configuring the input domain and output domain of the randomization mechanism optimized for non-statistical tasks includes:
[0053] Set the input domain and output domain to be spherical, and the radius of the output domain is larger than the radius of the input domain by a number of Minkowski distances.
[0054] In one embodiment, using the randomization mechanism optimized for non-statistical tasks to select a first randomization value with a first probability and a second randomization value with a second probability includes:
[0055] Determine the selection range according to the true value, and the first randomization value is within the selection range of the true value; select the second randomization value outside the selection range of the true value; wherein, the first probability is greater than the second probability.
[0056] In one embodiment, using the improved personalized shuffling model to encrypt and decrypt the user's input data includes:
[0057] The user adds noise to the input data and encapsulates it into the public key encrypted message of the computing server.
[0058] Use the improved personalized shuffling model to shuffle the public key encrypted message and send the shuffling result to the computing server.
[0059] Based on the computing server, decrypt the shuffled public key encrypted message and perform permutation equivariant calculation.
[0060] The computing server outputs the results of permutation-equivariant computation.
[0061] From the results of permutation-equivariant computation, the user decrypts the entries related to their own public key.
[0062] In one embodiment, the public-key encrypted message includes the user's input data, the noise added by the user, and the public key.
[0063] In one embodiment, the public key is used to allow the computing server to encrypt the results of permutation-equivariant computation and serve as an anonymous identifier for the key owner.
[0064] In one embodiment, shuffling the public-key encrypted message using an improved personalized shuffling model includes:
[0065] Using the improved personalized shuffling model to shuffle the data and the message of the public-key encrypted message, and encrypting the user's input data using the public key of the computing server.
[0066] In one embodiment, the computing server outputting the results of permutation-equivariant computation includes:
[0067] Listing the results of equivariant computation in the form of the user's public key and the encrypted computation result pair.
[0068] To facilitate understanding of the above technical solution of the present invention, the working principle of the present invention in the actual process will be described in detail below.
[0069] The PIC model and the Minkowski randomizer proposed by the present invention solve different technical challenges respectively. PIC model: Expands the application scope of the shuffling model to support personalized output, not limited to statistical estimation; Minkowski randomizer: Provides a randomization mechanism optimized for non-statistical tasks, achieving a balance between privacy protection and utility.
[0070] The present invention aims to explore the application possibilities of the shuffling model in non-statistical privacy computing tasks, especially those applications that are difficult to effectively handle by existing methods, such as spatial crowdsourcing, advertising allocation, combinatorial optimization, location-based social systems, and incentivized federated learning. By introducing new technical means, such as random keys and noise addition, the present invention can achieve high-efficiency and high-utility data processing and computation while protecting user privacy.
[0071] The present invention proposes a privacy computing framework based on the shuffle model. PIC can achieve personalized output while protecting privacy and amplify privacy through shuffling. The present invention proposes a specific protocol for implementing the PIC method. By using one-time public keys, the protocol enables users to receive their own outputs without compromising anonymity, which is crucial for privacy amplification. In addition, an optimal randomizer - Minkowski response designed for the PIC model is proposed to improve practicality. The security and privacy properties of the PIC protocol are proven. Theoretical analysis and experiments demonstrate the ability of PIC to handle non-statistical computing tasks, as well as the efficacy of PIC and the Minkowski randomizer in achieving better practicality than existing solutions. The main inventive points include:
[0072] PIC protocol
[0073] As an overall framework for privacy amplification, the PIC model is designed independently of specific randomizers. Therefore, it can not only support the Minkowski randomizer but also be used in conjunction with other randomization mechanisms.
[0074] 1. Each user adds noise to their own data and encapsulates it (along with the one-time public key) into a public-key encrypted message for the computing server.
[0075] 2. The shuffler shuffles the encrypted messages and then sends them to the computing server.
[0076] 3. The server can decrypt these messages and perform permutation-equivariant computations.
[0077] 4. The server publishes the computation results, each of which is encrypted with the corresponding user's one-time public key.
[0078] 5. Users can download the entire list and decrypt the entries related to their own public keys, thus maintaining anonymity. Users can also use the public keys to establish a secure communication channel with other matching parties to complete the PIC task.
[0079] Minkowski response
[0080] The Minkowski randomizer is an independent technical innovation point proposed in the present invention. Its design aims to optimize the data randomization performance in the PIC model, but its algorithm itself can be adapted to a wider range of privacy computing tasks. Therefore, this randomizer can be used alone or in combination with other privacy protection models, with high generality and innovation.
[0081] The new randomization mechanism is designed specifically for personalized computing tasks and the steps are as follows:
[0082] 1. The output domain of the randomizer is defined as a ball whose radius is r Minkowski distances larger than the radius of the input domain, and the magnitude of r is determined by the privacy budget.
[0083] 2. The randomizer selects a value close to the true value (within r) with a relatively high probability and a value from elsewhere with a relatively low probability to achieve high utility.
[0084] 3. The present invention proves that the upper bound of the error of the Minkowski response matches the lower bound of the error of all possible randomizers in the PIC model, thus achieving asymptotic optimality.
[0085] 4. Compared with using existing LDP (Local Differential Privacy) randomizers in the PIC model, the Minkowski randomizer can provide significantly better utility, and this advantage becomes obvious even when the number of users is not very large (on the order of 102).
[0086] The PIC method includes:
[0087] Parameters: m ∈ N, groups G 1 , G 2 , …, G m , where n i = |G i |, representing the number of squares in group G i ; the data randomization mechanism for group G i is R i ; server S.
[0088] Function: Accept the n = ∑ i∈[m] n i input values and the description of the function f to be computed on the server, and perform the following steps:
[0089] 1. For all i ∈ [m] and j ∈ [|G i |], compute x′ i,j ← R i (x i,j ).
[0090] 2. Sample m random permutations π 1 , π 2 , …, π m , where π i : |n i | → |n i |, shuffle the input values and obtain:
[0091] Compute:
[0092] 3. Send to user u i,j , i ∈ [m] and j ∈ [n i . Additionally, send L and f(L) to server S.
[0093] where m ∈ N represents there are m user groups G 1 , G 2 , …, G m . n i =|G i | represents the number of users in the i-th user group G i is n i . R i represents the data randomization mechanism used by the i-th user group G i . S represents a server. represents the input data provided by all users, where x i,j represents the input data of the j-th user in the i-th user group. f represents the function that the server wants to calculate. x' i,j ←R i (x i,j ) means that the input data of user x i,j is processed through the randomization mechanism R i to obtain the privacy-protected data x' i,j . π 1 , π 2 ,..., π m represent m random permutations respectively adopted for m user groups.
[0094] represents the list of privacy-protected data after random permutation. represents the output result obtained by calculating the function f for the list L of privacy-protected data after random permutation. represents, for the j-th user u i,j in user group i, the message x”i,j obtained after being processed by the randomizer R i,j , and the result of rearranging these messages through a random permutation π i . [L, f(L)] means sending the intermediate result L and the final result f(L) in the calculation process to server S.
[0095] The Shuffle method includes:
[0096] Parameter: n ∈ N: natural number. n parties: P 1 , P 2 , …, P n. A server S. The set C of corrupted parties. The leakage L(π)=[i,π(i)], where i∈C represents the information of the permutation π used.
[0097] Method: After receiving n inputs {x 1 ,P 2 ,…,P n} from P i : i∈[n]
[0098] Sample a random permutation π∈S n , where S n represents a random permuter that permutes the messages y 1 ,y 2 ,…,y n from n users. Specifically, Sn is a function S:Yn→Yn that applies a random permutation to the n input messages to generate a new permuted message sequence.
[0099] Define {y i} i∈[n] such that y i =x π(i) .
[0100] Send {y i} i∈[n] to the server S. Additionally, since S∈C, send L(π) to the adversary S.
[0101] Here, n∈N represents that there are n parties P 1 ,P 2 ,...,P n . S represents a server. C represents the set of corrupted parties. L(π)=[i,π(i)], where i∈C means that for the corrupted party i∈C, the server S can learn its mapping π(i) in the random permutation π. {x i} i∈[n] represents the input data provided by each of the n parties. π∈S n represents the random permutation applied to the n input data. {y i} i∈[n] represents the output data sequence obtained after the random permutation π, where y i =x π(i) , and y i represents the message after the user i processes it through the local randomizer R, and y i =R(x i ). x π(i) represents the message after being shuffled by the random permuter S. Here, π is a random permutation function.
[0102] Finally, this function describes a simple data obfuscation process: the participating parties provide input data, obfuscate the data using a random permutation, and send the obfuscated data to the server. If the server is corrupted, part of the permutation information L(π) is also leaked to the server.
[0103] The encrypted data processing steps of the PIC method are as follows:
[0104] The participating parties add noise to their own data and encrypt it before sending it to the server.
[0105] The server shuffles the encrypted data.
[0106] The server computes the function f on the shuffled and sanitized user inputs to generate personalized outputs for each user.
[0107] The server publishes the computation results on a public bulletin board, listed in the form of the users' one-time public keys and encrypted computation result pairs.
[0108] Each user downloads the list from the public bulletin board, finds the entry corresponding to their public key, and decrypts it to obtain the personalized computation result.
[0109] Specification of the public key encryption scheme:
[0110] Π = (Gen, Enc, Dec); where Gen represents the key generation algorithm, used to generate a public key and private key pair.
[0111] Enc represents the encryption algorithm, used to encrypt a message with the public key.
[0112] Dec represents the decryption algorithm, used to decrypt a message with the private key.
[0113] The security parameter λ, representing the security level of the system, usually determines the key length.
[0114] The public key pk c , represents the public key generated by the server.
[0115] User group: G i (i ∈ [m]), represents the i-th user group, where the value range of i is from 1 to m.
[0116] Key pair: (pk i,j , sk i,j ) ← Gen(λ), represents the public key pk i,j and private key sk i,j generated by each user.
[0117] Input data: x i,j , represents the private input data of the j-th user in the i-th group.
[0118] Randomization mechanism: R i , representing the data randomization mechanism for the i-th user group.
[0119] Perturbed data: x′ i,j ←R i (x i,j ), representing the data of user u i,j after randomization.
[0120] Encrypted message: x″ i,j ←Enc pkc (pk i,j |x′ i,j ), representing the encrypted perturbed data x′ c using the server's public key pk i,j and the user's public key pk i,j .
[0121] Shuffling function: F Shuffle , representing the function for shuffling messages.
[0122] Permutation: π, representing a random permutation of a set of data.
[0123] List: L, representing the list of all groups after the server's decryption.
[0124] Computation function: f, representing the computation function performed by the server on the decrypted list.
[0125] One of the implementation methods in specific applications:
[0126] 1. The server publishes global parameters including:
[0127] (1) The specification Π=(Gen,Enc,Dec) of the public key encryption scheme; (2) The security parameter λ; (3) Its own public key pkc generated by calling Gen(λ); (4) For each user group G i (i ∈ [m]), the data randomization mechanism R i .
[0128] 2. The user generates a key pair:
[0129] Denote the j-th user in group G i as u i,j . Each user generates a key pair (pk i,j , sk i,j )←Gen(λ), and then each user randomizes their private data x Shuffle using the mechanism function F i,j ;
[0130] n ∈ N; server S; set of corrupted parties C; for the permutation π used, the leakage L(π) = {[i, π(i)]}, i ∈ C.
[0131] 3. Data Shuffling and Encryption:
[0132] Function F Shuffle Receives the perturbed data z from all users i , and performs a random shuffle, defining {y i} i∈ n such that y i = x π ( i ) and sends {y i} i∈ n to server S. Additionally, if S ∈ C, then L(π) is sent to the adversary S.
[0133] Function F Shuffle R i and obtains x' i,j ← R i (x i,j ). Then, the purified inputs are concatenated with their own public keys and encrypted with the server's public key .
[0134] 4. Message Shuffling:
[0135] Samples a random permutation function π. Defines y i as the shuffled output.
[0136] Sends the shuffled message y i to server S. If S is corrupted, the permutation function π is also sent to the adversary S.
[0137] Through such a random shuffling process, privacy protection can be enhanced because the server cannot identify the original data of each user. At the same time, this random shuffling also provides anonymized inputs for subsequent calculations.
[0138] 5. Server Processing:
[0139] The server decrypts each shuffled message and obtains a list L of all m groups as follows: Then, computes the function f on L to generate outputs for each anonymous user: In , represents the public key of the j-th user in the first group, represents the private data of the j-th user in the first group after data randomization, n 1 is the number of users in the first group, represents the public key of the j-th user in the m-th group, represents the private data of the j-th user in the m-th group after data randomization, n m represents the number of users in the m-th group.
[0140] Each user also includes a one-time public key in the encrypted message. This one-time public key serves two purposes: allowing the server to encrypt the calculation result so that it can only be decrypted by the owner of the corresponding private key. Acting as an anonymous identifier of the key owner.
[0141] 6. Result publication:
[0142] The server publishes the calculation results as a list to the public bulletin board: Each entry consists of a public key and the calculation result encrypted using that public key. in represents the public key of the j-th user in the i-th group, π i (j) represents the position of the j-th user in the i-th group in the random permutation π, represents using the public key of the j-th user in the i-th group to encrypt the calculation result n i The number of users in the i-th group.
[0143] 7. User decrypts the result:
[0144] Users can download the entire list and decrypt the entries related to their own public keys, remaining anonymous. Users can also use their public keys to establish a secure communication channel with other matching parties, ultimately completing the PIC task.
[0145] The second implementation method in specific applications:
[0146] PIC is used to parameterize the list of parties that have been compromised by an adversary and allied with the server. For these, the adversary should know the correspondence of their messages before and after shuffling, so Shuffle leaks this part of the permutation to the adversary.
[0147] 1. The server publishes global parameters, including:
[0148] (1) The specification Π = (Gen, Enc, Dec) of the public-key encryption scheme, (2) The security parameter λ, (3) Its own public key pk generated by calling Gen(λ) c , (4) For each user group G i (i ∈ [m]), the data randomization mechanism R i .
[0149] 2. For the group G iThe j-th user in is denoted as u i,j . Each user generates a key pair (pk i ,j, sk i ,j) ← Gen(λ). Then, each user randomizes their private data x Shuffle using the mechanism function F i,j .
[0150] n ∈ N; n parts are P 1 , P 2 , ···, P n ; server S; corrupted party set C; for the permutation π used, the leakage L(π) = {[i, π(i)]}, i ∈ C. Functionality: When receiving n inputs x 1 , P 2 , ···, P n respectively, sample a random permutation π ∈ S i . Define y n ,i ∈ [n] such that y i = x i . Send y π(i) ,i ∈ [n] to the server S. Additionally, if S ∈ C, send L(π) to the adversary S i .
[0151] Function F Shuffle R i and obtain x' i,j ← R i (x i ,j). Then, the purified inputs are concatenated with their own public keys and encrypted with the server's public key x'' i,j ← Enc pkc (pk i,j |x' i,j ). R i represents the local randomizer used by the i-th group of users
[0152] 3. Each user in group G i calls F with x'' i,j , and F Shuffle outputs the shuffled messages Shuffle to the server, where π is the (secret) random permutation used during F -1 of group G i . Shuffle
[0153] 4. The server decrypts each group of shuffled messages and obtains a list L of all m groups
[0154]
[0155] Then, calculate the function f on L and generate an output for each anonymous user:
[0156]
[0157] 5. The server posts the calculation result as a pair list to the public bulletin board:
[0158] Let i represent a certain user group, j represent a certain user in the user group, m represent there are m groups of users, and n i represent the total number of users within the group, and π 1 (j) represents a random permutation function used to shuffle the order of users to ensure privacy. pk i , π i (j) represents the public key of the user, which is used to encrypt data and remains bound to the corresponding user after shuffling. The personalized result calculated by the server for the user, represents the encrypted calculation result.
[0159] Each user also includes a one-time public key in the encrypted message. This one-time public key has two purposes: allowing the server to encrypt the calculation result so that it can only be decrypted by the owner of the corresponding private key. Acting as an anonymous identifier for the key owner.
[0160] 6. The server publishes a list, and each entry consists of a public key and the calculation result encrypted using that public key.
[0161] 7. Each user downloads the list, finds the entry with their own public key in the list, and decrypts the payload to obtain the calculation result.
[0162] 8. Users can also use their public keys to establish a secure communication channel with other matching parties to finally complete the PIC task.
[0163] Traditional differential privacy techniques are vulnerable to analysis and inference by attackers during data processing and calculation, thus failing to effectively protect user privacy. In addition, existing differential privacy models often have problems of high computational complexity and low efficiency when dealing with large-scale data.
[0164] The present invention proposes a differential privacy calculation method (PIC) based on a shuffle model. By introducing a shuffle operation during data processing, it effectively increases the randomness and unpredictability of the data, thereby improving the data privacy protection ability.
[0165] The key advantage of PIC over existing shuffling models is that it can support a wider range of permutation-equivariant computations that require personalized outputs for each user, while still enjoying the privacy amplification benefits of the shuffling model. This makes PIC a valuable tool in various practical applications for processing personal data. The advantages specifically include:
[0166] (1) PIC model: By means of an innovative encryption and anonymization scheme, it solves the pain points of privacy protection for non-statistical tasks and has broad application potential.
[0167] (2) Minkowski randomizer: Optimizes the utility-privacy trade-off of randomization, significantly improving the privacy protection strength and the accuracy of personalized computations, and performing excellently even in the case of a limited number of users.
[0168] (3) Enhanced privacy protection: By introducing the shuffling technique, it increases the randomness of the data, making it difficult for attackers to infer individual information through the input-output relationship, thus effectively enhancing the data privacy protection ability.
[0169] (4) Improved computational efficiency: The proposed method optimizes the computational process while ensuring privacy protection, reduces the computational complexity, and improves the data processing efficiency, especially in large-scale data processing scenarios.
[0170] (5) Wide applicability: This invention introduces a new paradigm called PIC, which extends the shuffling model of differential privacy to support a wider range of permutation-equivariant computations, not just statistical estimation. This is important because many real-world applications require personalized outputs for each user, rather than just a single aggregated output, which is the case for statistical estimation tasks.
[0171] Specific applications include:
[0172] (1) Machine learning and data mining: PIC can generate personalized outputs while protecting privacy, which is crucial for applications such as location-based social systems and federated learning with incentive mechanisms. The privacy amplification benefits of the shuffling model can significantly improve the utility of these applications compared to local differential privacy (LDP) models.
[0173] (2) Combinatorial optimization: In spatial crowdsourcing and advertising allocation, PIC can match users with workers / advertisers based on the users' private information and provide a personalized best-matching list for each party. Such personalized outputs are crucial for these combinatorial optimization tasks.
[0174] (3) Information Retrieval: For mobile search and location-based systems, PIC can generate personalized query results (such as nearby restaurants or neighboring users) that depend on the private information of the enquirer while protecting their privacy.
[0175] (4) Incentive Mechanism: In federated learning, PIC can calculate personalized incentives (such as monetary tokens) based on the contributions of users (such as Shapley values, a distribution method in game theory) while protecting their privacy. This is crucial for encouraging good participation.
[0176] The PIC paradigm achieves personalized output while maintaining privacy and enjoys the privacy amplification benefits of the shuffle model. This invention proposes a specific protocol to implement the PIC paradigm, allowing users to receive personalized output without compromising anonymity, which is crucial for the privacy amplification effect. It is also proven that PIC can handle non-statistical computing tasks. Theoretical analysis and empirical evaluation show that PIC and the proposed Minkowski randomizer can achieve better utility than existing solutions.
[0177] In summary, the PIC paradigm introduced in this invention greatly expands the applicability of the differential privacy shuffle model beyond statistical estimation, enabling a wide range of personalized computing tasks while maintaining privacy and achieving efficient reuse.
[0178] Theoretical Verification and Experimental Support
[0179] Through comparative experiments, the proposed method is superior to existing differential privacy models in terms of privacy protection strength and computational efficiency. The specific experimental results are as follows: Compared with traditional methods, the method of this invention reduces the computing time by about 30% and the privacy leakage probability by more than 50%. Through multiple experimental verifications, the differential privacy computing method based on the shuffle model provided by this invention shows good stability and applicability under different data scales and application scenarios. Through detailed theoretical analysis and experimental data, the effectiveness and superiority of the proposed technical solution are proven, further enhancing the credibility and practicality of this method in practical applications.
[0180] In summary, the differential privacy computing method based on the shuffle model proposed in this invention has made significant technological progress and improvement in solving the privacy protection and computational efficiency problems faced by traditional differential privacy technologies, providing new ideas and solutions for future privacy protection research and applications. The core innovations of this invention include the differential privacy computing framework based on the shuffle model (PIC model) and the specially designed Minkowski randomizer. These two technologies have their own independent application scenarios and innovation values.
[0181] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A differential privacy computing method based on a shuffle model, characterized in that: The method includes: S1. Build a personalized shuffle model; optimize the data randomization performance of the personalized shuffle model using a randomization mechanism optimized for non-statistical tasks to obtain an improved personalized shuffle model; S2. Use the improved personalized shuffling model to encrypt and decrypt the user's input data.
2. According to claim 1, a differential privacy computing method based on a shuffle model is characterized in that: The constructing of a personalized shuffle model comprises: Expand the application scope of the shuffle model and enable the shuffle model to support personalized output; Make the shuffle model independent of the randomizer.
3. The differential privacy computing method based on the shuffle model according to claim 1, characterized in that: The data randomization performance of the personalized shuffling model is optimized by using the randomization mechanism optimized for non-statistical tasks, and the improved personalized shuffling model includes: An input domain and an output domain of a randomization mechanism optimized for non-statistical tasks are configured; a first randomization value is selected with a first probability and a second randomization value is selected with a second probability using the randomization mechanism optimized for non-statistical tasks; The configured randomization mechanism optimized for non-statistical tasks is added to the personalized shuffling model to obtain an improved personalized shuffling model.
4. The differential privacy computing method based on the shuffle model according to claim 3, characterized in that: The input domain and output domain of the randomization mechanism configured for non-statistical tasks optimization include: The input and output domains are set to be spherical, and the radius of the output domain is larger than the radius of the input domain by a number of Minkowski distances.
5. The differential privacy computing method based on the shuffle model according to claim 3 is characterized in that: The selecting a first randomization value with a first probability and selecting a second randomization value with a second probability by using a randomization mechanism optimized for non-statistical tasks comprises: Determine a selection range according to the true value, and the first randomized value is within the selection range of the true value; Selecting a second randomized value from outside the selection range of the true value; Among them, the first probability is greater than the second probability.
6. The differential privacy computing method based on the shuffle model according to claim 1, characterized in that: The method of using the improved personalized shuffle model to encrypt and decrypt the user's input data includes: The user adds noise to the input data and encapsulates it into a public key encrypted message on the computing server; Using the improved personalized shuffling model to shuffle the public key encrypted message, and sending the shuffling result to the computing server; Based on the computing server, decrypt the shuffled public key encrypted message and perform permutation equivariant calculation; Output the results of permutation equivariant calculations through the calculation server; From the result of the permutation equivariant computation, the user decrypts the entry associated with his own public key.
7. The differential privacy computing method based on the shuffle model according to claim 6, characterized in that: The public key encrypted message includes the user's input data, the noise added by the user and the public key.
8. The differential privacy computing method based on the shuffle model according to claim 7, characterized in that: The public key is used to allow the computing server to encrypt the permutation equivariant computing result and serve as an anonymous identifier of the key owner.
9. The differential privacy computing method based on the shuffle model according to claim 6, characterized in that: The method of shuffling the public key encrypted message using the improved personalized shuffling model includes: The improved personalized shuffling model is used to perform data scrambling and message shuffling on the public key encrypted message, and the user's input data is encrypted using the public key of the computing server.
10. The differential privacy computing method based on the shuffle model according to claim 6, characterized in that: Outputting the result of the permutation equivariant calculation through the calculation server includes: The results of the equivariant calculation are listed in the form of a pair of the user's public key and the encrypted calculation result.