Data query method, system and equipment based on shuffling differential privacy and medium
By introducing policy query sets and shuffling differential privacy technology on the user side, the communication overhead problem during large-scale continuous interval query is solved, and efficient privacy protection and data query accuracy are achieved.
Patent Information
- Application Number
- CN202510473626.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-15
AI Technical Summary
In data query scenarios based on shuffle differential privacy, the user needs to independently generate an answer message for each query and inject noise, resulting in a sharp increase in communication overhead during large-scale continuous interval queries.
By introducing a policy query set, which is a real subset of the original query set. The user only needs to answer each query item in the policy query set, and realize anonymization processing and noise control through shuffler and differential privacy algorithms.
It reduces the number of messages sent by the user, reduces communication overhead, and ensures the privacy protection of the overall data query process, improving query accuracy and efficiency.
Smart Images

Figure CN120011401A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of privacy protection technology, and in particular, to a data query method, system, electronic device, computer-readable storage medium, and computer program product based on shuffled differential privacy. Background Art
[0002] In data query scenarios based on shuffle differential privacy, related methods usually execute shuffle protocols independently for a single query, and the user needs to generate a separate answer message for each query and inject noise to protect individual privacy. However, when it comes to large-scale continuous interval queries, the user needs to participate in all query items, resulting in a linear increase in the number of messages sent by a single user and the query scale, especially when the data domain is large, the communication overhead increases sharply. Summary of the invention
[0003] In view of this, one or more embodiments of the present specification provide a data query method, system, electronic device, computer-readable storage medium, and computer program product based on shuffle differential privacy.
[0004] To achieve the above objectives, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, a data query system based on shuffle differential privacy is proposed, including a shuffler, an analyzer, and a user terminal holding private data; The user end is used to obtain a policy query set, which is a proper subset of an original query set, and the original query set is used to describe all interval queries that need to be answered in the data domain where the private data is located, and the query results of the query items in the original query set except the policy query set can be obtained by linearly combining the query results of at least two query items in the policy query set; answer each query item in the policy query set based on the private data held by itself and the differential privacy algorithm, generate an answer message of the policy query set, and send it to the shuffler; The shuffler is used to anonymize and shuffle the reply messages sent by the user end to form a message queue, and send the message queue to the analyzer; The analyzer is used to determine the query results of all query items in the policy query set based on the message queue and the differential privacy algorithm; based on the mapping relationship between the original query set and the policy query set, convert the query results of all query items in the policy query set into the query results of all query items in the original query set.
[0005] According to a second aspect of an embodiment of this specification, a data query method based on shuffled differential privacy is provided, which is applied to a user end in a data query system; the method includes: Obtaining a policy query set, wherein the policy query set is a proper subset of an original query set, the original query set is used to describe all interval queries that need to be answered in the data domain where the private data held by the user terminal is located, and the query results of the query items in the original query set other than the policy query set can be obtained by linearly combining the query results of at least two query items in the policy query set; Based on the private data held by itself and the differential privacy algorithm, each query item in the policy query set is answered, and an answer message of the policy query set is generated and sent to the shuffler in the data query system.
[0006] According to a third aspect of an embodiment of this specification, a data query method based on shuffled differential privacy is provided, which is applied to an analyzer in a data query system; the method includes: Receive a message queue for a strategy query set sent by a shuffler in the data query system; wherein the strategy query set is a proper subset of an original query set, the original query set is used to describe all interval queries that need to be answered in the data domain where the private data is located, and the query results of query items in the original query set other than the strategy query set can be obtained by linearly combining the query results of at least two query items in the strategy query set; Determine query results for all query items in the policy query set based on the message queue and the differential privacy algorithm; Based on the mapping relationship between the original query set and the policy query set, query results of all query items in the policy query set are converted into query results of all query items in the original query set.
[0007] According to a fourth aspect of the embodiments of this specification, there is provided an electronic device, including: processor; a memory for storing processor-executable instructions; When the processor executes the executable instructions, it is used to implement the method described in the second aspect or the third aspect.
[0008] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in the second aspect or the third aspect are implemented.
[0009] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program, which implements the steps of the method described in the second aspect or the third aspect when executed by a processor.
[0010] The technical solutions provided by the embodiments of this specification may have the following beneficial effects: In the embodiments of the present specification, a policy query set is used, which is a proper subset of the original query set, and some query results in the original query set can be restored by a linear combination of multiple query results in the policy query set, so that the user end only needs to answer each query item in the policy query set, without having to answer each query item in the original query set, thereby reducing communication overhead; by introducing shuffling technology and differential privacy algorithm, anonymization processing and noise control of user answer messages are achieved, thereby ensuring that the overall data query process meets privacy protection requirements; after obtaining the query results of all query items in the policy query set, the analyzer can restore the complete original query answer from the query results of all query items in the policy query set through the mapping relationship between the original query set and the policy query set.
[0011] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a schematic diagram of an application scenario of data query based on shuffled differential privacy provided by an exemplary embodiment.
[0013] Figure 2 It is a structural diagram of a data query system based on shuffled differential privacy provided by an exemplary embodiment.
[0014] Figure 3 It is a schematic diagram of the relationship and data flow among an original query set, a policy query set and a message queue provided by an exemplary embodiment.
[0015] Figure 4 It is a schematic diagram of a binary tree provided by an exemplary embodiment.
[0016] Figure 5 It is a flowchart of a data query method based on shuffled differential privacy provided by an exemplary embodiment.
[0017] Figure 6 It is a flowchart of another data query method based on shuffled differential privacy provided by an exemplary embodiment.
[0018] Figure 7 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment. DETAILED DESCRIPTION
[0019] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0020] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0021] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0022] Here are some explanations of the terms mentioned in this specification: 1. Differential Privacy: Differential Privacy (DP) is a privacy enhancing technology (PETs) that aims to reconcile the contradiction between data utilization efficiency and personal privacy protection. It ensures that when publishing the statistical data of the database, it is impossible to accurately trace the information of any specific user by artificially adding controllable noise in the data processing link. The core idea of differential privacy is that for any two similar databases (that is, only one individual's record is different), the query results processed by the differential privacy algorithm should be highly similar, so that the observer cannot determine whether the information of a specific individual is included. This feature is achieved by injecting random noise into the query results, ensuring the privacy protection of data analysis while still maintaining the usefulness of the data.
[0023] 2. Shuffled differential privacy: Mainly used for large-scale distributed data collection, analysis, and processing. Shuffled differential privacy assumes that there is a shuffler between the user end and the analyzer. By shuffling messages, the analyzer cannot distinguish which user end sent which message, so as to achieve anonymous communication. The implementation of the shuffler can rely on specific routing protocols, trusted security hardware, etc. Shuffled differential privacy can greatly improve the accuracy of the results while ensuring privacy security.
[0024] Therefore, shuffle differential privacy has the following advantages: (1) Enhanced privacy protection without the need to fully trust a single entity: Users do not need to directly trust the data collector, because the shuffling process can be performed by a neutral third party, or decentralized shuffling can be achieved through protocols such as secure multi-party computing and anonymous network communication, further dispersing trust risks. (2) Reduce noise and increase data accuracy: Since noise is added to the data as a whole after shuffling, the total amount of noise required to be added is reduced, thereby improving the accuracy of data analysis. (3) Wide applicability: Applicable to a variety of scenarios, especially those environments that require effective data analysis while protecting user privacy.
[0025] 3. Counting queries are a simple query type commonly used in database queries. Its purpose is to count the number of data records that meet specific conditions. A basic SQL counting query format is such as "SELECTCOUNT(*) FROM table WHERE condition;" where table is the name of the query table and condition is the condition for filtering data. The COUNT(*) function will count the number of rows in the table that meet the specified conditions.
[0026] Interval query is a type of counting query, which aims to count the number of data records that meet a certain continuous interval.
[0027] In the interval query scenario under shuffled differential privacy, each user holds private data. Based on a given query item (such as the number of data records in a certain interval), all users jointly answer the question based on their own private data (whether it meets the query interval) and add appropriate differential privacy noise protection.
[0028] For example, see Figure 1, each user sends a real message "1" if the private data it holds meets the interval indicated by the query item, otherwise it does not send it; and sends a noise message "1" with a certain probability. The shuffler collects the messages sent by all users, randomly shuffles the order of the messages, and passes the shuffled messages to the analyzer to ensure the anonymity of the user identity. After receiving all the messages sent by the shuffler, the analyzer obtains a query result that meets the differential privacy requirements.
[0029] Based on the problems in related technologies, please refer to Figure 2 as well as Figure 3 , the embodiment of this specification provides a data query system based on shuffle differential privacy, including a shuffler, an analyzer, and a user end holding private data. Among them, "holding" means the ownership of the data, such as Figure 2 As shown, there can be multiple user terminals, and different user terminals include private data that is not disclosed to the outside. It can be understood that this embodiment does not impose any restrictions on the storage method of private data belonging to the user terminal, and specific settings can be made according to actual application scenarios.
[0030] The user end is the data holder, whose main responsibility is to answer a set of queries based on the private data it holds. The user end is used to obtain the policy query set, which is a proper subset of the original query set. The original query set is used to describe all interval queries that need to be answered in the data domain where the private data is located, and the query results of the query items in the original query set except the policy query set can be obtained by linearly combining the query results of at least two query items in the policy query set; based on the private data it holds and the differential privacy algorithm, it answers each query item in the policy query set, generates an answer message for the policy query set, and sends it to the shuffler.
[0031] Among them, see Figure 3, the strategic query set refers to a subset obtained by optimizing the original query set, which contains fewer query items than the original query set. The original query set is used to describe all interval queries that may need to be answered in the data domain where the private data is located. For example, in the data domain [1,4], the original query set includes 10 continuous interval queries, namely [1,1], [1,2], [1,3], [1,4], [2,2], [2,3], [2,4], [3,3], [3,4], [4,4]. The strategic query set is a true subset optimized from the original query set. The query items in the strategic query set are carefully designed so that the query results of the query items in the original query set except the subset can be obtained by linearly combining the results of at least two query items in the strategic query set. By adopting the strategic query set, the user end only needs to answer fewer query items than the original query set, which can greatly reduce the communication cost and noise accumulation, and improve the overall query accuracy and efficiency.
[0032] The shuffler is used to anonymize and shuffle the reply messages sent by the user end to form Figure 3 The message queue shown in Figure 1 is then sent to the analyzer. The main function of the shuffler is to disrupt the order of messages sent by each client, preventing the analyzer from inferring the source of a message through the message order or other information, thereby further enhancing privacy protection. Shuffling can effectively isolate the direct association between a single client and its reply message, ensuring that the anonymity requirements of the differential privacy mechanism are met.
[0033] like Figure 3 As shown, the analyzer is used to determine the query results of all query items in the policy query set based on the message queue and the differential privacy algorithm; based on the mapping relationship between the original query set and the policy query set, the query results of all query items in the policy query set are converted into the query results of all query items in the original query set. The mapping relationship ensures that the results of the original query set can be derived even when protecting privacy, and the error of the results of the original query set is within a controllable range, so as to achieve efficient answers to all interval queries described by the original query set.
[0034] In this system, a policy query set is used, which is a proper subset of the original query set, and some query results in the original query set can be restored by a linear combination of multiple query results in the policy query set, so that the user end only needs to make answers based on the policy query set, reducing communication overhead; by introducing shuffling technology and differential privacy algorithm, anonymization and noise control of user answer messages are achieved, thereby ensuring that the overall data query process meets strict privacy protection requirements; through the mapping relationship between the original query set and the policy query set, the complete original query answer can be restored from the query results of all query items in the policy query set.
[0035] In some embodiments, the policy query set can be generated in the following manner: the data domain involved in the original query set is used as the root node of the binary tree, and each parent node of the binary tree is split into two child nodes until the interval length of the child node is 1, and the splitting is stopped; wherein, the interval of any parent node is the result of merging the intervals of the two child nodes of the parent node; finally, based on all the nodes in the binary tree, the query items in the policy query set are determined. In this embodiment, constructing a binary tree can capture the hierarchical structure of interval queries in the data domain, and the query results of the parent node can be restored through the linear combination of its child nodes. This can not only maintain the integrity of the information, but also optimize using the hierarchical relationship, so as to achieve the purpose of restoring all target queries at a lower communication cost.
[0036] For example, the complete data domain [1, B ] as the root node, and each parent node [ a , b ] is split into two child nodes: the left child node is , the right child node is: , when the interval length is 1, stop splitting and generate leaf nodes. All non-leaf nodes and leaf nodes constitute the strategy query set A.
[0037] For example, see Figure 4 In the data domain [1,4], the original query set includes 10 continuous interval queries, namely [1,1], [1,2], [1,3], [1,4], [2,2], [2,3], [2,4], [3,3], [3,4], [4,4]. These query items cover all interval combinations and can fully reflect the statistical information of each sub-interval in the data domain. By constructing Figure 4The binary tree shown takes the entire data domain [1,4] as the root node of the binary tree. Corresponding to the interval [1,4], [1,4] is split into two child nodes: [1,2] and [3,4]. The child node [1,2] is further split to obtain [1,1] and [2,2]. The child node [3,4] is further split to obtain [3,3] and [4,4]. Finally, based on all the nodes in the constructed binary tree, the strategy query set is determined. The strategy query set obtained in this way contains 7 query items instead of the original 10 query items.
[0038] In other embodiments, in addition to using a binary tree structure to construct a policy query set, other multiple optimization methods may be used to construct a policy query set. This embodiment does not impose any restrictions on this, and specific settings may be made according to actual application scenarios.
[0039] Exemplarily, a multi-level tree structure can be used to divide the data domain, where each node represents a data interval. A policy query set can be constructed based on some nodes in the tree. The user end only answers some nodes in the tree, and the query results of any interval can be reconstructed through the linear relationship between parent nodes and child nodes.
[0040] For example, a wavelet transform method can be used to construct a strategic query set, where the data domain is decomposed into coefficients of different frequencies through a wavelet basis function (such as Haar wavelet), and each coefficient corresponds to a basic query (i.e., a query with an interval length of 1). Other non-basic queries (i.e., queries with an interval length greater than 1) can be expressed as a linear combination of these wavelet bases.
[0041] Exemplarily, the sparsity of data may be utilized to select key query items that have a greater impact on the results from the original query set as query items in the strategic query set.
[0042] In some embodiments, the generation of the policy query set does not rely on the private data of the user end, but only requires the structural information of the target query set, such as the data domain size, the query type (such as interval query), etc. Then in one possible implementation, the policy query set can be optimized by the relevant equipment of the data service provider or the relevant equipment of the trusted third party, and then the policy query set is distributed as a public parameter to the analyzer and all user ends when the system is initialized. In another possible implementation, the policy query set can be optimized by the analyzer, and then the policy query set is distributed as a public parameter to all user ends when the system is initialized.
[0043] In some embodiments, the execution process of the shuffler, analyzer, and user end in the data query system based on shuffle differential privacy is further described below: After obtaining the policy query set (different query items in the policy query set correspond to different query identifiers), each user terminal generates a real message carrying the query identifier related to the query item for each query item in the policy query set if it is determined to be in compliance with the query item based on its own private data, and generates a noise message carrying the query identifier related to the query item according to the noise sending probability corresponding to the query item, which is determined based on the differential privacy algorithm; then all real messages and all noise messages for the policy query set are sent to the shuffler.
[0044] In this embodiment, since the user terminal only needs to answer the query items in the policy query set (rather than all the queries in the original query set), the number of messages generated and sent by the user is greatly reduced, which has obvious advantages in communication cost and storage, especially in a big data environment, and can significantly reduce the communication burden of each user. In addition, each query item needs to add noise to meet the differential privacy requirements. Answering fewer query items in the policy query set means that the overall amount of noise added is reduced, thereby reducing the impact of noise accumulation on the accuracy of the query results. Even if a linear combination is used to restore the original query results, the noise accumulation effect is relatively small, thereby improving the accuracy of the final query.
[0045] Among them, for any noise message corresponding to a query item, the user terminal follows the Bernoulli distribution with the noise sending probability corresponding to the query item as a parameter: if and only if the result of the Bernoulli test is 1, the noise message corresponding to the query item is generated. Bernoulli distribution is a discrete probability distribution, and its random variable takes only two values: usually 1 represents success and 0 represents failure. Its probability distribution is determined by the parameter p (that is, the noise sending probability), that is, the probability of taking 1 is p, and the probability of taking 0 is 1-p. In this embodiment, the Bernoulli distribution is used to randomly determine whether to generate a noise message, ensuring that the noise added by each user terminal when answering is random and independent, which helps to achieve differential privacy protection and prevent attackers from inferring the user's real data through deterministic models. Because the generation of the noise message corresponding to each query item depends only on one Bernoulli test (a noise message is generated when the result is 1, otherwise it is not generated), the probability of generating a noise message is precisely controlled, which can reduce unnecessary message sending, thereby reducing the overall communication overhead.
[0046] The shuffler can perform anonymization shuffling on the real messages and noise messages received from all the users, obtain a message queue and send it to the analyzer.
[0047] After receiving the message queue, the analyzer classifies and counts the messages in the message queue based on the query identifiers carried by each message in the received message queue. The analyzer can classify each message in the message queue into the query item in the corresponding policy query set based on these query identifiers, thereby obtaining the number of messages for each query item in the policy query set. Since the user end adds noise messages according to the set noise sending probability in addition to the real answer when answering the query item, the number of messages of the query item obtained by the analyzer actually includes a mixed result of real data and noise data. In order to obtain the most accurate query results, the analyzer needs to denoise the number of messages for each query item based on the noise sending probability corresponding to each query item to obtain the query results for each query item.
[0048] Specifically, for each query item, the analyzer can estimate the estimated number of noise messages corresponding to the query item based on the noise sending probability corresponding to the query item and the total number of user terminals. For example, the estimated number of noise messages is the product of the noise sending probability corresponding to the query item and the total number of user terminals. Then, the analyzer uses the estimated number of noise messages corresponding to the query item to denoise the number of messages in the query item to obtain the query result of the query item, that is, the query result of the query item is the difference between the number of messages in the query item and the estimated number of noise messages corresponding to the query item.
[0049] In this embodiment, by determining the estimated number of noise messages and performing denoising, the analyzer can effectively reduce the impact of noise on the final query results, improve the accuracy of query statistics, and obtain more reliable data statistics while protecting privacy. Since the noise addition on the user side is random, the analyzer does not know which specific data is true, thereby ensuring that the user's privacy will not be directly leaked, and the denoising process is performed on the basis of overall statistics, which does not affect individual privacy protection.
[0050] In some embodiments, to ensure that the noise of each query item under shuffled differential privacy is properly controlled, the following steps can be used to determine the noise transmission probability corresponding to each query item: obtain the total privacy budget indicated by the differential privacy algorithm, which represents the upper limit of the privacy loss allowed during the entire query process; and determine the maximum number of private data held by the user end that meets the query item; illustratively, the policy query set is represented by a policy matrix, and a row in the policy matrix is used to describe a query item; the maximum number of private data held by the user end that meets the query item can be determined based on the L1 norm of the policy matrix. Specifically, the maximum value of the sum of the absolute values of the elements in each column of the policy matrix reflects the maximum number of times the user data is counted in all queries, that is, the maximum sensitivity. Then, based on the maximum number of matches obtained in the previous step, the total privacy budget is evenly distributed to each query item to obtain the privacy budget corresponding to each query item. This allocation method ensures that each query item meets the corresponding privacy requirements when answered separately. Next, based on the privacy budget corresponding to each query item and the total number of user terminals, the noise sending probability corresponding to the query item is calculated; wherein the noise sending probability corresponding to the query item is negatively correlated with the privacy budget corresponding to the query item and the total number of user terminals.
[0051] For example, suppose the total privacy budget indicated by the differential privacy algorithm is the parameter value of ε and the parameter value of δ. ε is the main privacy parameter in differential privacy, which is used to measure the strength of privacy protection; δ is another privacy parameter in differential privacy, which is used to allow a certain probability of privacy leakage. It is a non-negative real number, indicating that in rare cases, the privacy protection mechanism may not fully meet the requirements of differential privacy. The policy query set is represented by the policy matrix A. The values of the matrix elements in the policy matrix A are 0 / 1. Then the privacy budget corresponding to each query item is , , represents the L1 norm of the strategy matrix A. Assuming the total number of user terminals is n, and the noise transmission probability corresponding to each query item is p′, a similar formula can be used as , or other equivalent formulas, which are not restricted in this embodiment. At this time, since ε′ and δ′ are uniformly distributed, when the privacy budget allocated to each query item is large, the required amount of noise will be reduced, so that p′ is small; and when the total number of user terminals N is large, the proportion of noise contributed by a single user can also be reduced. Therefore, the noise transmission probability p′ is negatively correlated with the privacy budget corresponding to the query item and the total number of users, that is, the higher the privacy budget and the number of users, the lower the noise transmission probability.
[0052] The noise transmission probability is calculated based on public privacy parameters, the number of users, and the policy query set, and does not involve any user private data. In one possible implementation, the noise transmission probability corresponding to each query item can be calculated by the relevant equipment of the data service provider or the relevant equipment of the trusted third party, and then the noise transmission probability corresponding to the query item is distributed as a public parameter to the analyzer and all user terminals when the system is initialized. In another possible implementation, the noise transmission probability corresponding to the query item can be calculated by the analyzer, and then the noise transmission probability corresponding to the query item is distributed as a public parameter to all user terminals when the system is initialized.
[0053] In some embodiments, the strategic query set is represented by a strategic matrix (denoted as A), the original query set is represented by an original matrix (denoted as W), and a row in the strategic matrix and the original matrix is used to describe a query item; the query results of all query items in the strategic query set are represented by a first result matrix (denoted as ) indicates that the rows in the first result matrix correspond to the rows in the strategy matrix one by one. Since there is a predetermined mapping relationship between the original query set W and the strategy query set A, it can be considered that the original query results can be restored by linear combination of the strategy query results. Specifically, the inverse matrix of the strategy matrix A can be constructed (or if A is not invertible, then the pseudo-inverse is usually used to achieve a similar effect).
[0054] After obtaining the first result matrix based on the above steps, the analyzer can transform the first result matrix based on the original matrix and the inverse matrix of the strategy matrix to obtain a second result matrix (denoted as ), different rows in the second result matrix represent the query results of different query items in the original query set. Specifically, the second result matrix is the product of the original matrix, the inverse matrix of the strategy matrix, and the first result matrix, that is, By using the inverse matrix of A, the query results in the strategy query set Restore the query results in the original query set through linear combination , ensuring that all original query items can still be accurately answered under differential privacy protection. In this way, the system can not only use the strategic query set to reduce the communication burden of each user end, but also restore the complete original query results through post-processing, thereby improving query efficiency and accuracy while ensuring differential privacy protection.
[0055] The following is an example of the data query based on shuffled differential privacy mentioned in the embodiments of this specification: (1) About the original query set and the strategic query set.
[0056] Assume that the original query set describes all possible continuous interval queries in the data domain [1,4], with a total of 10 query items. Represented by 0 / 1, each query item is represented by a 4-dimensional row vector, indicating which data values belong to the query. Specifically, the meaning of 0 / 1 in each dimension of the 4-dimensional row vector is as follows: the first dimension indicates whether the data value 1 is included in the query, 1 indicates included, 0 indicates not included; the second dimension indicates whether the data value 2 is included in the query, 1 indicates included, 0 indicates not included; the third dimension indicates whether the data value 3 is included in the query, 1 indicates included, 0 indicates not included; the fourth dimension indicates whether the data value 4 is included in the query, 1 indicates included, 0 indicates not included. For example, the original query set can be represented as the following 10*4 matrix: ; Each row corresponds to a query item, the first row corresponds to the query [1,1]; the second row corresponds to the query [1,2]; the third row corresponds to the query [1,3]; the fourth row corresponds to the query [1,4]; the fifth row corresponds to the query [2,2]; the sixth row corresponds to the query [2,3]; the seventh row corresponds to the query [2,4]; the eighth row corresponds to the query [3,3]; the ninth row corresponds to the query [3,4]; the tenth row corresponds to the query [4,4].
[0057] In order to reduce the user communication burden and noise accumulation, this embodiment uses a binary tree method to optimize the original query set to generate a smaller policy query set. The policy query set can be represented as the following 10*4 matrix: ; The first row corresponds to the query [1,4]; the second row corresponds to the query [1,2]; the third row corresponds to the query [3,4]; the fourth row corresponds to the query [1,1]; the fifth row corresponds to the query [2,2]; the sixth row corresponds to the query [3,3]; the seventh row corresponds to the query [4,4].
[0058] (2) Local calculation process on the user side.
[0059] Assume that the private data held by the user is a specific value. Taking data value 2 as an example, the user first converts its data into a one-hot vector. The dimension of the one-hot vector is determined based on the size of the data domain, that is, the dimension of the one-hot vector is 4. For the data value 2 held by the user, the one-hot vector (denoted as x) is: Then, the user side calculates the local response, that is, multiplying the one-hot vector by the strategy matrix A to obtain the answer vector (denoted as y), y=x*A, which means that ; The answer vector corresponds to the strategy matrix A one-to-one. For example, the first row in the answer vector is "1", which means that it meets the first query item in the strategy matrix A, that is, it is in the interval [1,4].
[0060] After the user calculates y locally, it will generate a real message with the corresponding query identifier for each query item: if a position in y is 1, the user will generate a real message with the corresponding query identifier; at the same time, each query item will also decide whether to generate a noise message through a Bernoulli test (i.e., it follows the Bernoulli(p′) distribution) according to the preset noise sending probability (parameters determined by the differential privacy algorithm). Finally, the user sends all real and noise messages to the shuffler.
[0061] (3) After receiving messages from all users, the shuffler randomly shuffles (anonymizes) these messages to form a message queue and sends the queue to the analyzer.
[0062] (4) The analyzer classifies and counts the messages based on the query identifier carried by each message in the message queue, and obtains the number of messages for each query item in the policy query set. Subsequently, combined with the noise sending probability and the total number of users, the analyzer can calculate the estimated number of noise messages corresponding to each query item. In order to obtain an unbiased estimate, the analyzer denoises the number of messages for each query item and obtains the denoised number of messages for each query item in the policy query set, which is the query result for each query item in the policy query set. Finally, according to the pre-established mapping relationship between the original query set and the policy query set, the query results in the policy query set are converted into query results in the original query set. The analyzer finally outputs the query results of each query item in the original query set, which are close to the true value under privacy guarantee.
[0063] The various technical features in the above embodiments can be arbitrarily combined as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, any combination of the various technical features in the above embodiments also falls within the scope of this specification.
[0064] In some embodiments, see Figure 5 This specification also provides a data query method based on shuffled differential privacy, which is applied to the user end of the above data query system; the method includes: In S501, a policy query set is obtained, which is a proper subset of the original query set. The original query set is used to describe all interval queries that need to be answered in the data domain where private data held by the user end is located, and the query results of query items in the original query set except the policy query set can be obtained by linearly combining the query results of at least two query items in the policy query set.
[0065] In S502, each query item in the policy query set is answered based on the private data held by itself and the differential privacy algorithm, and an answer message of the policy query set is generated and sent to the shuffler in the data query system.
[0066] In some embodiments, different query items in the policy query set correspond to different query identifiers.
[0067] Based on the private data it holds and the differential privacy algorithm, it answers each query item in the policy query set and generates an answer message for the policy query set, including: for each query item in the policy query set, if it is determined to meet the query item based on the private data it holds, generating a real message carrying a query identifier related to the query item; and generating a noise message carrying a query identifier related to the query item according to the noise sending probability corresponding to the query item, the noise sending probability being determined based on the differential privacy algorithm; and sending all real messages and all noise messages for the policy query set to a shuffler.
[0068] In some embodiments, for any noise message corresponding to a query item, the user terminal follows a Bernoulli distribution with the noise sending probability corresponding to the query item as a parameter: if and only if the Bernoulli test result is 1, a noise message corresponding to the query item is generated.
[0069] The implementation process of the above method is detailed in the description process of the corresponding user end in the above system, which will not be repeated here.
[0070] In some embodiments, see Figure 6 This specification also provides a data query method based on shuffled differential privacy, which is applied to the analyzer in the above data query system; the method includes: In S601, a message queue for a policy query set sent by a shuffler in a data query system is received; wherein the policy query set is a proper subset of the original query set, the original query set is used to describe all interval queries that need to be answered in the data domain where private data is located, and the query results of query items in the original query set except the policy query set can be obtained by linearly combining the query results of at least two query items in the policy query set.
[0071] In S602 , query results of all query items in the policy query set are determined based on the message queue and the differential privacy algorithm.
[0072] In S603, based on the mapping relationship between the original query set and the policy query set, the query results of all query items in the policy query set are converted into the query results of all query items in the original query set.
[0073] In some embodiments, different query items in the policy query set correspond to different query identifiers.
[0074] The query results of all query items in the policy query set are determined based on the message queue and the differential privacy algorithm, including: classifying and counting the messages in the message queue based on the query identifiers carried by each message in the message queue to obtain the number of messages for each query item in the policy query set; denoising the number of messages for each query item based on the noise sending probability corresponding to each query item to obtain the query results for each query item.
[0075] In some embodiments, the number of messages for each query item is denoised based on the noise sending probability corresponding to each query item to obtain query results for each query item, including: for each query item, based on the noise sending probability corresponding to the query item and the total number of user terminals, an estimated number of noise messages corresponding to the query item is estimated, and the estimated number of noise messages corresponding to the query item is denoised using the estimated number of noise messages corresponding to the query item to obtain the query result for the query item.
[0076] In some embodiments, the policy query set is represented by a policy matrix, the original query set is represented by an original matrix, and a row in the policy matrix and the original matrix is used to describe a query item; the query results of all query items in the policy query set are represented by a first result matrix, and the rows in the first result matrix correspond one-to-one to the rows in the policy matrix.
[0077] Based on the mapping relationship between the original query set and the strategy query set, the query results of all query items in the strategy query set are converted into the query results of all query items in the original query set, including: based on the original matrix and the inverse matrix of the strategy matrix, the first result matrix is converted to obtain a second result matrix, and different rows in the second result matrix represent the query results of different query items in the original query set.
[0078] In some embodiments, the second result matrix is the product of the original matrix, the inverse matrix of the strategy matrix, and the first result matrix.
[0079] The implementation process of the above method is detailed in the description process of the corresponding analyzer in the above system, which will not be repeated here.
[0080] In some embodiments, the embodiments of this specification also provide an electronic device, including: a processor; a memory for storing processor executable instructions; wherein the processor implements any of the above methods by running the executable instructions.
[0081] Figure 7 is a schematic structural diagram of a device provided by an exemplary embodiment. Figure 7At the hardware level, the device includes a processor 702, an internal bus 704, a network interface 706, a memory 708, and a non-volatile memory 710, and may also include hardware required for other functions. One or more embodiments of this specification may be implemented based on software, such as the processor 702 reading the corresponding computer program from the non-volatile memory 710 into the memory 708 and then running it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0082] In some embodiments, the data query device based on shuffle differential privacy can be applied to Figure 7 The device shown in the figure can realize the technical solution of this specification. The device can include: A policy query set acquisition module is used to acquire a policy query set, which is a proper subset of an original query set. The original query set is used to describe all interval queries that need to be answered in the data domain where private data is located, and the query results of query items in the original query set except for the policy query set can be obtained by linearly combining the query results of at least two query items in the policy query set.
[0083] The query answering module is used to answer each query item in the policy query set based on its own private data and differential privacy algorithm, generate an answer message for the policy query set, and send it to the shuffler in the data query system.
[0084] In some embodiments, the data query device based on shuffle differential privacy can be applied to Figure 7 The device shown in the figure can realize the technical solution of this specification. The device can include: A message queue receiving module is used to receive a message queue for a policy query set sent by a shuffler in a data query system; wherein the policy query set is a proper subset of the original query set, the original query set is used to describe all interval queries that need to be answered in the data domain where private data is located, and the query results of query items in the original query set except the policy query set can be obtained by linearly combining the query results of at least two query items in the policy query set.
[0085] The query result determination module is used to determine the query results of all query items in the policy query set based on the message queue and the differential privacy algorithm.
[0086] The query result mapping module is used to convert the query results of all query items in the strategy query set into the query results of all query items in the original query set based on the mapping relationship between the original query set and the strategy query set.
[0087] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0088] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor executable instructions; wherein the processor implements the steps of the method described in any of the above embodiments by running the executable instructions.
[0089] Based on the same concept as the above method, this specification also provides a computer-readable storage medium on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.
[0090] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0091] Based on the same concept as the above method, this specification also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any of the above embodiments when executed by a processor.
[0092] The above description is merely a preferred embodiment of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.
Claims
1. A data query system based on shuffle differential privacy, including a shuffler, an analyzer, and a user terminal holding private data; The user end is used to obtain a policy query set, which is a proper subset of an original query set, and the original query set is used to describe all interval queries that need to be answered in the data domain where the private data is located, and the query results of the query items in the original query set except the policy query set can be obtained by linearly combining the query results of at least two query items in the policy query set; answer each query item in the policy query set based on the private data held by itself and the differential privacy algorithm, generate an answer message of the policy query set, and send it to the shuffler; The shuffler is used to anonymize and shuffle the reply messages sent by the user end to form a message queue, and send the message queue to the analyzer; The analyzer is used to determine the query results of all query items in the policy query set based on the message queue and the differential privacy algorithm; Based on the mapping relationship between the original query set and the policy query set, query results of all query items in the policy query set are converted into query results of all query items in the original query set.
2. According to the system of claim 1, different query items in the policy query set correspond to different query identifiers; The user end is specifically used for generating, for each query item in the policy query set, a real message carrying a query identifier related to the query item if it is determined based on the private data held by the user end that the query item meets the query item, and generating a noise message carrying a query identifier related to the query item according to the noise sending probability corresponding to the query item, wherein the noise sending probability is determined based on the differential privacy algorithm; and sending all real messages and all noise messages for the policy query set to the shuffler; The shuffler is specifically used to perform anonymization shuffling on the real messages and noise messages received from all the user terminals, obtain a message queue and send it to the analyzer.
3. According to the system of claim 2, for any noise message corresponding to a query item, the user terminal follows a Bernoulli distribution with the noise sending probability corresponding to the query item as a parameter: a noise message corresponding to the query item is generated if and only if the Bernoulli test result is 1.
4. According to the system described in claim 2, the analyzer is specifically used to classify and count the messages in the message queue based on the query identifier carried by each message received in the message queue, so as to obtain the number of messages for each query item in the policy query set; and to denoise the number of messages for each query item based on the noise sending probability corresponding to each query item, so as to obtain the query results for each query item.
5. According to the system of claim 4, the analyzer is specifically used to determine, for each of the query items, the estimated number of noise messages corresponding to the query item based on the noise sending probability corresponding to the query item and the total number of user terminals, and use the estimated number of noise messages corresponding to the query item to denoise the number of messages in the query item to obtain the query result of the query item.
6. According to the system of any one of claims 2 to 5, the noise sending probability corresponding to the query item is determined by: Obtaining a total privacy budget indicated by the differential privacy algorithm, and determining a maximum number of private data held by the user terminal that match the query item; Based on the maximum matching number, the total privacy budget is evenly distributed to each query item to obtain a privacy budget corresponding to each query item; Based on the privacy budget corresponding to each query item and the total number of user terminals, the noise transmission probability corresponding to the query item is calculated; wherein, The noise sending probability corresponding to the query item is negatively correlated with the privacy budget corresponding to the query item and the total number of user terminals.
7. The system according to claim 6, wherein the policy query set is represented by a policy matrix, and a row in the policy matrix is used to describe a query item; in, The maximum number of private data held by the user terminal that matches the query item is determined based on the L1 norm of the policy matrix.
8. The system according to claim 1, wherein the policy query set is generated by: The data domain involved in the original query set is used as the root node of a binary tree, and each parent node of the binary tree is split into two child nodes until the interval length of the child node is 1 and the splitting is stopped; wherein, The interval of any parent node is the merged result of the intervals of the two child nodes of the parent node; Based on all nodes in the binary tree, query items in the strategic query set are determined.
9. The system according to claim 1, wherein the policy query set is represented by a policy matrix, the original query set is represented by an original matrix, and a row in the policy matrix and the original matrix is used to describe a query item; The query results of all query items in the strategy query set are represented by a first result matrix, and the rows in the first result matrix correspond one-to-one to the rows in the strategy matrix; The analyzer is specifically used to transform the first result matrix based on the original matrix and the inverse matrix of the strategy matrix to obtain a second result matrix, and different rows in the second result matrix represent query results of different query items in the original query set. 10 . The system according to claim 9 , wherein the second result matrix is the product of the original matrix, the inverse matrix of the strategy matrix, and the first result matrix.
11. A data query method based on shuffled differential privacy, applied to a user end in a data query system; the method comprises: Obtaining a policy query set, wherein the policy query set is a proper subset of an original query set, the original query set is used to describe all interval queries that need to be answered in the data domain where the private data held by the user terminal is located, and the query results of the query items in the original query set other than the policy query set can be obtained by linearly combining the query results of at least two query items in the policy query set; Based on the private data held by itself and the differential privacy algorithm, each query item in the policy query set is answered, and an answer message of the policy query set is generated and sent to the shuffler in the data query system.
12. The method according to claim 11, wherein different query items in the policy query set correspond to different query identifiers; The answering of each query item in the policy query set based on the private data held by the user and the differential privacy algorithm to generate an answer message of the policy query set includes: For each query item in the policy query set, if it is determined based on the private data held by itself that the query item is met, a real message carrying a query identifier related to the query item is generated; as well as Generate a noise message carrying a query identifier related to the query item according to a noise sending probability corresponding to the query item, wherein the noise sending probability is determined based on the differential privacy algorithm; All true messages and all noise messages for the policy query set are sent to the shuffler.
13. According to the method of claim 12, for any noise message corresponding to a query item, the user terminal follows a Bernoulli distribution with the noise sending probability corresponding to the query item as a parameter: if and only if the Bernoulli test result is 1, a noise message corresponding to the query item is generated.
14. A data query method based on shuffle differential privacy, applied to an analyzer in a data query system; the method comprises: Receive a message queue for a strategy query set sent by a shuffler in the data query system; wherein the strategy query set is a proper subset of an original query set, the original query set is used to describe all interval queries that need to be answered in the data domain where the private data is located, and the query results of the query items in the original query set other than the strategy query set can be obtained by linearly combining the query results of at least two query items in the strategy query set; Determine query results for all query items in the policy query set based on the message queue and the differential privacy algorithm; Based on the mapping relationship between the original query set and the policy query set, query results of all query items in the policy query set are converted into query results of all query items in the original query set.
15. The method according to claim 14, wherein different query items in the policy query set correspond to different query identifiers; The determining the query results of all query items in the policy query set based on the message queue and the differential privacy algorithm includes: Based on the query identifiers carried by the messages in the message queue, the messages in the message queue are classified and counted to obtain the number of messages for each query item in the policy query set; The number of messages for each query item is denoised based on the noise sending probability corresponding to each query item to obtain a query result for each query item.
16. The method according to claim 15, wherein the denoising process is performed on the number of messages of each query item based on the noise transmission probability corresponding to each query item to obtain the query result of each query item, comprising: For each of the query items, based on the noise sending probability corresponding to the query item and the total number of user terminals, the estimated number of noise messages corresponding to the query item is estimated, and the estimated number of noise messages corresponding to the query item is used to denoise the number of messages in the query item to obtain the query result of the query item.
17. The method according to claim 14, wherein the policy query set is represented by a policy matrix, the original query set is represented by an original matrix, and a row in the policy matrix and the original matrix is used to describe a query item; The query results of all query items in the strategy query set are represented by a first result matrix, and the rows in the first result matrix correspond one-to-one to the rows in the strategy matrix; The converting, based on the mapping relationship between the original query set and the policy query set, the query results of all query items in the policy query set into the query results of all query items in the original query set includes: Based on the original matrix and the inverse matrix of the strategy matrix, the first result matrix is transformed to obtain a second result matrix, and different rows in the second result matrix represent query results of different query items in the original query set.
18. The method according to claim 17, wherein the second result matrix is the product of the original matrix, the inverse matrix of the strategy matrix, and the first result matrix.
19. An electronic device comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as claimed in any one of claims 11 to 18 by running the executable instructions.
20. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method as claimed in any one of claims 11 to 18.
21. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 11 to 18.
Citation Information
Patent Citations
Computer-implemented privacy engineering system and method
CN109716345A
Data privacy protection method and device based on random noise and medium
CN118350042A
Computer-implemented privacy engineering system and method
US20200327252A1