Differential privacy protection method and device, storage medium and computer equipment

By calculating the information entropy of attributes and normalized mutual information, dividing attribute groups, and calculating the optimal privacy budget based on the sensitivity coefficient, the problem of low data utility in traditional differential privacy technology is solved, and the optimal balance between privacy budget and data utility is achieved.

CN120124096APending Publication Date: 2025-06-10HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510149483.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

While protecting data privacy, traditional differential privacy technology leads to reduced data utility, and it is impossible to effectively weigh privacy protection and data utility.

Method used

By calculating the information entropy and normalized mutual information of each attribute, dividing the attribute grouping, and calculating the optimal privacy budget based on the direct and indirect sensitivity coefficients within the packet, adding the same level of Laplace mechanism noise to the packet.

Benefits of technology

It realizes the optimal balance between privacy while protecting data, reducing unnecessary noise addition, improving data utility, and achieving the best balance between privacy budget and data utility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124096A_ABST
    Figure CN120124096A_ABST
Patent Text Reader

Abstract

The invention relates to a differential privacy protection method and device, a storage medium and computer equipment, and the method comprises the steps: marking sensitive attributes for an original data set, distributing direct sensitive coefficients, and generating a corresponding attribute set; calculating the information entropy of each attribute and the normalized mutual information between every two attributes according to the attribute set, and calculating the indirect sensitivity coefficient of each attribute in other attribute sets about the sensitive attribute set according to the normalized mutual information; dividing the original data set into groups according to normalized mutual information between every two attributes to obtain corresponding groups; and on the basis of the direct sensitivity coefficient, the indirect sensitivity coefficient and the information entropy, counting the average normalized information entropy and the maximum sensitivity coefficient of the attributes in the groups, and obtaining the optimal privacy budget of the corresponding groups according to the average normalized information entropy, the maximum sensitivity coefficient and a preset adjustment factor. And according to the optimal privacy budget of each group, Laplacian mechanism noise with the same level is added into the group, so that the effectiveness of the data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a differential privacy protection method, apparatus, storage medium, and computer device. Background Art

[0002] With the rapid development of the Internet of Things technology, various intelligent devices have continuously penetrated into people's daily lives. However, in complex real-world application scenarios, the data collected by Internet of Things devices is huge in volume and diverse in type, inevitably accompanied by potential risks such as data leakage and privacy infringement. The problem of privacy security will weaken users' trust. Therefore, more and more researchers have begun to devote themselves to researching privacy protection technologies. Among them, differential privacy technology is a technology that protects individual privacy by adding an appropriate amount of noise during data publishing or querying. Its core idea is to ensure that even if the individual data in a dataset is modified, no specific information about any individual can be inferred from the change in the query result.

[0003] The degree of differential privacy protection is determined by the privacy budget. A smaller privacy budget represents stronger privacy protection, but it will also add more noise, resulting in the utility of the data being affected. Therefore, how to balance privacy protection and utility has become an important issue that needs to be considered in differential privacy technology. Traditional differential privacy technologies usually add the same level of noise to all attributes of the data, resulting in some low-sensitivity attributes adding too much noise and thus losing unnecessary data utility. Another traditional method is to add different levels of noise to all attributes separately, both of which destroy the correlation between the data attributes after noise addition, and the privacy of potentially sensitive attributes is not necessarily guaranteed, having the disadvantage of low data utility. Summary of the Invention

[0004] Based on this, in view of the above problems, it is necessary to provide a differential privacy protection method, apparatus, storage medium, and computer device that can improve the data utility.

[0005] The first aspect of this application provides a differential privacy protection method, including:

[0006] Obtain the original data set to which noise is to be added, mark the sensitive attributes of the original data set and assign direct sensitivity coefficients to generate a corresponding attribute set; the attribute set includes a sensitive attribute set and other attribute sets;

[0007] Calculate the information entropy of each attribute and the normalized mutual information between all pairs of attributes according to the attribute set, and calculate the indirect sensitivity coefficient of each attribute in the other attribute set with respect to the sensitive attribute set according to the normalized mutual information;

[0008] Divide the original dataset into groups according to the normalized mutual information between pairs of attributes, and obtain the corresponding groups;

[0009] Based on the direct sensitivity coefficient, the indirect sensitivity coefficient, and the information entropy, calculate the average normalized information entropy and the maximum sensitivity coefficient of the attributes within the group. According to the average normalized information entropy, the maximum sensitivity coefficient, and a preset adjustment factor, obtain the optimal privacy budget for the corresponding group;

[0010] According to the optimal privacy budget of each group, add Laplace mechanism noise at the same level to the group, and merge the groups with added noise to obtain a noisy dataset and output it.

[0011] In one embodiment, calculating the information entropy of each attribute and the normalized mutual information between all pairs of attributes according to the attribute set includes:

[0012]

[0013] Among them, H(s k ) is the information entropy of attribute s k , p(x k,l ) represents the probability that the value of attribute s k is x k,l , L is the number of median values in the value range of attribute s k ; I(s k ; s p ) is the mutual information between attribute s k and attribute s p , p(x k,l , x p,q ) is the joint probability distribution of attribute s k and attribute s p , indicating the probability that the attribute pair (s k , s p ) takes the value (x k,l , x p,q ), p(x p,q ) represents the probability that the value of attribute s p is x p,q , Q is the number of median values in the value range of attribute s p ; NMI(s k , s p ) is the normalized mutual information between attribute s k and attribute s p , H(s p ) is the information entropy of attribute s p .

[0014] In one embodiment, calculating the indirect sensitivity coefficient of each attribute in the other attribute set with respect to the sensitive attribute set according to the normalized mutual information includes:

[0015] m a =max(m b ×NMI(s a ,s b ))

[0016] where m a is the indirect sensitivity coefficient of attribute s a in the other attribute set, m b is the direct sensitivity coefficient of attribute s b in the sensitive attribute set, max represents taking the maximum value of the product of the direct sensitivity coefficients of each attribute in the sensitive attribute set and the normalized mutual information between each attribute in the sensitive attribute set and attribute s a .

[0017] In one embodiment, partitioning the original data set into corresponding groups according to the normalized mutual information between pairs of attributes includes:

[0018] Select an attribute to create a new group;

[0019] Traverse the remaining attributes, and include the attributes whose normalized mutual information with each attribute in the newly created group is greater than a given threshold into the newly created group until no remaining attributes can be included in the newly created group;

[0020] Select an attribute from the remaining attributes to create a new group, and repeat the step of traversing the remaining attributes and including the attributes whose normalized mutual information with each attribute in the newly created group is greater than a given threshold into the newly created group until no remaining attributes can be included in the newly created group until all attributes are included in the corresponding groups.

[0021] In one embodiment, based on the direct sensitivity coefficient, the indirect sensitivity coefficient, and the information entropy, statistically calculating the average normalized information entropy and the maximum sensitivity coefficient of the attributes within the group includes:

[0022] Calculate the normalized information entropy of each attribute within the group, and the formula is as follows:

[0023]

[0024] where H norm (s) is the normalized information entropy of attribute s in the group, H(s) is the information entropy of attribute s in the group, H max (s) is the maximum information entropy in the group, and L represents the number of values in the value range of attribute s;

[0025] Calculate the average normalized information entropy of the attributes within a group, and the formula is as follows:

[0026]

[0027] Among them, H avg-norm (G i ) is the average normalized information entropy of group G i , s i,r represents the r-th attribute of group G i , R is the number of attributes within group G i , 1 ≤ r ≤ R, 1 ≤ i ≤ g, and g is the number of groups;

[0028] Statistically calculate the maximum sensitivity coefficient within a group, and the formula is as follows:

[0029]

[0030] Among them, m' i is the maximum sensitivity coefficient within group G i , s r is the attribute of group G i , m Sr is the direct sensitivity coefficient or indirect sensitivity coefficient of attribute s r .

[0031] In one embodiment, the obtaining of the optimal privacy budget for the corresponding group according to the average normalized information entropy, the maximum sensitivity coefficient, and a preset adjustment factor includes:

[0032] Calculate the privacy budget weight of the group, and the formula is as follows:

[0033] ω i = α(1 - m′ i ) + (1 - α)H avg-norm (G i )

[0034] Among them, ω i is the privacy budget weight of group G i , α is the adjustment factor, and 0 ≤ α ≤ 1;

[0035] Calculate the optimal privacy budget of the group, and the formula is as follows:

[0036]

[0037] Among them, ε i is the optimal privacy budget of group G i , ε is the given total privacy budget, and g is the number of groups.

[0038] In one embodiment, adding Laplace mechanism noise at the same level to a group according to the optimal privacy budget of each group includes:

[0039]

[0040] where ε i is the privacy budget of group G i , Ds r represents the subset of the original data set corresponding to attribute s i within group G r , M(Ds r ) is the subset of data after adding noise to attribute s i within group G r , f represents the query function, Δf is the sensitivity of the query function f corresponding to attribute s r , and Lap represents the Laplace function.

[0041] The second aspect of the present application provides a differential privacy protection device, including:

[0042] A data acquisition module, configured to acquire the original data set to which noise is to be added, mark sensitive attributes for the original data set and assign direct sensitivity coefficients, and generate a corresponding attribute set; the attribute set includes a sensitive attribute set and other attribute sets;

[0043] A data calculation module, configured to calculate the information entropy of each attribute according to the attribute set, and the normalized mutual information between all pairs of attributes, and calculate the indirect sensitivity coefficient of each attribute in the other attribute set with respect to the sensitive attribute set according to the normalized mutual information;

[0044] An attribute grouping module, configured to divide the original data set into groups according to the normalized mutual information between pairs of attributes to obtain corresponding groups;

[0045] A grouping processing module, configured to statistically calculate the average normalized information entropy and the maximum sensitivity coefficient of the attributes within the group based on the direct sensitivity coefficient, the indirect sensitivity coefficient, and the information entropy, and obtain the optimal privacy budget of the corresponding group according to the average normalized information entropy, the maximum sensitivity coefficient, and a preset adjustment factor;

[0046] A grouping noise adding module, configured to add Laplace mechanism noise at the same level to the group according to the optimal privacy budget of each group, merge the groups to which noise has been added respectively to obtain a noise-added data set, and output it.

[0047] The third aspect of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0048] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0049] For the above differential privacy protection method, device, storage medium and computer device, sensitive attributes are marked for the original data set and direct sensitivity coefficients are assigned to generate corresponding attribute sets; the information entropy of each attribute and the normalized mutual information between all pairs of attributes are calculated according to the attribute sets, and the indirect sensitivity coefficients of each attribute in other attribute sets with respect to the sensitive attribute set are calculated according to the normalized mutual information; the original data set is grouped according to the normalized mutual information between pairs of attributes to obtain corresponding groups; based on the direct sensitivity coefficients, indirect sensitivity coefficients and information entropy, the average normalized information entropy and the maximum sensitivity coefficient of the attributes within the group are statistically calculated, and according to the average normalized information entropy, the maximum sensitivity coefficient and a preset adjustment factor, the optimal privacy budget for the corresponding group is obtained, realizing the balance between information entropy and sensitivity coefficient, and can more flexibly combine the actual usage situation to generate the optimal privacy budget. According to the optimal privacy budget of each group, Laplace mechanism noise at the same level is added to the group, and the groups with added noise are merged to obtain a noisy data set and output. It can protect the correlation between attributes, reduce unnecessary noise addition, improve the utility of data, and achieve the optimal balance between privacy budget and data utility. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is an application environment diagram of the differential privacy protection method in an embodiment;

[0051] Figure 2 It is a flowchart of the differential privacy protection method in an embodiment;

[0052] Figure 3 It is an example diagram of the process for obtaining attribute sensitivity coefficients in an embodiment;

[0053] Figure 4 It is an example diagram of the group noise addition process in an embodiment;

[0054] Figure 5 It is a structural block diagram of the differential privacy protection device in an embodiment;

[0055] Figure 6 It is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0057] The differential privacy protection method provided by this application can be applied to, for example, Figure 1 the application environment shown as follows. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The terminal 102 obtains the original data set to which noise needs to be added and sends it to the server 104. The server 104 marks the sensitive attributes of the original data set and assigns direct sensitivity coefficients to generate the corresponding attribute set; the attribute set includes a sensitive attribute set and other attribute sets; calculates the information entropy of each attribute according to the attribute set, and the normalized mutual information between all pairs of attributes, and calculates the indirect sensitivity coefficient of each attribute in the other attribute set with respect to the sensitive attribute set according to the normalized mutual information; divides the original data set into groups according to the normalized mutual information between pairs of attributes to obtain the corresponding groups; based on the direct sensitivity coefficient, indirect sensitivity coefficient, and information entropy, statistically calculates the average normalized information entropy and the maximum sensitivity coefficient of the attributes within the group, and obtains the optimal privacy budget for the corresponding group according to the average normalized information entropy, the maximum sensitivity coefficient, and a preset adjustment factor; adds Laplace mechanism noise at the same level to the group according to the optimal privacy budget of each group, and merges the groups that have been added noise respectively to obtain the noisy data set. The server 104 can also output and feedback the noisy data set to the terminal 102 so that the terminal 102 can display the noisy data set to the user. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0058] In one embodiment, as Figure 2 shown, a differential privacy protection method is provided, including:

[0059] Step S110: Obtain the original data set to which noise needs to be added, mark the sensitive attributes of the original data set and assign direct sensitivity coefficients to generate the corresponding attribute set.

[0060] Specifically, the attribute set includes a sensitive attribute set and other attribute sets. The original data set D has n attributes, constituting the attribute set S = {s 1 , s 2 , s 3 , …, s i , …, s n}, initialize the sensitivity coefficient set M of the attributes = {m 1 , m 2 , m3 ,…,m i …,m n}, where one attribute corresponds to one sensitivity coefficient, m i = 0, 1 ≤ i ≤ n. Select m attributes from the attribute set S according to a set rule or randomly, mark them as sensitive attributes and put them into the sensitive attribute set S m , which is the sensitive attribute set S m in the sensitivity coefficients m corresponding to the attributes j are assigned to obtain the direct sensitivity coefficients, where 0 < m j < 1, 1 ≤ j ≤ n. The direct sensitivity coefficients for the sensitive attributes can be randomly assigned values between (0, 1), and m j corresponding to the attribute s j ∈ S m , The remaining attributes then form the other attribute set S t , S t ∪ S m = S.

[0061] Step S120: Calculate the information entropy of each attribute according to the attribute set, as well as the normalized mutual information between all pairs of attributes, and calculate the indirect sensitivity coefficients of each attribute in the other attribute set with respect to the sensitive attribute set. First, calculate the information entropy of each attribute, and then further calculate the normalized mutual information between all pairs of attributes as the basis for judging attribute correlation, and calculate the indirect sensitivity coefficients of each attribute in the other attribute set according to the correlation. By calculating the mutual information between data attributes, grouping attributes with high correlation together and adding the same level of noise to the attributes in the same group can ensure that the attributes in the same group have the same level of privacy protection. Compared with the current direct partitioning method with a given group size, it can protect the correlation between attributes and improve data utility.

[0062] Information entropy is used to measure the uncertainty of a system or the average amount of information. The greater the information entropy, the greater the uncertainty, and the lower the required level of privacy protection. For the k-th attribute s k , its value range is X k = {x k,1 , x k,2 , x k,3 , …, x k,l , …, x k,L}, where 1 ≤ l ≤ L, and L is the number of values in the value range of the attribute s k . The value range X k of the attribute s k is determined according to the subset of the original data set D corresponding to the attribute s k . The data with the same attribute in the original data set D is used as the subset of the original data set D corresponding to this attribute.

[0063] Mutual information is used to measure the amount of information shared between two random variables, that is, the correlation between data. For any two attributes s k and attribute s p in the attribute set S, the number of values in the corresponding value ranges are L and Q respectively. After calculating the mutual information of attribute s k and attribute s p according to the attribute set, the normalized mutual information of attribute s k and attribute s p can be further calculated.

[0064] Specifically, in step S120, the information entropy of each attribute and the normalized mutual information between all pairs of attributes are calculated according to the attribute set, including:

[0065]

[0066] Among them, H(s k ) is the information entropy of attribute s k , p(x k,l ) represents the probability that the value of attribute s k is x k,l , L is the number of values in the value range of attribute s k ; I(s k ; s p ) is the mutual information of attribute s k and attribute s p , p(x k,l , x p,q ) is the joint probability distribution of attribute s k and attribute s p , representing the probability that the attribute pair (s k , s p ) takes the value (x k,l , x p,q ), p(x p,q ) represents the probability that the value of attribute s p is x p,q , Q is the number of values in the value range of attribute s p ; NMI(s k , s p ) is the normalized mutual information of attribute s k and attribute s p , H(s p ) is the information entropy of attribute s p . The normalized mutual information can quantify the correlation between two attributes into the value range [0, 1].

[0067] Furthermore, for any attribute s a, whose indirect sensitivity coefficient can be expressed as attribute s a About the sensitive attribute set S m The maximum value of the direct sensitivity coefficient corresponding to each attribute in the sensitive attribute set and the product of the normalized mutual information between the two is the maximum correlation sensitivity. In step S120, the indirect sensitivity coefficient of each attribute in the other attribute set with respect to the sensitive attribute set is calculated according to the normalized mutual information, including:

[0068] m a =max(m b ×NMI(s a ,s b ))

[0069] Among them, m a Attribute s in other attribute sets a Indirect sensitivity coefficient, m b Attribute s is the sensitive attribute set b The direct sensitivity coefficient of max represents the direct sensitivity coefficient of each attribute in the sensitive attribute set, which is different from the direct sensitivity coefficient of each attribute in the sensitive attribute set and attribute s a The product of the normalized mutual information between takes the maximum value, s a ∈S t ,s b ∈S m .

[0070] Figure 3 The figure shows an example of the attribute sensitivity coefficient acquisition process. After selecting attribute marks from the original attribute set S to mark sensitive attributes and assigning direct sensitivity coefficients, the indirect sensitivity coefficients are calculated based on the correlation to determine the sensitivity coefficients corresponding to each attribute in other attribute sets. At this point, all attributes in the attribute set S have determined the relevant sensitivity coefficients, which will be used to subsequently determine the maximum sensitivity coefficient of each group.

[0071] Step S130: Divide the original data set into groups according to the normalized mutual information between the attributes to obtain corresponding groups. Each attribute has different privacy protection requirements. When the sensitivity number of the grouped data is low, it means that only a low level of noise needs to be added to improve the utility of the overall data. This method can add differentiated noise to groups with different privacy requirements, and unlike simply dividing sensitive attributes, it can more accurately determine the amount of noise added.

[0072] In one embodiment, step S130 includes: selecting an attribute to create a new group; traversing the remaining attributes, and classifying the attributes whose normalized mutual information with each attribute in the newly created group is greater than a given threshold into the newly created group until no remaining attribute can be classified into the newly created group; selecting an attribute from the remaining attributes to create a new group, and repeating the step of traversing the remaining attributes and classifying the attributes whose normalized mutual information with each attribute in the newly created group is greater than a given threshold into the newly created group until no remaining attribute can be classified into the newly created group, until all attributes are classified into corresponding groups. Specifically, the steps of grouping are as follows:

[0073] S31. First, classify all attributes into the attribute set S to be grouped n .

[0074] S32. Select the first attribute from the attribute set S to be grouped n to remove, create a new group G i and classify this attribute into G i , initially set i = 1, and when creating a new group later, i = i + 1.

[0075] S33. Traverse the attribute set S to be grouped n , and judge whether the NMI between the current attribute s c and each attribute in the newly created group G i is greater than the given threshold. When any NMI is greater than the given threshold, classify the attribute s c into the group G i .

[0076] S34. Repeat step S33 until no remaining attribute in the attribute set S to be grouped n can be classified into the group G i .

[0077] S35. Repeat steps S32 to S34 until the attribute set S to be grouped n is empty. After the attributes are classified, the subsets of the original data set D corresponding to each attribute are also classified into the corresponding groups together, and the original data set D is completely divided into groups.

[0078] Step S140: Based on the direct sensitivity coefficient, the indirect sensitivity coefficient, and the information entropy, calculate the average normalized information entropy and the maximum sensitivity coefficient of the attributes within the group, and obtain the optimal privacy budget for the corresponding group according to the average normalized information entropy, the maximum sensitivity coefficient, and a preset adjustment factor. Among them, the average information entropy and the maximum sensitivity coefficient of the attributes within the group are calculated according to information theory, and finally the adjustment factor is introduced to determine the privacy budget required for this group, realizing the balance between the information entropy and the sensitivity coefficient. Both the information entropy and the sensitivity coefficient can measure the degree of privacy protection. Compared with the current method for determining the privacy budget, it can more flexibly combine the actual usage situation to generate the optimal privacy budget.

[0079] Specifically, in step S140, based on the direct sensitivity coefficient, the indirect sensitivity coefficient, and the information entropy, the average normalized information entropy and the maximum sensitivity coefficient of the attributes within the group are statistically calculated, including:

[0080] Calculate the normalized information entropy of each attribute within the group. The formula is as follows:

[0081]

[0082] where H norm (s) is the normalized information entropy of attribute s in the group, H(s) is the information entropy of attribute s in the group, and H max (s) is the maximum information entropy in the group, and L represents the number of values in the value range of attribute s.

[0083] Calculate the average normalized information entropy of the attributes within the group. The formula is as follows:

[0084]

[0085] where H avg-norm (G i ) is the average normalized information entropy of group G i , s i,r represents the r-th attribute of group G i , R is the number of attributes within group G i , 1 ≤ r ≤ R, 1 ≤ i ≤ g, and g is the number of groups.

[0086] Statistically calculate the maximum sensitivity coefficient within the group. The formula is as follows:

[0087]

[0088] where m' i is the maximum sensitivity coefficient within group G i , s r is the attribute of group G i , m Sr is the direct sensitivity coefficient or the indirect sensitivity coefficient of attribute s r , and max represents taking the maximum value of the sensitivity coefficients corresponding to each attribute in group G i .

[0089] Furthermore, in step S140, based on the average normalized information entropy, the maximum sensitivity coefficient, and a preset adjustment factor, the optimal privacy budget for the corresponding group is obtained, including:

[0090] Calculate the privacy budget weight of the group. The formula is as follows:

[0091] ω i = α(1 - m′ i)+(1-α)H avg-norm (G i )

[0092] where ω i is the privacy budget weight of group G i , α is the adjustment factor, and 0 ≤ α ≤ 1.

[0093] Calculate the optimal privacy budget for each group, and the formula is as follows:

[0094]

[0095] where ε i is the optimal privacy budget of group G i , ε is the given total privacy budget, and g is the number of groups.

[0096] Step S150: According to the optimal privacy budget of each group, add Laplace mechanism noise at the same level to the group, and merge the groups with added noise respectively to obtain the noisy dataset and output it. Specifically, according to the optimal privacy budget of each group, adding Laplace mechanism noise at the same level to the group includes:

[0097]

[0098] In the above noise addition formula, ε i is the privacy budget of group G i , Ds r represents the subset of the original dataset corresponding to attribute s i in group G r , M(Ds r ) is the subset of the data after adding noise to attribute s i in group G r , f represents the query function, and Lap represents the Laplace function. Δf is the sensitivity of the query function f corresponding to attribute s r , that is, the maximum change in the query result, and its expression is as follows:

[0099]

[0100] where ||·|| is the Manhattan norm, Ds r and Ds r ' are two adjacent datasets that differ by only one element.

[0101] Figure 4 is an example of the group noise addition process. After grouping the relevant attributes for each attribute, add differential noise in units of groups. After adding noise to each group separately, merge all groups to obtain the final noisy dataset, and the noisy dataset can be output to the terminal and returned to the data requester, or output to the memory for storage.

[0102] The above-mentioned differential privacy protection method based on information theory can protect the correlation between attributes by calculating the mutual information between data attributes, grouping the attributes with high correlation into one group, and adding the same level of noise to the attributes in the same group. By introducing a sensitivity coefficient, directly marking and assigning the sensitivity coefficients of sensitive attributes in advance, and then calculating the indirect sensitivity coefficients of other attributes according to the attribute correlation, different attributes have different privacy protection requirements. When the sensitivity coefficient of a group is low, it means that only a low level of noise needs to be added, thereby improving the utility of the data. Different from the existing simple division of sensitive attributes, it can more accurately determine the amount of noise added. By using a personalized privacy budget determination method, calculating the average information entropy and the highest sensitivity coefficient of the attributes within the group according to information theory, and finally introducing an adjustment factor to determine the privacy budget required for this group, achieving a balance between information entropy and sensitivity coefficient, it can more flexibly combine the actual usage situation to generate an optimal privacy budget. The application of the above technology can protect the correlation between attributes, reduce unnecessary noise addition, improve the utility of the data, and achieve the optimal balance between privacy budget and data utility.

[0103] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0104] Based on the same inventive concept, the embodiments of the present application also provide a differential privacy protection device for implementing the above-mentioned differential privacy protection method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the differential privacy protection device provided below can refer to the limitations on the differential privacy protection method in the above text, and will not be repeated here.

[0105] In one embodiment, as Figure 5 shown, a differential privacy protection device is provided, including: a data acquisition module 110, a data calculation module 120, an attribute grouping module 130, a grouping processing module 140, and a grouping noise addition module 150, where:

[0106] The data acquisition module 110 is configured to acquire the original data set to which noise is to be added, mark the sensitive attributes of the original data set and assign direct sensitivity coefficients, and generate the corresponding attribute set; the attribute set includes a sensitive attribute set and other attribute sets.

[0107] The data calculation module 120 is configured to calculate the information entropy of each attribute according to the attribute set, and the normalized mutual information between every two attributes, and calculate the indirect sensitivity coefficient of each attribute in the other attribute set with respect to the sensitive attribute set according to the normalized mutual information.

[0108] The attribute grouping module 130 is configured to divide the original data set into groups according to the normalized mutual information between every two attributes to obtain the corresponding groups.

[0109] The grouping processing module 140 is configured to statistically calculate the average normalized information entropy and the maximum sensitivity coefficient of the attributes within the group based on the direct sensitivity coefficient, the indirect sensitivity coefficient, and the information entropy, and obtain the optimal privacy budget for the corresponding group according to the average normalized information entropy, the maximum sensitivity coefficient, and a preset adjustment factor.

[0110] The grouping noise addition module 150 is configured to add Laplace mechanism noise at the same level to the group according to the optimal privacy budget of each group, and merge the groups to which noise has been added respectively to obtain a noise-added data set and output it.

[0111] Each module in the above differential privacy protection device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form so that the processor can call and execute the operations corresponding to each of the above modules.

[0112] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a differential privacy protection method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0113] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0114] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the above method.

[0115] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps of the above method.

[0116] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps of the above method.

[0117] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The above computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the various embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0118] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0119] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A differential privacy protection method, characterized in that: include: Acquire an original data set to which noise is to be added, mark sensitive attributes of the original data set and assign direct sensitivity coefficients to generate a corresponding attribute set; the attribute set includes a sensitive attribute set and other attribute sets; Calculating the information entropy of each attribute and the normalized mutual information between all attributes according to the attribute set, and calculating the indirect sensitivity coefficient of each attribute in the other attribute set with respect to the sensitive attribute set according to the normalized mutual information; Dividing the original data set into groups according to normalized mutual information between attributes to obtain corresponding groups; Based on the direct sensitivity coefficient, the indirect sensitivity coefficient and the information entropy, the average normalized information entropy and the maximum sensitivity coefficient of the attributes in the group are counted, and according to the average normalized information entropy, the maximum sensitivity coefficient and the preset adjustment factor, the optimal privacy budget of the corresponding group is obtained; According to the optimal privacy budget of each group, the same level of Laplace mechanism noise is added to the group, and the groups to which noise has been added are merged to obtain the noisy data set and output.

2. The method according to claim 1, characterized in that The calculating the information entropy of each attribute and the normalized mutual information between all attributes according to the attribute set includes: Among them, H(s k ) is the attribute s k The information entropy of p(x k,l ) indicates attribute s k The value is x k,l The probability of L is attribute s k The number of values ​​in the range of I(s k ;s p ) is the attribute s k and attributes p The mutual information of k,l , x p,q ) is the attribute s k and attributes p The joint probability distribution of attribute pairs (s k ,s p ) takes the value of (x k,l , x p,q ), p( xp,q ) indicates attribute s p The value is x p,q The probability of Q is attribute s p The number of values ​​in the range of NMI(s k ,s p ) is the attribute s k and attributes p The normalized mutual information of p ) is the attribute s p Information entropy.

3. The method according to claim 2, characterized in that The calculating, according to the normalized mutual information, the indirect sensitivity coefficient of each attribute in the other attribute set with respect to the sensitive attribute set comprises: m a =max(m b ×NMI(s a ,s b )) Among them, m a Attribute s in other attribute sets a Indirect sensitivity coefficient, m b Attribute s is the sensitive attribute set b The direct sensitivity coefficient of max represents the direct sensitivity coefficient of each attribute in the sensitive attribute set, which is different from the direct sensitivity coefficient of each attribute in the sensitive attribute set and attribute s a The product of the normalized mutual information between them takes the maximum value.

4. The method according to claim 1, characterized in that: The method of dividing the original data set into groups according to the normalized mutual information between the attributes to obtain corresponding groups includes: Select an attribute to create a new group; Traverse the remaining attributes and assign the attributes whose normalized mutual information with each attribute in the newly created group is greater than a given threshold to the newly created group, until the remaining attributes cannot be assigned to the newly created group; Select an attribute from the remaining attributes to establish a new group, and repeat the traversal of the remaining attributes, and assign the attributes whose normalized mutual information with each attribute in the newly established group is greater than a given threshold to the newly established group, until the remaining attributes cannot be assigned to the newly established group, until all attributes are assigned to the corresponding groups.

5. The method according to claim 1, characterized in that The counting of the average normalized information entropy and the maximum sensitivity coefficient of the attributes in the group based on the direct sensitivity coefficient, the indirect sensitivity coefficient and the information entropy includes: Calculate the normalized information entropy of each attribute in the group. The formula is as follows: Among them, H norm (s) is the normalized information entropy of attribute s in the group, H(s) is the information entropy of attribute s in the group, and H max (s) is the maximum information entropy in the group, and L represents the number of values ​​in the value range of attribute s; Calculate the average normalized information entropy of the attributes within the group. The formula is as follows: Among them, H avg-norm (G i ) is group G i The average normalized information entropy of , s i,r Indicates group G i The rth attribute of i The number of internal attributes, 1≤r≤R, 1≤i≤g, g is the number of groups; The maximum sensitivity coefficient within the statistical group is calculated using the following formula: Among them, m′ i For group G i Maximum sensitivity coefficient within, s r For group G i The property of m Sr For attribute s r Direct sensitivity coefficient or indirect sensitivity coefficient.

6. The method according to claim 5, characterized in that The obtaining, according to the average normalized information entropy, the maximum sensitivity coefficient and the preset adjustment factor, an optimal privacy budget of the corresponding group includes: Calculate the privacy budget weight of the group using the following formula: oh i =α(1-m′ i )+(1-α)H avg-norm (G i ) Among them, ω i For group G i The privacy budget weight, α is the adjustment factor, 0≤α≤1; Calculate the optimal privacy budget for a group using the following formula: Among them, ε i For group G i The optimal privacy budget of , ε is the given total privacy budget, and g is the number of groups.

7. The method according to any one of claims 1 to 6, characterized in that: According to the optimal privacy budget of each group, adding the same level of Laplace mechanism noise to the group includes: Among them, ε i For group G i Privacy budget, Ds r Indicates group G i Internal attributes r The corresponding subset of the original data set, M(Ds r ) is group G i Internal attributes r The data subset after adding noise, f represents the query function, Δf is the attribute s corresponding to the query function f r sensitivity, Lap represents the Laplace function.

8. A differential privacy protection device, characterized in that: include: A data acquisition module is used to acquire an original data set to which noise is to be added, mark sensitive attributes of the original data set and assign direct sensitivity coefficients to generate a corresponding attribute set; the attribute set includes a sensitive attribute set and other attribute sets; A data calculation module, used to calculate the information entropy of each attribute and the normalized mutual information between all attributes according to the attribute set, and calculate the indirect sensitivity coefficient of each attribute in the other attribute set with respect to the sensitive attribute set according to the normalized mutual information; An attribute grouping module, used to divide the original data set into groups according to the normalized mutual information between the attributes, and obtain corresponding groups; A group processing module, used for counting the average normalized information entropy and the maximum sensitivity coefficient of the attributes in the group based on the direct sensitivity coefficient, the indirect sensitivity coefficient and the information entropy, and obtaining the optimal privacy budget of the corresponding group according to the average normalized information entropy, the maximum sensitivity coefficient and a preset adjustment factor; The group noise adding module is used to add the same level of Laplace mechanism noise to the group according to the optimal privacy budget of each group, merge the groups to which the noise has been added to obtain the noisy data set and output it.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.