A privacy-enhanced method for simulating and generating relational tabular data

The method addresses inefficiencies and privacy issues in multi-table data simulation by identifying and merging highly associated attributes, using Markov random fields with differential privacy to maintain data associations and privacy in complex relational networks.

CN119622822BActive Publication Date: 2025-07-15HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510161837.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-15
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

When generating multi-table data simulation, the prior art cannot effectively maintain the correlation between multiple tables, and lacks privacy enhancement protection, especially under the many-to-many foreign key relationship and there is a risk of privacy leakage.

Method used

By mining the correlation properties between tables, Markov random field model is used to combine Laplace, exponential and Gaussian mechanisms to generate simulated data, meet differential privacy requirements, retain the association relationship and provide privacy protection.

Benefits of technology

The effectiveness of multi-table simulation data is improved, ensuring that the simulated data remains authentic and effective while protecting privacy, and solving the risk of privacy leakage under the many-to-many foreign key relationship.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119622822B_ABST
    Figure CN119622822B_ABST
Patent Text Reader

Abstract

The present invention provides a method for simulating and generating privacy-enhanced relational tabular data, which mines highly correlated attributes in the linked table L and the single tables U and V with foreign key associations, and merges the attributes with the linked table L to obtain the attributes in the corresponding U and V tables for k attributes; according to the foreign key correspondence of the linked table L, the obtained attributes are spliced with the linked table L to obtain the merged table #imgabs0#. According to the attributes of the linked table L, the merged table #imgabs1# is split by columns to obtain the simulated and generated linked table #imgabs2#; according to the synthesis result of the linked table #imgabs3#, the Markov random field model is used to simulate and generate the table #imgabs4#; according to the synthesis result of the linked table #imgabs5#, the Markov random field model is used to simulate and generate the table #imgabs6#. When generating simulation data, the utility of the simulation data is improved, ensuring that the simulation data can still maintain its authenticity and effectiveness while protecting privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of tabular data processing, and particularly relates to a method for simulating and generating privacy-enhanced relational tabular data. Background Art

[0002] In various artificial intelligence applications, it is usually necessary to process and integrate multi-tables. Different from the traditional single-table data scenario, multi-table data often has foreign key associations. This requires that the simulated data can not only reflect the statistical characteristics of the original single-table data, but also maintain data consistency and relevance in a complex multi-table relationship network. Existing methods connect the parent table and the child table into a single table through foreign keys, and then use single-table algorithms for simulated data synthesis. However, the connection of irrelevant attributes not only reduces the running efficiency of the simulation algorithm, but also can only maintain the association relationship between the "U table and the L table" or the "V table and the L table", which has an adverse effect on the utility of the simulated data. It is not only inapplicable to the tabular data with "many-to-many" foreign key associations, but also does not apply privacy-enhanced technologies to protect the privacy of the model and the generated data. Summary of the Invention

[0003] To solve the above problems, the present invention provides a method for simulating and generating privacy-enhanced relational tabular data.

[0004] The present invention is implemented as follows. A method for simulating and generating privacy-enhanced relational tabular data, the method includes the following steps:

[0005] Step S1: Mine the highly relevant attributes in the link table L and the single tables U and V with foreign key associations, and merge the attributes with the link table L to obtain the attributes in the corresponding U and V tables for k attributes;

[0006] Step S2: According to the foreign key correspondence of the link table L, splice the attributes obtained in step S1 with the link table L to obtain a merged table , and according to the attributes of the link table L, split the merged table by columns to obtain the simulated and generated link table ;

[0007] Step S3: According to the synthesis result of the link table , use the Markov random field model to simulate and generate table ;

[0008] Step S4: According to the synthesis result of the link table , use the Markov random field model to simulate and generate table .

[0009] A further technical solution of the present invention is: The step S1 includes the following steps:

[0010] Step S11: Merge all attributes of Table U and Table V according to the foreign key matching relationship of Link Table L;

[0011] Step S12: Traverse the non-primary key attributes of Table U , and combine them with each non-primary key attribute of Table V to measure the correlation of each attribute pair using the relevance evaluation function;

[0012] Step S13: Apply the exponential mechanism: Use the relevance evaluation function as the utility function, and use as the output probability of the attribute pair , and randomly output the attribute pair k times from the set of attribute pairs in Step S12;

[0013] Step S14: Traverse the results of the k - time output, and respectively collect the corresponding attributes of Table U and Table V {U m |1 ≤ m ≤ k}, {V m |1 ≤ m ≤ k}.

[0014] A further technical solution of the present invention is that the said Step S2 includes the following steps:

[0015] Step S21: According to the foreign key correspondence of Link Table L, splice the attributes {U m |1 ≤ m ≤ k}, {V m |1 ≤ m ≤ k} obtained in Step S1 with Table L to obtain the merged table ;

[0016] Step S22: Use the Markov random field model to model the original Table U, and add Laplace noise during the modeling process to make it satisfy ε 21 - differential privacy to obtain the graph structure;

[0017] Step S23: According to the existing graph structure, use mirror gradient descent to iteratively learn the parameters of the Markov random field model, and apply the Gaussian mechanism during the process of learning the parameters to make it satisfy differential privacy;

[0018] Step S24: Calculate the gap between the marginal distribution of this model and the marginal distribution in the original data , and apply the Laplace mechanism to add noise to make it satisfy ε 23 - differential privacy, select the marginal distribution corresponding to the maximum value from the noisy version of the gap , and add the edges corresponding to the marginal distribution to the graph structure of the Markov random field, and perform the parameter learning of Step S23;

[0019] Step S25: Set the privacy budget for step S2 as , ε 21 as the privacy budget for step S21, ε 22 as the privacy budget for step S22, and ε 23 as the privacy budget for step S23. According to the serial mechanism of differential privacy, step S2 satisfies differential privacy;

[0020] Step S26: According to the attributes of the original link table L, split the merged table by columns to obtain the simulated generated link table .

[0021] A further technical solution of the present invention is: After obtaining the generation model in step S2, sample on the generation model to obtain the simulated data of the merged table . The algorithm has two termination conditions. One is that if the number of rows of the simulated data is to be maintained the same as that of the original table, the algorithm terminates when the algorithm generates simulated data with the same number of rows as the original link table L. The other is that if the foreign keys are to be maintained consistent, when the algorithm generates simulated data with the same number of rows as the original link table L, check whether the foreign keys of the original link table L exist in the simulated data table . If all exist, the algorithm terminates. If there are any missing, continue to sample using the generation model until all the foreign keys of the original link table L exist.

[0022] A further technical solution of the present invention is: If the tables U, V and the link table L have a one-to-many foreign key relationship, it is necessary to perform aggregation operations on the attributes related to the two tables respectively. If the attribute is a categorical attribute, sample and generate the categorical data according to the proportion of each category within the group after grouping by the foreign key ID; if the attribute is a numerical attribute, replace the simulated data with its corresponding average according to the foreign key ID. Since the sampling ratio and the average are both calculated based on the simulated data, according to the properties of differential privacy post-processing, no privacy budget is consumed.

[0023] A further technical solution of the present invention is: The said step S3 includes the following steps:

[0024] Step S31: Use the Markov random field model to model the original table U, and add Laplace noise during the modeling process to make it satisfy ε 31 -differential privacy to obtain the graph structure;

[0025] Step S32: According to the existing graph structure, use the mirror gradient descent to iteratively learn the parameters of the Markov random field model. The process of learning the parameters applies the Gaussian mechanism to make it satisfy differential privacy;

[0026] Step S33: Calculate the gap between the marginal distribution of this model and the marginal distribution in the original data , and apply the Laplace mechanism to add noise to satisfy ε 33 -differential privacy, and select the marginal distribution corresponding to the maximum value from the noisy version of the gap and add the edges corresponding to the marginal distribution to the graph structure of the Markov random field;

[0027] Step S34: To maintain the consistency of the simulation table and the simulation link table , first fill the U ID obtained in Step S2 and the related attributes of table U into and keep them unchanged during the sampling process;

[0028] Step S35: Given the Markov random field model and the attributes filled from the simulation link table , sample to generate the remaining attributes of table U;

[0029] Step S36: Let the privacy budget of Step S3 be , ε 31 be the privacy budget of Step S31, ε 32 be the privacy budget of Step S32, ε 33 be the privacy budget of Step S33. According to the serial mechanism of differential privacy, Step S3 satisfies differential privacy.

[0030] A further technical solution of the present invention is that the said Step S4 includes the following steps:

[0031] Step S41: Use the Markov random field model to model the original table V, and add Laplace noise during the modeling process to satisfy ε 41 -differential privacy to obtain the graph structure;

[0032] Step S42: According to the existing graph structure, use mirror gradient descent to iteratively learn the parameters of the Markov random field model, and apply the Gaussian mechanism during the process of learning the parameters to satisfy differential privacy;

[0033] Step S43: Calculate the gap between the marginal distribution of this model and the marginal distribution in the original data , and apply the Laplace mechanism to add noise to satisfy ε 43 -differential privacy, and select the marginal distribution corresponding to the maximum value from the noisy version of the gap and add the edges corresponding to the marginal distribution to the graph structure of the Markov random field;

[0034] Step S44: To maintain the consistency of the simulation table and the simulation link table , first fill the V IDAnd fill the relevant attributes of Table V into and keep them unchanged during the sampling process;

[0035] Step S45: Under the given Markov random field model and the attributes filled from the simulation link table sample to generate the remaining attributes of Table V;

[0036] Step S46: Let the privacy budget of Step S4 be , ε 41 be the privacy budget of Step S41, ε 42 be the privacy budget of Step S42, and ε 43 be the privacy budget of Step S43. According to the serial mechanism of differential privacy, Step S4 satisfies differential privacy.

[0037] A further technical solution of the present invention is: It further includes Step S5: Based on the algorithm flow of Steps S1 to S4, according to the serial mechanism of differential privacy, the overall algorithm satisfies differential privacy.

[0038] The beneficial effects of the present invention are: By analyzing the associated attributes between multiple tables with foreign key relationships, select the attributes with significant association relationships to better maintain the association between the sub-table and multiple parent tables during the generation of simulation data, thereby improving the utility of the simulation data. The differential privacy mechanism is integrated, including the exponential mechanism, the Laplace mechanism, and the Gaussian mechanism, to provide privacy protection for the generation model and simulation data during the process of data simulation generation, prevent information leakage, and ensure that the simulation data can still maintain its authenticity and effectiveness while protecting privacy. Brief Description of the Drawings

[0039] Figure 1 is the flow chart of the method of the present invention;

[0040] Figure 2 is an example of relational table data with many-to-many foreign keys of the present invention;

[0041] Figure 3 is the multi-table data simulation generation process of the present invention for many-to-many foreign keys;

[0042] Figure 4 is the process of simulating and generating the link table of the present invention ;

[0043] Figure 5 is the process of simulating and generating Table U of the present invention;

[0044] Figure 6 is the process of simulating and generating Table V of the present invention. Detailed Embodiments

[0045] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0046] The present invention provides a method for generating simulated relational tabular data with enhanced privacy. For relational tabular data with many-to-many foreign keys, the method can preserve the association of attributes between multiple tables under foreign keys when generating simulated data, overcoming the limitations of traditional methods that are only applicable to "one-to-one" and "one-to-many".

[0047] Terminology description:

[0048] Differential privacy: Differential privacy is a privacy-enhancing technology with a mathematical theoretical basis. Adding or deleting any single data record will not significantly affect the query results. Therefore, even if an attacker knows all records in a dataset except one, the leakage of that record will not occur. The relevant definitions related to this patent are as follows:

[0049] Definition 1, ε-differential privacy: For any two datasets D1 and D2 with the same data structure, where D1 and D2 differ by only one record (adjacent datasets). Let there be a random algorithm M, and the set of all possible output results of this algorithm is O. For any output result on datasets D1 and D2, if the random algorithm M satisfies:

[0050] where Pr[*] represents the probability of an event occurring, and the parameter ε represents the privacy protection budget, then the algorithm M satisfies ε-differential privacy protection.

[0051] Definition 2, Sensitivity: Given a query function , D is the input dataset, and R d is the output dataset. For any pair of D1 and D2, the sensitivity of the function f is:

[0052] where is and the first-order distance between. Sensitivity measures the impact on the corresponding output results when the input dataset changes. Based on sensitivity, we can give the implementation mechanism of differential privacy.

[0053] Definition 3, Laplace mechanism: Given a dataset D and a privacy budget ε, the sensitivity of the function f is , when the output of f satisfies:

[0054] then the algorithm A is said to satisfy ε-differential privacy, where is random noise that satisfies the Laplace distribution.

[0055] Definition 4, Exponential Mechanism: Given a dataset D and a privacy budget ε, the output of the randomized algorithm M is an entity object r ∈ R, μ(D, R) is the utility function, is the sensitivity of the function μ(D, R). If r is selected and output from the input with a probability proportional to , then the algorithm M satisfies ε-differential privacy.

[0056] Definition 5, Serial Composition Mechanism: For sub-algorithms that satisfy differential privacy respectively, and the processes of any two sub-algorithms are independent of each other, then the overall output of this algorithm satisfies differential privacy.

[0057] Foreign Key: A foreign key in a database is a field that references the primary key in another table and is used to establish and represent the relationship between tables. The main relationship types of foreign keys are as follows: one-to-one, one-to-many, and many-to-many relationships.

[0058] To facilitate the description of the algorithms involved in the present invention, for the relational table data of many-to-many foreign keys, a triple B = (U, V, L) is used to represent the results that need to be modeled by the generative model. Among them, the elements U and V respectively represent the table data generated by the two foreign keys. The element L, as the representation of the many-to-many foreign key relationship, not only has the association relationship of the attributes between the previous two single tables, but also has the attributes under the many-to-many foreign key association (such as Figure 2 shown).

[0059] A kind of many-to-many foreign key relational table data is as Figure 2 shown. The attributes User ID and Movie ID are the primary keys of the user table (U) and the movie table (V) respectively, uniquely identifying the data entries in the user table and the movie table. The scoring table is the link table L in the many-to-many foreign key relationship. There not only exist two foreign keys User ID and Movie ID representing the foreign key association between the user table and the movie table, but also contain other necessary attributes, such as the scoring value, etc. The existing generative models first model the single table user table U or the movie table V, and then merge the user table U or the movie table V with the scoring table L into a single table through the foreign key relationship, and then perform data synthesis based on the single table simulation algorithm. The existing methods have to deal with a contradiction: when modeling the scoring table L, the existing methods cannot decide whether to base on the data of the user table U or the data of the movie table V, and it is difficult to ensure that the generated scoring table L conforms to the logical consistency of the user table U and the movie table V. To overcome this problem, the present invention adopts a new generation order: first generate the scoring table L, and then generate the user table U and the movie table V respectively based on the already generated scoring table L. The specific process is as Figure 3as shown

[0060] The following is a detailed introduction to each step:

[0061] The solution of step S1 is used to mine the inter-table association attribute pairs. Input the single tables U, V with foreign key associations and the link table L, and output the attributes in the corresponding U and V tables for k attribute pairs.

[0062] The purpose of mining the association attribute pairs is to find the highly correlated attributes in tables U and V and merge them with the link table L, which not only avoids the excessively high data dimension caused by merging all attributes but also prevents the negative impact of irrelevant attribute pairs on the effectiveness of the simulation algorithm. As Figure 3 shown, its algorithm process is as follows: According to the foreign key matching relationship of the link table L, merge all the attributes of tables U and V; traverse the non-primary key attributes of table U , and combine them with the non-primary key attributes of each table V , and use the association evaluation function to measure the correlation of each attribute pair ; The Laplace and Gaussian mechanisms are mainly applicable to the privacy protection of numerical data, while the exponential mechanism can be applied to selective problems, such as selecting the best option under privacy protection from a set of discrete outputs. Apply the exponential mechanism: use the association evaluation function as the utility function (its sensitivity is 2), take as the output probability of the attribute pair , and randomly output k attribute pairs from the set of attribute pairs; traverse the results of the k outputs, and respectively collect the corresponding attributes of tables U and V {U m |1 ≤ m ≤ k}, {V m |1 ≤ m ≤ k}

[0063] Example of the application of step S1: As shown in Table 1, U and table V have non-primary key attributes U1, U2 and V1, V2 respectively, and 4 different attribute pairs can be obtained by combination. Based on the association evaluation function S(U i , V j ), calculate the association of different attribute pairs under foreign key association respectively. According to the exponential mechanism, each attribute pair (U i , V j ) is randomly output with a probability proportional to

[0064] Table 1:

[0065]

[0066] ​Let the number of output times \(k = 2\), and the results randomly output twice according to probability be \((U1, V2)\) and \((U2, V1)\). Then the corresponding attributes of the collected tables U and V are \(\{U1, U2\}\) and \(\{V1\}\).

[0067] Step S2 is used to simulate and generate a link table , input the link table L and the attributes in the corresponding U and V tables of k attribute pairs, and output the simulated and generated link table .

[0068] According to the foreign key correspondence relationship of the link table L, splice the attributes \(\{U\) m | \(1\leq m\leq k\}\), \(\{V\) m | \(1\leq m\leq k\}\) obtained in step S1 with the table L to obtain a merged table .

[0069] As Figure 4 shown, under the limitation of differential privacy, use the Markov random field model to simulate and generate the merged table U. The specific steps are as follows: Use the Markov random field model to model the original table U, and add Laplace noise during the modeling process to make it satisfy \(\epsilon\) 21 - differential privacy to obtain a graph structure; According to the existing graph structure, use mirror gradient descent to iteratively learn the parameters of the Markov random field model. The process of learning parameters applies the Gaussian mechanism to make it satisfy differential privacy; In each iteration process, in order to improve the model accuracy of the Markov random field as much as possible, the present invention calculates the gap between the marginal distribution of the model and the marginal distribution in the original data , and applies the Laplace mechanism to add noise to make it satisfy \(\epsilon\) 23 - differential privacy. Select the marginal distribution corresponding to the maximum value from the noisy version gap , and add the edges corresponding to the marginal distribution to the graph structure of the Markov random field. Let the privacy budget of step S2 be , according to the serial mechanism of differential privacy, step S2 satisfies differential privacy.

[0070] After obtaining the generation model, sample on the generation model to obtain the simulated data of the merged table . According to specific business requirements, the algorithm has two termination conditions: If it is required to maintain the same number of rows in the simulated data as in the original table, the algorithm terminates when the algorithm generates simulated data with the same number of rows as the original link table L; If it is required to maintain the foreign keys consistent, when the algorithm generates simulated data with the same number of rows as the original link table L, check whether the foreign keys of the original link table L exist in the simulated data table . If all exist, the algorithm terminates. If there are any missing, continue to sample using the generation model until all the foreign keys of the original link table L exist.

[0071] In addition, since there may be a one-to-many foreign key relationship between table U, table V and the link table L, it is necessary to perform aggregation operations on the attributes related to the two tables respectively: if the attribute is a categorical attribute, such as the gender of the user, etc., then according to the foreign key ID grouping, the proportion of each category within the group is sampled to generate this categorical data. If the attribute is a numerical attribute, such as the age of the user, etc., then the simulation data is replaced with its corresponding average according to the foreign key ID. Since both the sampling ratio and the average are calculated based on the simulation data, according to the properties of differential privacy post-processing, this step does not consume privacy budget.

[0072] Finally, according to the attributes of the original link table L, the merged table is split by column to obtain the simulated link table .

[0073] Step S3 generates a single table according to the synthesis result of the link table L , inputs the original data table U and the attributes U' corresponding to table U in the simulated link table , and outputs the simulated generated table .

[0074] Suppose the privacy budget of step S3 is ε3. Under the limitation of differential privacy, the Markov random field model is used to simulate and generate the merged table U. The specific steps are as follows:

[0075] Use the Markov random field model to model the original table U, and add Laplace noise during the modeling process to make it satisfy ε 31 -differential privacy to obtain the graph structure;

[0076] According to the existing graph structure, use mirror gradient descent to iteratively learn the parameters of the Markov random field model. The process of learning the parameters applies the Gaussian mechanism to make it satisfy differential privacy;

[0077] In each iteration process, in order to improve the model accuracy of the Markov random field as much as possible, the present invention calculates the gap between the marginal distribution of the model and the marginal distribution in the original data , and applies the Laplace mechanism to add noise to make it satisfy ε 33 -differential privacy. Select the marginal distribution corresponding to the maximum value from the noisy version of the gap , and add the edges corresponding to the marginal distribution to the graph structure of the Markov random field;

[0078] As Figure 5 shown, in order to maintain the consistency of the simulated table and the simulated link table , first fill the U ID obtained in step S2 and the attributes related to table U into and remain unchanged during the sampling process.

[0079] Under the given Markov random field model and the attributes filled from the simulation link table the remaining attributes of the U table are sampled and generated.

[0080] Let the privacy budget of step S3 be , according to the serial mechanism of differential privacy, step S3 satisfies differential privacy.

[0081] Step S4 generates a single table V according to the synthesis result of the link table L. Input: the original data table V and the attributes V' corresponding to the table V in the simulation link table Output: the simulated generated table .

[0082] Let the privacy budget of step S4 be ε4. Under the limitation of differential privacy, the combined table V is simulated and generated using the Markov random field model. The specific steps are as follows:

[0083] Use the Markov random field model to model the original table V, and add Laplace noise during the modeling process to make it satisfy ε 41 - differential privacy to obtain the graph structure;

[0084] According to the existing graph structure, use mirror gradient descent to iteratively learn the parameters of the Markov random field model. The process of learning the parameters applies the Gaussian mechanism to make it satisfy differential privacy;

[0085] In each iteration process, in order to improve the model accuracy of the Markov random field as much as possible, the present invention calculates the gap between the marginal distribution of the model and the marginal distribution in the original data , and applies the Laplace mechanism to add noise to make it satisfy ε 43 - differential privacy. Select the marginal distribution corresponding to the maximum value from the noisy version of the gap and add the edges corresponding to the marginal distribution to the graph structure of the Markov random field;

[0086] As Figure 6 shown, in order to maintain the consistency between the simulated table and the simulation link table , first fill the V ID obtained in step S2 and the attributes related to the table V into and remain unchanged during the sampling process;

[0087] Under the given Markov random field model and the attributes filled from the simulation link table the remaining attributes of the V table are sampled and generated.

[0088] Let the privacy budget for step S4 be , according to the serial mechanism of differential privacy, step S4 satisfies differential privacy.

[0089] Based on the algorithmic process of steps S1 - S4, according to the serial mechanism of differential privacy, the overall algorithm satisfies differential privacy.

[0090] Existing methods connect the parent table and the child table into a single table through foreign keys, and then use a single - table algorithm for synthetic data generation. However, the connection of irrelevant attributes not only reduces the running efficiency of the simulation algorithm, but also can only maintain the association relationship between the "U - table and L - table" or "V - table and L - table", which has an adverse effect on the utility of the simulation data. The present invention proposes a method for generating multi - table data with enhanced privacy. This method mines the association relationship between the attributes of tables with foreign - key relationships under differential privacy, and retains the attributes with significant association relationships and the link table for synthetic data generation, while maintaining the association relationship between the link table L and the two parent tables U and V; existing relational table - based synthetic data generation algorithms can neither be applied to table data with "many - to - many" foreign - key associations, nor apply privacy - enhancement technologies to provide necessary privacy protection for the model and the generated data. Based on the entire algorithmic process, the present invention applies differential privacy mechanisms such as the exponential mechanism, the Laplace mechanism, and the Gaussian mechanism, protecting the privacy and security of the synthetic data and the generation model in theory of differential privacy.

[0091] Aiming at the problems of low association efficiency and poor data utility caused by merging multiple tables into a single table in traditional methods, the present invention proposes a method for generating multi - table synthetic data that can effectively retain the association relationship. By analyzing the associated attributes between multiple tables with foreign - key relationships, attributes with significant association relationships are selected to better maintain the association between the child table and multiple parent tables (i.e., the L - table and the U - table, V - table) during synthetic data generation, thereby improving the utility of the synthetic data. To address the risk of potential privacy leakage in existing methods under "many - to - many" foreign - key relationships, the present invention integrates differential privacy mechanisms into the algorithmic process, including the exponential mechanism (associated attribute pairs output), the Laplace mechanism (constructing a graph model structure, calculating the marginal distribution gap), and the Gaussian mechanism (mirror gradient descent). These mechanisms provide privacy protection for the generation model and the synthetic data during the process of data synthetic generation to prevent information leakage, ensuring that the synthetic data can still maintain its authenticity and effectiveness while protecting privacy.

[0092] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for simulating and generating privacy-enhanced relational tabular data, characterized in that The method includes the following steps: Step S1: Mine highly correlated attributes in the link table L and the single tables U and V with foreign key associations, and merge the attributes with the link table L to obtain the attributes in the corresponding U and V tables for k attributes; Step S2: According to the foreign key correspondence of the link table L, splice the attributes obtained in Step S1 with the link table L to obtain a merged table , according to the attributes of the link table L, the merged table is split by columns to obtain a simulated generated link table ; Step S3: According to the synthesis result of the link table , use the Markov random field model to simulate and generate the table ; Step S4: According to the synthesis result of the link table , use the Markov random field model to simulate and generate the table ; The step S2 includes the following steps: Step S21: According to the foreign key correspondence of the link table L, splice the attributes {U m | 1 ≤ m ≤ k}, {V m | 1 ≤ m ≤ k} obtained in Step S1 with the table L to obtain a merged table ; Step S22: Model the original table U using a Markov random field model, and add Laplace noise during the modeling process to make it satisfy ε 21 - differential privacy to obtain a graph structure; Step S23: According to the existing graph structure, use mirror gradient descent to iteratively learn the parameters of the Markov random field model. The Gaussian mechanism is applied in the process of learning the parameters to satisfy differential privacy; Step S24: Calculate the difference between the marginal distribution of the model and the marginal distribution in the original data , and apply the Laplace mechanism to add noise to make it satisfy ε 23 -differential privacy. Select the marginal distribution corresponding to the maximum value from the noisy version of the difference , and add the edge corresponding to the marginal distribution to the graph structure of the Markov random field, and perform parameter learning in step S23; Step S25: Set the privacy budget of step S2 as , ε 21 is the privacy budget of step S21, ε 22 is the privacy budget of step S22, ε 23 is the privacy budget of step S23. According to the serial mechanism of differential privacy, step S2 satisfies differential privacy. After obtaining the generative model, sample on the generative model to obtain the simulation data of the merged table ; Step S26: According to the attributes of the original link table L, split the merged table by columns to obtain the link table generated by simulation .

2. A method for simulating and generating privacy-enhanced relational table data according to claim 1, wherein The step S1 includes the following steps: Step S11: Merge all the attributes of table U and table V according to the foreign key matching relationship of the link table L; Step S12: Traverse the non-primary key attributes of table U , and combine them with the non-primary key attributes of each table V . Use the correlation evaluation function to measure the correlation of each pair of attributes ; Step S13: Apply the exponential mechanism: Use the relevance evaluation function as the utility function, and use as the output probability of the attribute pair , as the sensitivity, and randomly output the attribute pair from the set of attribute pairs in step S12 k times ; Step S14: Traverse the results output k times, and respectively collect the corresponding attributes of Table U and Table V {U m | 1 ≤ m ≤ k}, {V m | 1 ≤ m ≤ k}.

3. A method for simulating and generating privacy-enhanced relational table data according to claim 1, wherein After obtaining the generation model in step S2, sample on the generation model to obtain the merged table of simulation data. The algorithm has two termination conditions. One is that if the number of rows of the simulation data is to be maintained the same as that of the original table, the algorithm terminates when it generates simulation data with the same number of rows as the original link table L. The other is that if the foreign keys are to be maintained consistent, when the algorithm generates simulation data with the same number of rows as the original link table L, check whether the foreign keys of the original link table L exist in the simulation data table . If all exist, the algorithm terminates. If there are any missing, continue to sample using the generation model until all the foreign keys of the original link table L exist 4. A method for simulating and generating privacy-enhanced relational table data according to claim 1, characterized in that If the relationship between table U, table V and the link table L is a one-to-many foreign key relationship, it is necessary to perform aggregation operations on the relevant attributes of the two tables respectively. If the attribute is a categorical attribute, sample and generate the categorical data according to the proportion of each category within the group after grouping by the foreign key ID; if the attribute is a numerical attribute, replace the simulation data with its corresponding average according to the foreign key ID. Since both the sampling ratio and the average are calculated based on the simulation data, according to the properties of differential privacy post-processing, no privacy budget is consumed.

5. A method for simulating and generating privacy-enhanced relational table data according to claim 1, characterized in that The step S3 includes the following steps: Step S31: Model the original table U using a Markov random field model and add Laplace noise during the modeling process to satisfy ε 31 - differential privacy to obtain a graph structure; Step S32: According to the existing graph structure, use mirror gradient descent to iteratively learn the parameters of the Markov random field model. The Gaussian mechanism is applied in the process of learning the parameters to satisfy differential privacy; Step S33: Calculate the difference between the marginal distribution of the model and the marginal distribution in the original data , and apply the Laplace mechanism to add noise to make it satisfy ε 33 -differential privacy, select the marginal distribution corresponding to the maximum value from the noisy version of the difference , and add the edge corresponding to the marginal distribution to the graph structure of the Markov random field; Step S34: To maintain the simulation table and the simulation link table consistent, first fill the U ID obtained in step S2 and the related attributes of table U into and keep them unchanged during the sampling process; Step S35: Under the given Markov random field model and the attributes filled from the simulation link table sample to generate the remaining U-table attributes; Step S36: Set the privacy budget of step S3 as , ε 31 is the privacy budget of step S31, ε 32 is the privacy budget of step S32, ε 33 is the privacy budget of step S33. According to the serial mechanism of differential privacy, step S3 satisfies differential privacy.

6. A method for simulating and generating privacy-enhanced relational table data according to claim 1, characterized in that, The step S4 includes the following steps: Step S41: Model the original table V using a Markov random field model and add Laplace noise during the modeling process to make it satisfy ε 41 - differential privacy to obtain a graph structure; Step S42: According to the existing graph structure, use mirror gradient descent to iteratively learn the parameters of the Markov random field model. The Gaussian mechanism is applied in the process of learning the parameters to make it satisfy differential privacy; Step S43: Calculate the difference between the marginal distribution of the model and the marginal distribution in the original data , and apply the Laplace mechanism to add noise to make it satisfy ε 43 -differential privacy, select the marginal distribution corresponding to the maximum value from the noisy version of the difference , and add the edge corresponding to the marginal distribution to the graph structure of the Markov random field; Step S44: To maintain the simulation table and the simulation link table consistent, first fill the V ID obtained in step S2 and the related attributes of table V into and keep them unchanged during the sampling process; Step S45: Under the given Markov random field model and the attributes filled in from the simulation link table sample and generate the remaining V-table attributes; Step S46: Set the privacy budget of step S4 as , ε 41 is the privacy budget of step S41, ε 42 is the privacy budget of step S42, ε 43 is the privacy budget of step S43. According to the serial mechanism of differential privacy, step S4 satisfies differential privacy.

7. A method for simulating and generating privacy-enhanced relational table data according to claim 1, characterized in that It further includes step S5: Based on the algorithmic processes of steps S1 to S4, according to the serial mechanism of differential privacy, the overall algorithm satisfies differential privacy.

Citation Information

Patent Citations

  • Relational table data generation method and system

    CN116089504A