Privacy-enhanced structured data simulation and generation method and system

By constructing a nonlinear Gaussian Bayesian network and using a variational gradient descent method, combined with Monte Carlo estimation algorithm, privacy-enhanced simulation data of structured data is generated, and the problems of insufficient privacy protection and poor data quality in the existing technology are solved, and high-quality privacy-enhanced simulation data generation is achieved.

WO2025107789A1PCT designated stage expired Publication Date: 2025-05-30HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Application Number
PCT/CN2024/115033
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-08-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively protect privacy when generating simulated data of structured data, especially when facing attacks with all background knowledge. After the traditional method introduces a differential privacy mechanism, it is easy to cause pattern crashes and poor quality of generated data.

Method used

A privacy-enhanced structured data simulation generation method is adopted to normalize data, build a probability graph model of a nonlinear Gaussian Bayesian network through data standardization, and infer the correlation relationship between features using the variational gradient descent method. The noise amount is automatically obtained in combination with the Monte Carlo estimation algorithm to meet the requirements of differential privacy.

Benefits of technology

It effectively protects the privacy of structured data, avoids the uncertain impact of different binning strategies on the results, retains the sequential properties of continuous numerical characteristics, and improves the quality of simulated data and the mathematical guarantee of privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024115033_30052025_PF_FP_ABST
    Figure CN2024115033_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a privacy-enhanced structured data simulation and generation method and system. The method comprises: step 1, a data conversion phase: performing normalization preprocessing on data; step 2, a probabilistic graphical model construction phase: on the basis of a Bayesian form, constructing the posterior distribution for variational inference from the data having undergone normalization preprocessing in step 1, obtaining an association relationship between the features of structured data by using a Stein variational gradient descent method, and when introducing differential privacy noise, using a Monte Carlo estimation algorithm to automatically obtain the amount of noise that needs to be added for each update step; and step 3: a data generation phase: using the association relationship obtained in step 2 as a metric set to generate more accurate simulation data than real data. The beneficial effect of the present invention is that: the method of the present invention avoids gradient clipping when applying DP-SGD, thereby avoiding the selection of clipping parameters and reducing the adverse effects of gradient clipping on an inference process.
Need to check novelty before this filing date? Find Prior Art

Description

A privacy-enhanced structured data simulation generation method and system Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a privacy-enhanced structured data simulation generation method and system. Background Art

[0002] With the rapid development of emerging technologies such as artificial intelligence and mobile internet, massive amounts of data are increasingly playing an indispensable role in various scientific and technological applications, becoming digital assets essential for the operations of relevant organizations and enterprises. As a widely used data type, structured data plays a vital role in fields such as the internet, finance, and healthcare. Furthermore, much unstructured data can be transformed into structured data after processing, such as parsing user web browsing history to identify user IP addresses, access times, and types of websites visited. While many AI applications, such as advertising recommendations and medical diagnosis, rely on structured data, analysis of such public datasets often leads to privacy breaches, such as revealing the age and educational background of individuals. In recent years, academics have invested increasing effort in this area, aiming to facilitate data analysis in a privacy-preserving manner. One key strategy is to generate simulated data.

[0003] Currently, a wealth of research has been conducted on simulation data generation. Methods using traditional privacy-preserving strategies (k-anonymization, l-diversity, and t-similarity) are unable to defend against attacks based on full background knowledge. Methods based on deep generative models can generate high-quality simulation data without considering privacy enhancement. However, the introduction of differential privacy mechanisms can easily lead to mode collapse, resulting in unsatisfactory quality of model-generated data. To balance privacy and data utility, researchers have attempted to capture the low-order distribution of structured data to approximate the original data distribution and add random noise to the low-order distribution to achieve differential privacy defenses. While these efforts have achieved remarkable results, there is still room for improvement.

[0004] Defects in existing technologies: For example: 1. Patent document CN109784091A, a tabular data privacy protection method that integrates differential privacy GAN and PATE models. The method proposed in this document uses DP-SGD and PATE models to train the generative model and teacher classifier respectively, and uses the classifier to identify the simulation data generated by the differential privacy generative model, selects data with consistent generated labels to train the student model, and finally releases the generative model and student classifier to meet the needs of downstream tasks for data synthesis and data analysis. Defects: (1) This method generates simulation data through DP-SGD and PATE models to meet the needs of differential privacy. However, due to the problems of mode collapse in GAN, the application of DP-SGD method often results in poor quality of generated simulation data. (2) In addition, this method involves the training and combination of multiple models such as teacher model and generator model, which has high computational cost and parameter selection requirements. 2. Patent document CN115455668A, a method, device, and electronic device for generating simulated data for tabular data, transforms the original data using a cumulative distribution table. Then, by obtaining the mean and covariance matrix of each column, it generates jointly Gaussian distributed data. Finally, it generates simulated data by looking up the inverse of the cumulative distribution table. This method generates simulated data with high efficiency, requires little time, and produces high-quality data in a distributed environment. However, it has drawbacks: This method is applicable in federated learning environments. Its main advantage lies in its efficiency. However, because it does not consider privacy protection, the generated simulated data may still leak personal sensitive data. 3. Patent document CN114385868A, a method and device for generating relational tabular data based on a generative adversarial network, uses classifiable data that uniquely identifies an entity as an entity identification attribute and conditional information. A random noise vector is used as input, and the generated relational tabular data is generated using a pre-trained data generation model. A drawback: This method generates simulated data using a DPGAN to meet differential privacy requirements. However, due to gradient clipping and differential privacy noise, the generated simulated data is of poor quality. 4. Patent document CN112836830A, a method for parallel training of federated gradient boosting decision trees, uses privacy-preserving tabular data to counter the generative network and KD-tree to cluster and sample synthetic and local samples, creating a mixed sample set close to the overall data distribution. Based on these global mixed samples and local samples, federated learning is combined to complete the training of the gradient boosting decision tree, and histogram optimization and other methods are used to reduce communication costs. However, there are defects: (1) The purpose of this patent is to improve the training efficiency of federated learning, which is different from the purpose of this patent; (2) This method involves the training and combination of multiple models such as teacher models and generator models, which has high requirements for computational cost and parameter selection. 5. Patent document CN107886009A, a method for generating big data to prevent privacy leakage.This document uses random numbers to represent the eigenvalues ​​of each feature based on the probability distribution of the original data features to avoid sensitive privacy information in the data, and then uses the nearest neighbor model to generate simulation data. The defects are: (1) The generation of features is independent and cannot capture the correlation between attributes; (2) This method can only generate discrete data; (3) This method protects real data by using random numbers, which lacks the mathematical guarantee of privacy compared with the differential privacy mechanism. 6. Patent document CN110377725A, data generation method, device, computer equipment and storage medium. This document extracts information from text data, generates candidate key information (including keywords, key phrases, etc.), and generates simulation data records containing semantic text information by integrating the key information. The defects are: (1) This method cannot generate simulation data that is more widely applicable and contains continuous numerical features and discrete categorical features; (2) This method does not adopt a privacy security protection strategy. 7. Patent document CN115169252A, a system and method for generating structured simulation data, uses a Bayesian network to mine associations between features, simultaneously processing both discrete and continuous features, and generating structured simulation data using a generative adversarial network. However, this method has the following drawbacks: (1) When mining associations between features, it requires binning continuous data, which inevitably results in information loss; (2) This method does not employ a privacy and security protection strategy.

[0005] Summary of the Invention

[0006] In order to solve the problems in the prior art, the present invention provides a privacy-enhanced structured data simulation generation method and system.

[0007] The present invention provides a privacy-enhanced structured data simulation generation method, comprising the following steps:

[0008] Step 1, data conversion stage: normalize and preprocess structured data;

[0009] Step 2, probabilistic graphical model construction phase: construct a probabilistic graphical model of the structured data that has been normalized and preprocessed in step 1, use the variational gradient descent method to infer the correlation between the features of the structured data, and use the Monte Carlo estimation algorithm to automatically obtain the amount of noise required to be added for each update step;

[0010] Step 3, data generation stage: using the association relationship obtained in step 2 as a measurement set to generate simulation data.

[0011] As a further improvement of the present invention, the step 1 includes:

[0012] Normalization steps for numerical data: perform Min-Max normalization on each numerical type of data, limiting the range of all numerical values ​​to (0,1);

[0013] Normalization steps for categorical data: Map the numerical domain of categorical data to encode the categorical data between (0,1).

[0014] As a further improvement of the present invention, the categorical data normalization step specifically includes:

[0015] Step 1: If the category feature is an ordinal feature, sort the categories according to the ordinal relationship of the attributes; otherwise, sort them in descending order according to the size of the category proportion;

[0016] Step 2: Divide the interval [0,1] into parts according to the cumulative probability of each category;

[0017] Step 3: Find the interval [a,b]∈[0,1] corresponding to the category to be converted and select its middle value As a mapping of the category to the numeric domain.

[0018] As a further improvement of the present invention, the step 2 includes:

[0019] Step S1: using a nonlinear Gaussian Bayesian network to construct a probabilistic graphical model for the structured data pre-processed by the normalization step 1;

[0020] Step S2: Use a three-layer dense convolutional network to learn the nonlinear functional relationship between the nodes of the nonlinear Gaussian Bayesian network. The number of neurons in each layer is d is the number of features, ReLU activation function is used between layers, and sigmoid activation function is used in the last layer of the probabilistic graph model;

[0021] Step S3: Use Stein's variational gradient descent method to infer the approximate posterior distribution in the probabilistic graphical model to obtain a more accurate description of the correlation between structured data features;

[0022] Step S4: Use the Monte Carlo estimation algorithm to automatically obtain the sensitivity of the log-likelihood function of the posterior probability p(Θ, D|G) at each update step of the latent variable Z, and add the required amount of Laplace noise according to the serial combination mechanism of differential privacy.

[0023] As a further improvement of the present invention, step S3 is specifically as follows:

[0024] Define a Bayesian network (G, θ) to simulate the joint density p(x) of a set of d rows of structured data x = x1:d, where the directed acyclic graph G encodes the conditional independence of x, and the parameter θ defines the numerical relationship parameter of each variable under the given conditions of the parent node in the DAG.

[0025] As a further improvement of the present invention, step S3 is specifically as follows:

[0026] Given a set of hidden variable particles of size M And the current iteration number t, through iterative update infer the latent variable Z:

[0027] Among them, η t Represents the amplitude of the perturbation, the purpose of the perturbation is to minimize the KL divergence; To update to the mth particle in the tth round, is the mth particle in the t+1th round;

[0028] φ t (x) is a smooth function describing the disturbance direction, for φ t ,have

[0029] in, It is a customizable kernel function, which aims to measure the The distance between represents the gradient of the logarithmic posterior probability density of the latent variable Z;

[0030] Use a simple kernel function, as shown in the following formula, where h is the bandwidth of the kernel function and is an adjustable parameter.

[0031] Among them, k(X,Y) is the kernel function with parameters X and Y, is the square of the 2-norm of XY, the gradient of the logarithmic posterior probability density of the latent variable Z According to the factor decomposition formula summarized by the Bayesian model, we can get

[0032] in, represents the gradient of the logarithmic posterior probability density of the latent variable Z, which is equivalent to represents the gradient of the logarithmic prior probability density of the latent variable Z, Represents the logarithmic conditional probability density of data D when the latent variable Z is given, After mathematical transformation, Ep(G|Z)[p(D,θ|Z)] is the expectation of the conditional probability density p(D,θ|Z) of the data D when the latent variable Z is given, calculated based on the probability p(G|Z) of the latent variable Z to the graph G.

[0033] As a further improvement of the present invention, step S4 further includes:

[0034] Step Y1: Calculate the sensitivity Δ of the posterior probability of a single numerical feature num and the sensitivity Δ to the posterior probability of a single category feature cat ;

[0035] Step Y2: Assume that the number of categorical features of dataset D is |C cat |, the number of numerical features is |C num |, when the latent variable Z is updated in a single step, the sensitivity of the log-likelihood function of the posterior probability p(Θ,D|G) is:

[0036] Where σ is a preset constant;

[0037] Step Y3: Assume that the total number of iterations is T, the number of particles is M, the number of Monte Carlo sampling is N, and the probability density function is The Laplace noise is Lap(λ). According to the serial combination mechanism of differential privacy, the log-likelihood function of the posterior probability p(θ, D|G) in each round is calculated to obtain the Laplace noise that needs to be added

[0038] As a further improvement of the present invention, the step three specifically includes:

[0039] Step A1: Given a feature subset C of a dataset and an edge query set Q c , privacy budget ε c , and a data set D containing m pieces of data = (x 1 ,...,x m ), based on the feature set subset C of a given data set, the marginal probability vector associated with its feature subset is defined as Where m is the number of rows in the dataset, i is the data row index, and x i is the i-th data, x c is a value of the domain under the metric set C, μ c is the marginal probability vector associated with the feature subset C, and we get c -Measurement of differential privacy Where Lap(λ) is a probability density function Laplace noise, λ refers to the parameter of the Lap() function, which represents an arbitrary formula, ΔQ cRepresents the edge query set Q c Sensitivity under the definition of differential privacy;

[0040] Step A2: For a certain edge μ, define the loss function L(x) = ||Qx-y||, and infer the simulation data that meets the metric set conditions to convert it into the optimized loss function L(x);

[0041] Step A3: Use the mirror gradient descent method to minimize the loss function L(x) so that the generated simulation data meets the constraints of the association relationship as much as possible.

[0042] As a further improvement of the present invention, in step A1, according to the serial combination mechanism, for all the privacy budgets ε c -Measurement of differential privacy

[0043] y c , the combined metric set y=(y c ) c∈C The required privacy budget for each metric y c The sum of the budget, that is, satisfying ε2=∑ c∈C ε c Differential privacy mechanism, the overall algorithm satisfies the ε1+ε2 differential privacy mechanism, where if the number of data items is m, then the sensitivity ||Q||1 is the sum of the absolute values ​​of matrix Q along the column direction, and the maximum value is taken as the 1-norm.

[0044] The present invention also discloses a privacy-enhanced structured data simulation generation system, comprising: a memory, a processor, and a computer program stored on the memory, wherein the computer program is configured to implement the steps of the structured data simulation generation method of the present invention when called by the processor.

[0045] The beneficial effects of the present invention are as follows: 1. The present invention proposes a structured data generation method for data security scenarios. By applying a differential privacy mechanism, it ensures the privacy of generated simulation data, which is beneficial for data analysis engineers to carry out various downstream tasks and give full play to the data utility. 2. The present invention analyzes the currently most effective marginal distribution-based method and finds that it must bin continuous numerical features when mining the correlation between features. The method of the present invention avoids the uncertainty caused by different binning strategies on the experiment while retaining the sequential properties of the original continuous numerical values. 3. Traditional probabilistic graphical model construction methods such as Bayesian structure learning require obtaining samples in an exponential space with the number of nodes as the order to calculate the corresponding statistics, which often requires a fixed limit on the size of the parent node set. The present invention adopts an approximate inference method of a variational graph autoencoder, which does not require a limit on the size of the parent node set of each node, and the resulting correlation between features is more accurate. 4. When applying the differential privacy mechanism, the present invention uses a Monte Carlo algorithm to estimate the gradient of each iteration, avoiding gradient clipping when applying DP-SGD. This not only avoids the selection of clipping parameters but also alleviates the adverse effects of gradient clipping on the inference process. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] FIG1 is an example diagram of a privacy-enhanced simulation data scenario according to the present invention;

[0047] FIG2 is a flow chart of simulation data generation according to the present invention. DETAILED DESCRIPTION

[0048] Terminology Notes:

[0049] Differential privacy: Differential privacy is a privacy-enhancing technology based on mathematical theory that allows adding or removing any data record without significantly affecting query results. Therefore, even if an attacker knows all but one record in a dataset, that record remains undisclosed. Definitions relevant to this invention are as follows:

[0050] Definition 1 ε-Differential Privacy: For any two datasets D1 and D2 with the same data structure, D1 and D2 differ by only one record (neighboring dataset). Suppose there is a random algorithm M, the set of all possible output results of the algorithm is O, and any output result in datasets D1 and D2 is If the random algorithm M satisfies:

[0051] Where Pr[★] represents the probability of an event occurring, and the parameter ε represents the privacy protection budget, then the algorithm M satisfies ε-differential privacy protection.

[0052] Definition 2 Sensitivity: Given a query function f:D→R d , D is the input data set, Rd is the output data set. On any pair of D1 and D2, the sensitivity of function f is

[0053] Where ||f(D1)-f(D2)||1 is the first-order distance between f(D1) and f(D2). Sensitivity measures the impact of changes in the input dataset on the corresponding output. Based on sensitivity, the present invention can provide an implementation mechanism for differential privacy.

[0054] Definition 3: Laplace mechanism: Given a dataset D and a privacy budget ε, the sensitivity of function f is Δf, when the output of f satisfies:

[0055] Algorithm A is said to satisfy ε-differential privacy, where Lap(Δf / ε) is random noise that satisfies the Laplace distribution.

[0056] Definition 4 Gaussian mechanism: Given a dataset D and a privacy budget ε, the global sensitivity of function f is Δf. For any δ∈(0, 1), When the output of f satisfies:

[0057] A(D)=f(D)+N(0,σ 2 )

[0058] Then the algorithm A is said to satisfy (ε, δ)-differential privacy, where δ represents the relaxation term, for example, it is set to 10 -5 , which means that only 10 -5 The probability of violating strict differential privacy.

[0059] Definition 5: Serial combination mechanism: For each of i -Sub-algorithm A of differential privacy i (x) = O i , the processes of any two sub-algorithms are independent of each other, then the overall output of the algorithm is {O1, O2, ..., O m ,satisfy -Differential privacy.

[0060] Bayesian Network: A Bayesian network is a graphical model used to describe the dependencies and probability distributions between random variables. It represents conditional dependencies between variables using a directed acyclic graph, where each node represents a variable and the edges represent the dependencies. Bayesian networks, combined with probability distributions, can infer relationships between variables, learn from data, and predict probabilistic events. They are widely used in fields such as causality, forecasting, and decision analysis.

[0061] Stein's variational gradient descent: Stein's variational gradient descent method is a technique for approximating the inference of latent variables in probabilistic models. It constructs a special distribution, the variational distribution, to approximate the true posterior distribution, making complex inference problems tractable. Stein's variational gradient descent optimizes by minimizing the difference between the variational distribution and the true posterior, using random samples to estimate the expectation and gradient. This method has broad applications in fields such as Bayesian inference, machine learning, and probabilistic graphical models, enabling efficient inference and parameter estimation in complex models.

[0062] In the context of privacy-enhanced data generation, this paper innovatively proposes a structured data generation technique for complex numerical relationships, building on the currently effective marginal distribution-based approach. This method eliminates the need for manual binning of continuous data, avoiding the impact of different binning strategies on the results while preserving the sequential nature of continuous values ​​in the data. This paper proposes a privacy-enhanced structured data simulation generation method, the main contents of which are as follows:

[0063] 1. Privacy-enhanced structured data generation scenario

[0064] As shown in Figure 1, in various artificial intelligence applications, the simulation data method of the privacy enhancement mechanism can generate simulation data that meets privacy requirements, allowing data analysis engineers to perform various downstream tasks on the data and give full play to the data's utility.

[0065] Unlike images, text, and other data, research has shown that the relationships between features in tabular data are sparse. This type of method aims to approximate the high-dimensional distribution of the original data using a set of low-dimensional distributions related to the data dependencies. For a sufficiently accurate approximation, the resulting data will maintain high accuracy for nearly any type of query (linear or nonlinear). Because high-dimensional data is converted into a low-dimensional structure, this method can effectively avoid the "curse of dimensionality" problem of directly adding noise to the original data distribution.

[0066] Current differentially private structured data generation methods use statistics like mutual information to construct edges for each feature node, which in turn leads to probabilistic graph models like Bayesian networks, approximating the original data distribution. These methods rely on the calculation of accurate statistics like mutual information and require discretization of the data before constructing the simulation data generation model. This inevitably undermines the properties of continuous numerical features and makes it difficult to capture the correlations between continuous numerical features and other features.

[0067] 2. Privacy-enhanced simulation generation process for complex numerical features

[0068] The structured data generation method proposed in this paper takes real-world structured data and, while preserving the numerical order of the original data, infers a probabilistic graphical model based on Stein's variational gradient descent method. The Monte Carlo estimation algorithm automatically estimates the amount of noise required for each update step, avoiding the impact of gradient clipping on the results when using DP-SGD while achieving the requirements of compact differential privacy (ε-DP). The specific steps are as follows, and the specific process is shown in Figure 2:

[0069] The present invention discloses a privacy-enhanced structured data simulation generation method, comprising the following steps:

[0070] Step 1, data conversion stage: normalize and preprocess structured data;

[0071] Step 2: Probabilistic graphical model construction phase: Build a probabilistic graphical model of the structured data that has been normalized and preprocessed in step 1. Use the variational gradient descent method to infer the correlation between the features of the structured data, and use the Monte Carlo estimation algorithm to automatically obtain the amount of noise required for each update step.

[0072] Step 3, data generation stage: using the association relationship obtained in step 2 as a measurement set to generate simulation data.

[0073] The following describes the various parts of privacy-enhanced simulation data generation:

[0074] Data conversion:

[0075] The main purpose of simulation data is to simulate real-world conditions by generating simulated data. However, real-world data often comes from different physical quantities with different dimensions, numerical ranges, and distributions. Data normalization preprocessing can standardize the original data into a consistent form, which helps to convert data of different scales and ranges into a unified range. It can reduce the weight differences between features, thereby more evenly affecting the learning effect of the model. Specifically,

[0076] Steps to normalize numerical data:

[0077] Because different numerical features have different ranges and sizes, attributes with large value ranges may be considered to have stronger dependencies on other attributes. Before building a probabilistic graphical model, we first perform Min-Max normalization on the data of each numerical type, restricting all numerical values ​​to the range (0, 1).

[0078] Steps for normalizing categorical data:

[0079] Categorical data is another key type of data in tabular data. It needs to be mapped to a numerical domain before it can be modeled using a nonlinear Gaussian Bayesian network. The specific steps for normalizing categorical data include:

[0080] Step 1: If the category feature is an ordinal feature, sort the categories according to the ordinal relationship of the attributes; otherwise, sort them in descending order according to the size of the category proportion;

[0081] Step 2: Divide the interval [0,1] into parts according to the cumulative probability of each category;

[0082] Step 3: Find the interval [a,b]∈[0,1] corresponding to the category to be converted and select its middle value As a mapping of the category to the numeric domain.

[0083] The above method encodes categorical data between (0,1), so that the encoding result takes into account the proportion of different categories.

[0084] Probabilistic graphical model construction phase:

[0085] 1) Selection of Probabilistic Graphical Model

[0086] Step S1: To better model the complex numerical relationships present in structured data, the present invention uses a nonlinear Gaussian Bayesian network to construct a probabilistic graphical model for the structured data pre-processed in step 1. A Bayesian network (G, θ) is defined to simulate the joint density p(x) of a set of d rows of structured data x = x1:d. The directed acyclic graph G encodes the conditional independence of x; the parameter θ defines the numerical relationship parameter of each variable under given conditions of the parent node in the DAG.

[0087] Step S2: Unlike the linear Gaussian Bayesian network, the nonlinear Gaussian Bayesian network allows the functional dependency between nodes to be nonlinear. Its core is to use nonlinear functions to describe the dependency between nodes, which makes the probabilistic graph model better adaptable to nonlinear data. Since neural networks are good at modeling high-order nonlinear functional relationships, the present invention uses a three-layer DenseNet (dense convolutional network) to learn the nonlinear functional relationship between nodes. The number of neurons in each layer is (d is the number of features), and the ReLU activation function is used between layers. In order to map the output to the (0, 1) interval, the last layer of the probabilistic graphical model uses the sigmoid activation function.

[0088] 2) Inference methods of probabilistic graphical models

[0089] Step S3: Use the variational gradient descent method to infer the approximate posterior distribution in the probabilistic graphical model to obtain a more accurate description of the correlation between structured data features; the details are as follows:

[0090] Traditional methods require acquiring samples in an exponential space whose order is the number of nodes to calculate the corresponding statistics, often imposing hard limits on the size of the set of parent nodes. To overcome these hard limits in most structural learning methods, recent research has leveraged the mechanism of variational graph autoencoders to shift the posterior inference task to a latent space represented by a probabilistic graph. This paper also adopts this concept, using a variational graph autoencoder to perform variational inference on the associations of tabular data.

[0091] Variational graph autoencoder contains graph G = {0, 1} d*d , d represents the number of nodes. It models the generation process of graph G by introducing latent variables Z, where latent variables Z have two embedding matrices U, V∈R d*k , k is a parameter independent of d. U and V are initialized as Existing research works represent the discrete DAG structure in a continuous space and construct a mapping from the latent variable Z to the connection matrix G through the matrix inner product;

[0092] in, represents the sigmoid function and its temperature factor α, where the choice of k is independent of d.

[0093] For the observed data set, that is, the simulated data set D = {x0, x1...x n The goal of this paper is to infer the complete posterior probability of the probabilistic graphical model that models this dataset. According to the method of Friedman and Koller, assuming that there is a prior distribution p(G) of the directed acyclic graph and a prior distribution p(θ|G) of the Bayesian network parameters, the posterior distribution can be obtained by Bayes' theorem:

[0094] p(G, θ|D)∝p(G)p(θ|G)p(D|G, θ) #(3),

[0095] Without loss of generality, assuming that there is a latent variable Z to model the generative process of G, the Bayesian model is summarized as the following factorization:

[0096] p(Z,G,θ,D)∝p(Z)p(G|Z)p(θ|G)p(D|G,θ) #(4),

[0097] Among them, p(Z) can be expressed using the prior probability of the acyclic graph, as shown in Equation (5), so that the constructed probabilistic graphical model can ensure the property of directed acyclicity as much as possible; p(G|Z) is given by the mapping from the latent variable Z in the variational graph autoencoder to the graph G.

[0098] To infer the latent variable Z, related research has proposed a feasible solution based on Stein's variational gradient descent method. Stein's variational gradient descent (SVGD) is an optimization algorithm for variational inference, used to approximate the posterior distribution. This method combines the concepts of variational inference and gradient descent to efficiently approximate the posterior distribution in complex probabilistic models.

[0099] In Stein's variational gradient descent method, given a set of hidden variable particles of size M and the current iteration number t, the present invention can infer the latent variable Z through iterative update:

[0100] Among them, η t Represents the amplitude of the perturbation, the purpose of the perturbation is to minimize the KL divergence; To update to the mth particle in the tth round, is the mth particle in the t+1th round;

[0101] φ t (x) is a smooth function describing the disturbance direction, for φ t ,have

[0102] in, It is a customizable kernel function, which aims to measure the The distance between them, in SVGD theory, φ t (x) is a smooth function describing the disturbance direction, represents the gradient of the logarithmic posterior probability density of the latent variable Z;

[0103] The variational gradient descent method uses a simple kernel function, as shown in Equation (8), where h is the bandwidth of the kernel function as an adjustable parameter;

[0104] Among them, k(X,Y) is the kernel function with parameters X and Y, is the 2-norm square of XY. This kernel function is the Gaussian kernel function, which is the most widely used. It can map the original features to infinite dimensions.

[0105] The gradient of the logarithmic posterior probability density of the latent variable Z The factor decomposition formula (4) summarized according to the Bayesian model is:

[0106] in, represents the gradient of the logarithmic posterior probability density of the latent variable Z, which is equivalent to represents the gradient of the logarithmic prior probability density of the latent variable Z, Represents the logarithmic conditional probability density of data D when the latent variable Z is given, After mathematical transformation, E p (G|Z)[p(D,θ|Z)] is the expectation of the conditional probability density p(D,θ|Z) of the data D when the latent variable Z is given, calculated based on the probability p(G|Z) of the latent variable Z to the graph G. Note: and Both represent the gradient of the logarithmic posterior probability density of the latent variable Z. The latter Z adds superscripts and subscripts, which have specific meanings. This formula appears in formula (7), which is the gradient of φ in formula (6) t The expansion of , involving superscripts and subscripts, has been explained in the explanation of formula (6): To update to the mth particle in the tth round, is the mth particle in the t+1th round.

[0107] Step S4: Use the Monte Carlo estimation algorithm to automatically obtain the sensitivity of the log-likelihood function of the posterior probability p(θ, D|G) at each update of the latent variable Z, and add the required amount of Laplace noise according to the serial combination mechanism of differential privacy. The specific steps include:

[0108] Step Y1: Calculate the sensitivity Δ of the posterior probability of a single numerical feature num and the sensitivity Δ to the posterior probability of a single category feature cat ;

[0109] Specifically, in order to ensure that the generated graph structure does not expose private information, the present invention needs to introduce a differential privacy mechanism. During the gradient descent process, the existing differential privacy mechanism uses the DP-SGD (differential privacy stochastic gradient descent) method with gradient clipping to achieve privacy enhancement by applying the Gaussian mechanism. Generally speaking, DP-SGD requires manual adjustment of the clipping threshold to achieve the desired privacy level, and determining the appropriate threshold is challenging. Based on the existing gradient estimation method, the present invention uses a Monte Carlo-based scoring function gradient estimator to estimate the gradient in Equation (9), namely:

[0110] As can be seen from Equation (10), the original data set D is only involved in estimating the posterior probability p(θ, D|G). Under the framework of the nonlinear Gaussian Bayesian network, the present invention only needs to consider the sensitivity calculation of the logarithmic probability under the Gaussian distribution probability density function and add noise according to the Laplace mechanism. In order to calculate the sensitivity of the logarithm of the probability density function of the Gaussian distribution, the present invention needs to determine the maximum difference in the output when considering adjacent data sets D and D' with different values ​​(which only differ in the value of a single data point).

[0111] In a nonlinear Gaussian Bayesian network, the posterior probability can be defined as:

[0112] Let σ be a constant and represent the Gaussian distribution as f(x,μ), where x represents the input data point and μ represents the original data point. The probability density function of the Gaussian distribution is:

[0113] To calculate the sensitivity, the present invention considers two adjacent data sets D and D', where D' contains the added data x'. The sensitivity can be calculated as follows:

[0114] Among them, DN j is the DenseNet (non-linear function) of the jth feature, x i is the i-th row data, μ ij The data in the i-th row and j-th column for the current log-likelihood calculation.

[0115] Let the output DN of DenseNet be j (x i )≡o ij , take the logarithm of the probability density function, we have:

[0116] Since the first formula Being constant and independent of the data points, the present invention can ignore it in the sensitivity calculation. Therefore, the sensitivity becomes:

[0117] For numerical features, in order to eliminate the differences between dimensions and ensure that the contribution weights of all features to the posterior probability are similar, the present invention maps the value range of the data to the interval [0, 1] through regularization.

[0118] because:

[0119] |o-μ|≤1(0≤o,μ≤1) #(16),

[0120] For o′,μ′ corresponding to adding a piece of data, we have:

[0121] 0≤(o′-μ′) 2 ≤1 #(17),

[0122] Therefore, its contribution to sensitivity is:

[0123] For categorical features, the encoding of each category is based on the proportion of the category in the total. After modifying a piece of data, not only the source of the modified data x', but also the encoding of the same feature on other data will change. Therefore, the posterior probability p(x cat , D|G) sensitivity Δ cat It comes from two parts: (1) the change Δf1 caused by the modified data x'; (2) the coding change Δf2 caused by the change in category proportion.

[0124] For Δf1, similar to the numerical type, we have:

[0125] For Δf2, let the number of data rows be n and the number of categories be |C|. We need to count the differences in the results caused by the encoding changes of each row of data:

[0126] Therefore, the contribution of a single category feature to the sensitivity Δ cat is the sum of the two:

[0127] In summary, let the number of categorical features of the dataset D be |C cat |, the number of numerical features is |C num |, the present invention can obtain the sensitivity of the log-likelihood function of the posterior probability p(θ, D|G) when the latent variable Z is updated in a single step:

[0128] Where σ is a constant set in advance;

[0129] Step Y3: Finally, let the total number of iterations be T, the number of particles be M, the number of Monte Carlo sampling be N, and the probability density function be The Laplace noise is Lap(λ). According to the serial combination mechanism of differential privacy, in order to make the algorithm for building the probabilistic graphical model satisfy ε1-differential privacy, the log-likelihood function of the posterior probability p(θ,D|G) is calculated in each round. The Laplace noise that needs to be added is

[0130] Data generation phase:

[0131] Data generation method based on probabilistic graphical model and differential privacy mechanism

[0132] After steps 1 and 2, the present invention can obtain a low-dimensional distribution that describes the correlation between structured data features. In order to obtain simulation data that is more accurate than real data, the present invention uses the obtained correlation as a metric set. Specifically, it includes:

[0133] Step A1: Given the feature subset C of the dataset and the edge query set Q c, privacy budget ε c , and a data set D containing m pieces of data = (x 1 ,…,x m ), according to the feature set subset C of a given data set, the marginal probability vector associated with its feature subset can be defined as We can get the privacy budget - differential privacy (ε c -DP) Where Lap(λ) is a probability density function Laplace noise, where m is the number of rows in the dataset, i is the data row index, and x i is the i-th data, x c is a value of the domain under the metric set C, μ c is the marginal probability vector associated with the feature subset C. For simplicity, λ here refers to the parameter of the Lap() function, which represents an arbitrary formula, ΔQ c Represents the edge query set Q c Sensitivity under the definition of differential privacy. According to the serial combination mechanism, for all privacy budget ε c - Differential privacy measure y c , the combined metric y = (y c ) c∈C Satisfy the required privacy budget for each metric y c Budget c The sum of ε2=∑ c∈C ε c Differential privacy mechanism, the overall algorithm satisfies the ε1+ε2 differential privacy mechanism. Among them, if the number of data is m, then the sensitivity ||Q||1 is the sum of the absolute values ​​of matrix Q along the column direction, and the maximum value is taken as the 1-norm of Q.

[0134] Step A2: For a certain edge μ, define the loss function L(x) = ||Qx-y||, and infer that the simulation data that meets the metric set conditions is converted into the optimized loss function L(x).

[0135] Step A3: Related optimization methods have been studied maturely. The present invention uses mirror gradient descent to minimize the loss function so that the generated simulation data meets the constraints of the association relationship as much as possible.

[0136] 1. This paper proposes a method for generating structured data for complex numerical relationships. Users only need to provide real data without binning continuous data to obtain a probabilistic graphical model representing the correlation between features, avoiding the unnecessary impact of binning operations.

[0137] 2. This paper proposes a method for constructing probabilistic graphical models based on variational inference and differential privacy mechanisms. This method constructs the posterior distribution for variational inference using a Bayesian approach and learns the probabilistic graphical model using Stein's variational gradient descent method. When introducing differential privacy noise, a Monte Carlo estimation method is used to bound the sensitivity of each gradient descent update, thus avoiding the impact of gradient clipping on inference performance.

[0138] The beneficial effects of the present invention are as follows: 1. The present invention proposes a structured data generation method for data security scenarios. By applying a differential privacy mechanism, it ensures the privacy of generated simulation data, which is beneficial for data analysis engineers to carry out various downstream tasks and give full play to the data utility. 2. The present invention analyzes the currently most effective marginal distribution-based method and finds that it must bin continuous numerical features when mining the correlation between features. The method of the present invention avoids the uncertainty caused by different binning strategies on the experiment while retaining the sequential properties of the original continuous numerical values. 3. Traditional probabilistic graphical model construction methods such as Bayesian structure learning require obtaining samples in an exponential space with the number of nodes as the order to calculate the corresponding statistics, which often requires a fixed limit on the size of the parent node set. The present invention adopts an approximate inference method of a variational graph autoencoder, which does not require a limit on the size of the parent node set of each node, and the resulting correlation between features is more accurate. 4. When applying the differential privacy mechanism, the present invention uses a Monte Carlo algorithm to estimate the gradient of each iteration, avoiding gradient clipping when applying DP-SGD. This not only avoids the selection of clipping parameters but also alleviates the adverse effects of gradient clipping on the inference process.

[0139] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A privacy-enhanced structured data simulation generation method, characterized in that: The following steps are involved: Step 1, data conversion stage: normalize and preprocess structured data; Step 2, probabilistic graph model construction phase: construct a probabilistic graph model of the structured data preprocessed by normalization in step 1, use the variational gradient descent method to infer the correlation between the features of the structured data, and use the Monte Carlo estimation algorithm to automatically obtain the amount of noise required to be added for each update step; Step three, data generation stage: using the association relationship obtained in step two as a measurement set to generate simulation data.

2. The structured data simulation generation method according to claim 1, characterized in that: The step one comprises: Normalization steps for numerical data: perform Min-Max normalization on each numerical type of data and limit the range of all numerical values ​​to (0,1); Normalization steps for categorical data: Map the numerical domain of categorical data to encode the categorical data between (0,1).

3. The structured data simulation generation method according to claim 2, characterized in that: The categorical data normalization step specifically includes: Step 1: If the category feature is an ordinal feature, sort the categories according to the ordinal relationship of the attributes; otherwise, sort them in descending order according to the size of the category proportion; Step 2: Divide the interval [0,1] into parts according to the cumulative probability of each category; Step 3: Find the interval [a,b]∈[0,1] corresponding to the category to be converted and select its middle value As a mapping of the category in the numeric domain.

4. The structured data simulation generation method according to claim 2, characterized in that: The second step comprises: Step S1: constructing a probability graph model for the structured data preprocessed by the normalization in step 1 using a nonlinear Gaussian Bayesian network; Step S2: Use a three-layer dense convolutional network to learn the nonlinear functional relationship between the nodes of the nonlinear Gaussian Bayesian network. The number of neurons in each layer is d is the number of features, ReLU activation function is used between layers, and the last layer of the probabilistic graph model uses sigmoid activation function; Step S3: Use Stein's variational gradient descent method to infer the approximate posterior distribution in the probabilistic graphical model to obtain a more accurate description of the correlation between the structured data features; Step S4: Use the Monte Carlo estimation algorithm to automatically obtain the posterior probability p(Θ,D|G) of the latent variable Z at each update step The sensitivity of the log-likelihood function, the amount of Laplace noise required to be added according to the serial combination mechanism of differential privacy.

5. The structured data simulation generation method according to claim 4, characterized in that: The step S3 specifically includes: defining a Bayesian network (G, Θ) for simulating a set of d rows of structured data x=x 1:d The joint density p(x) of , where the directed acyclic graph G encodes the conditional independence of x, and the parameter Θ defines the numerical relationship parameter of each variable under the given conditions of the parent node in the DAG.

6. The structured data simulation generation method according to claim 4, characterized in that: The step S3 is specifically as follows: given a hidden variable particle set of size M And the current iteration number t, through iterative update infer the hidden variable Z: Among them, η t Represents the amplitude of the perturbation, the purpose of the perturbation is to minimize the KL divergence; To update the mth particle in the tth round, is the mth particle in the t+1th round; φ t (x) is a smooth function describing the disturbance direction. t ,have in, It is a customizable kernel function, which aims to measure the The distance between represents the gradient of the logarithmic posterior probability density of the latent variable Z; Use a simple kernel function, as shown in the following formula, where h is the bandwidth of the kernel function and is an adjustable parameter. Among them, k(X,Y) is the kernel function with parameters X and Y, is the square of the 2-norm of XY, for the latent variable The gradient of the logarithmic posterior probability density of Z According to the factorization formula summarized by the Bayesian model, we can get in, represents the gradient of the logarithmic posterior probability density of the latent variable Z, which is equivalent to represents the gradient of the logarithmic prior probability density of the latent variable Z, Represents the logarithmic conditional probability density of data D given the latent variable Z, After mathematical transformation, Ep(G|Z)[p(D,θ|Z)] is the expectation of the conditional probability density p(D,θ|Z) of the data D when the latent variable Z is given, calculated based on the probability p(G|Z) of the latent variable Z to the graph G.

7. The structured data simulation generation method according to claim 6, characterized in that: The step S4 further comprises: Step Y1: Calculate the sensitivity Δ of the posterior probability of a single numerical feature num and the sensitivity of the posterior probability of a single category feature Δ cat ; Step Y2: Assume that the number of categorical features in the dataset D is |C cat |, the number of numerical features is |C num |, we get the sensitivity of the log-likelihood function of the posterior probability p(Θ,D|G) when the latent variable Z is updated in a single step: Where σ is a preset constant; Step Y3: Assume that the total number of iterations is T, the number of particles is M, the number of Monte Carlo sampling is N, and the probability density function is The Laplace noise is Lap(λ), according to the serial Combination mechanism, calculate the log-likelihood function of the posterior probability p(Θ,D|G) in each round, and get the Laplace noise that needs to be added 8. The structured data simulation generation method according to claim 1, characterized in that: The step three specifically includes: Step A1: Given a feature subset C of a dataset and an edge query set Q c , privacy budget ε c , and a data set D containing m pieces of data = (x 1 , ..., x m ), according to the feature set subset C of a given data set, the marginal probability vector associated with its feature subset is defined as Where m is the number of rows in the data set, i is the data row index, and x i is the i-th data, x C is a value of the domain under the metric set C, μ C is the marginal probability vector associated with the feature subset C, and we get c -Measurement of Differential Privacy Where Lap(λ) is a probability density function Laplace noise, λ refers to the parameter of the Lap() function, which represents an arbitrary expression, ΔQ c Denotes the edge query set Q c Sensitivity under the definition of differential privacy; Step A2: For a certain edge μ, define the loss function L(x) = ||Qx-y||, and infer that the simulation data that meets the metric set conditions is converted into the optimized loss function L(x); Step A3: Use the mirror gradient descent method to minimize the loss function L(x) so that the generated simulation data meets the constraints of the association relationship as much as possible.

9. The structured data simulation generation method according to claim 8, characterized in that: In step A1, according to the serial combination mechanism, for all c - Differential privacy measure y c , the combined metric set y = (y c ) c∈C The required privacy budget for each metric y is c The sum of the budgets, that is, satisfying ε2=∑ c∈C ε c Differential privacy mechanism, the overall algorithm satisfies the ε1+ε2 differential privacy mechanism, where if the number of data is m, the sensitivity ||Q||1 is the sum of the absolute values ​​of the matrix Q along the column direction, and the maximum value is taken as the 1-norm.

10. A privacy-enhanced structured data simulation generation system, characterized in that: include: A memory, a processor and a computer program stored in the memory, wherein the computer program is configured to implement the steps of the structured data simulation generation method according to any one of claims 1 to 9 when called by the processor.

Citation Information

Patent Citations

  • Complex electromechanical system abnormal state detection method based on multi-source data

    CN111861272A

  • Medical privacy data identification method based on dynamic neural network

    CN114169007A

  • Structured simulation data generation system and generation method

    CN115169252A

  • Method and device for generating model by non-discrete data based on local differential privacy

    CN116244739A

  • Privacy-enhanced structured data simulation generation method and system

    CN117313160A

Cited By

  • Load impact type identification method and system for distributed power grid

    CN120408323A

  • Data completion method based on space-time privacy protection

    CN120579223A

  • BIM construction progress optimization method and system based on multi-agent autonomous decision

    CN120725623A

  • Differential privacy location protection method based on dynamic attenuation noise field

    CN120896973A

  • Differential privacy location protection method based on dynamic attenuated noise field

    CN120896973B