Generate High-Dimensional High-Utility Synthetic Data
Generate synthetic data in the federated learning framework through differential privacy global model and autoencoder, solving the problems of high-dimensional data privacy protection and utility, and achieving efficient data set generation and model training.
Patent Information
- Application Number
- CN202080085037.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-05-15
AI Technical Summary
The prior art is difficult to effectively collect and use high-dimensional user data in the federated learning framework, especially under the premise of retaining privacy protection, resulting in reduced data utility and increased computing costs.
The differential privacy global model is used to combine the autoencoder, and the generative autoencoder is trained on user equipment through the federated learning framework, mapping high-dimensional user data to low-dimensional feature space, and generating synthetic data using predefined distributions to avoid directly collecting original data.
While protecting user privacy, generating highly effective synthetic data sets can retain the statistical characteristics of high-dimensional data, which are suitable for classification, numerical and multimedia data, reducing the number of convergence iterations of model and improving data utility.
Smart Images

Figure CN114787826B_ABST
Abstract
Description
Technical Field
[0001] Aspects generally relate to methods for generating high-dimensional and high-utility synthetic data, and more specifically but not exclusively, to methods for generating such data in a federated learning framework in which differential privacy is applied at multiple stages of training iterations. Background Art
[0002] Services for user devices (e.g., mobile phones and smart devices) are very common. Such services can provide a large number of customized recommendations and information to users based on, for example, historical selections and / or historical data representing, for example, user profiles (e.g., age, gender, height, purchase history, etc.). The more available information there is as a reference point for users, the more accurate the personalized recommendations can be, which can improve, for example, user engagement in the service. Generally, the available information representing users and / or users' selections and preferences can be used as training data for services supported by artificial intelligence (AI) models, which are used to generate a set of personalized responses to queries from users or from the services being used.
[0003] To have any value in terms of accuracy and effectiveness, such AI services require a large amount of data from user devices for model training. With the rapid development of network and computer technologies, a wide variety of large amounts of multi-dimensional individual-specific data are generated on local devices. These data can contain rich univariate and multivariate statistical information, which can be used to build high-precision AI services. However, since these data are generated based on, for example, users' daily behaviors, direct collection may disclose sensitive information related to individuals and lead to serious privacy problems.
[0004] For this reason, local differential privacy (LDP) can be adopted to achieve privacy-preserving data collection. That is, before sending user data "outside the device" for, for example, training purposes, the user data can be randomized locally first. Broadly speaking, LDP algorithms ensure that the server used to build a model for implementing the service cannot see the original user data, but can know the overall statistics of the group. However, the LDP mechanism only supports the collection of low-dimensional data (e.g., of order about 10 dimensions), which limits the usefulness of the LDP mechanism and the utility of any information obtained from models trained using such data. Summary of the Invention
[0005] According to a first aspect, a computer-implemented method for generating high-dimensional high-utility synthetic data is provided, the method comprising: generating a differentially private global model using a global model, the differentially private global model defining an autoencoder for mapping high-dimensional user data to a low-dimensional feature space. In an implementation of the first aspect, the autoencoder may include two components: an encoder for projecting high-dimensional data into low-dimensional data, and a decoder for projecting the low-dimensional data back into high-dimensional data. The distribution of features in the low-dimensional space may be forced to follow a predefined distribution such as a standard Gaussian distribution. Once the autoencoder is trained, the decoder component can be used for data generation.
[0006] According to the first aspect, based on a plurality of differentially private local models received from a network of user devices defining a federated learning framework, the global model is iteratively optimized by: as part of an improvement iteration, broadcasting the differentially private global model to the network of user devices, receiving updated versions of the plurality of differentially private local models from the network of user devices, and using the differentially private global model to generate a set of synthetic data by, for example, based on a convergence threshold representing the convergence of the differentially private global model to a precision metric selected according to a loss function, using a predefined distribution to select a set of random latent features as input to the differentially private global model, thereby generating a set of synthetic data as output of the differentially private global model.
[0007] From a privacy perspective, federated learning using differential privacy provides a high level of protection for user data used to train a model that can be used to generate synthetic data. That is, training can be performed without collecting local (raw) data at the server. In addition, the differentially private models provided herein support the collection of high-dimensional data and subsequent generation of high-dimensional synthetic data. Generally, when considering high-dimensional data, as the number of categorical attributes increases linearly, the data domain (i.e., the number of possible combinations of all attributes) increases exponentially. In the case of a large data domain, directly randomizing the raw data (where the data domain is proportional to the length or estimation error of the anonymized data) results in significant communication costs and low data utility. Therefore, a model that can generate synthetic data without directly randomizing the raw data can be used to learn the statistical information of the raw data. The method provided herein can be applied to categorical, numerical, and multimedia data (e.g., image and video data, etc.). That is, the method can be applied to both structured data (e.g., [age, job, salary] collected from customers) and unstructured data (e.g., images, audio data, etc.). In one example, precoding and post-decoding can be used so that the autoencoder can be applied to categorical data (e.g., job).
[0008] The utility of a set of synthetic data can be evaluated using an attribute-based evaluation method. To reduce the number of iterations performed before model convergence, the evaluation of utility can be used in the mechanism for pre-tuning the model. This evaluation can include generating a measure of the divergence between the attribute distribution of the synthetic data and the attribute distribution of the real data. The utility of the set of synthetic data can be evaluated using a record-based evaluation method alone, or the record-based evaluation method can be combined with the attribute-based evaluation method to evaluate the utility of the set of synthetic data. The record-based evaluation method can include comparing the outputs of a pair of frameworks, where one framework is trained with real data (e.g., real data obtained from a public database) and the other framework is trained with synthetic data. In one example, the public database can be used to design the model framework and pre-tune the model. Federated learning and differential privacy can be used to further train the pre-tuned model according to the mechanisms described herein. Pre-tuning can reduce the number of iterations performed before model convergence. That is, the framework of the model can be designed around the framework of the public dataset. The public dataset can be used to simulate the data collection process and adjust the model parameters by evaluating the utility of the synthetic data generated using the model. In addition, the public dataset can be used to pre-train an autoencoder, which will help accelerate model convergence.
[0009] The framework according to one example includes a combination of an autoencoder, federated learning, and differential privacy, thus enabling the collection of high-dimensional data with strong privacy guarantees while preserving data utility.
[0010] In an implementation of the first aspect, multiple differentially private local models received from a user device network can be aggregated to generate a set of parameters, and these parameters can be used to update the global model. In one example, the differentially private global model is a generative autoencoder. The set of random latent features can be provided as input to the decoder of the autoencoder. The iterative improvement process can end when a convergence threshold is met, which represents the training loss associated with the differentially private global model.
[0011] For example, the global model can be initialized using a public database or a randomly generated synthetic database of user data. As described above, this initialization provides a mechanism for pre-tuning the model.
[0012] The autoencoder can follow a Gaussian distribution or a normal distribution, or any other suitable distribution (e.g., any other continuous probability distribution). Random Gaussian distribution data can be generated and input into the decoder component of the autoencoder to generate the set of synthetic data.
[0013] According to a second aspect, there is provided a user equipment forming a node in a federated learning framework. The user equipment includes a processor coupled to a memory. The processor is configured to: receive a first instance of the framework from a remote service to generate synthetic data representing a user profile, perform a modification to the first instance of the framework using local data by adjusting a set of parameters defining the first instance of the framework, thereby generating an updated instance of the framework, perform differential privacy on the updated instance of the framework to form a private local framework, and provide the private local framework to the remote service. The received first instance of the framework may define a differential privacy autoencoder. In an implementation of the second aspect, the framework may form a model that can be locally trained on user data on the user equipment. For example, differential privacy may be performed on the trained model by adding noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To more fully understand the present disclosure, reference is now made, by way of example only, to the following description taken in conjunction with the accompanying drawings, in which:
[0015] Figure 1 is a schematic representation of a method for generating high-dimensional high-utility synthetic data according to an example;
[0016] Figure 2 is according to an example of Figure 1 the method of;
[0017] Figure 3 is a schematic representation of data preprocessing and data postprocessing according to an example; and
[0018] Figure 4 is a schematic representation of the attribute distribution of exemplary data according to an example. DETAILED DESCRIPTION
[0019] The example embodiments are described in sufficient detail below so that those of ordinary skill in the art can implement and realize the systems and processes described herein. It is to be understood that the embodiments may be provided in many alternative forms and should not be construed as limited to the examples described herein. Thus, while the embodiments may be modified in various ways and take various alternative forms, the specific embodiments are shown in the drawings and described in detail below as examples. This is not limited to the particular forms disclosed. On the contrary, all modifications, equivalents, and alternatives falling within the scope of the appended claims should be included. Throughout the drawings and the appropriate detailed description, elements in the example embodiments are consistently denoted by the same reference numerals.
[0020] The terms used in this document to describe embodiments are not intended to limit the scope. The articles "a" and "the" are in the singular form and have a single referent, however, the use of the singular form herein should not exclude the existence of a plurality. In other words, unless the context clearly indicates otherwise, the elements referred to in the singular form can be one or more. It should be further understood that when the term "comprising" is used in this document, it is used to illustrate the existence of the described features, items, steps, operations, elements, and / or components, but does not exclude the existence or addition of one or more other features, items, steps, operations, elements, components, and / or their combinations.
[0021] Unless otherwise defined, all terms used in this document (including technical and scientific terms) shall be interpreted in accordance with the customs of the relevant field. It will be further understood that common terms are also interpreted in accordance with the customs in the relevant technology rather than in an idealized or overly formal sense, unless expressly defined herein.
[0022] In real-life scenarios, data collected from user devices for purposes such as training AI services includes many dimensions. In some cases, such data can include more than 100 dimensions, which makes the LDP method infeasible. For example, in the case of multi-dimensional data, randomization can be performed separately for each attribute locally (i.e., on the user device) to obtain a univariate distribution on the server side, where the data can be used for training purposes, for example. However, since the attributes in multi-dimensional data are usually correlated, applying LDP separately on a single dimension will break the possible multivariate correlations and result in information loss. If all attributes are treated as one attribute for estimating the joint distribution, only low-dimensional data (e.g., data with no more than about 10 dimensions) can be considered since the size of the domain increases exponentially with the number of dimensions. This will not only lead to a significant increase in the cost of computing and communication but may also result in a decrease in data utility.
[0023] According to one example, an autoencoder (e.g., a generative autoencoder) can be used to simulate the distribution of high-dimensional local data so as to generate reliable synthetic data. In a distributed setting, federated learning and differential privacy training can be employed to train the generative autoencoder. Thus, data synthesis can be performed without collecting real local data, which is beneficial for both privacy protection and data utility.
[0024] In one example, a differentially private global model defines an autoencoder for mapping high-dimensional user data to a low-dimensional feature space. This differentially private global model can capture the attribute distribution and correlations in high-dimensional data and can be used to generate high-utility synthetic data. The synthetic data generated in this way retains statistical characteristics similar to the real data and can be scaled up to replace the real data for data analysis and AI model training tasks. Additionally, since the generated data is entirely synthetic and cannot be associated with any specific individual, the generated data is no longer considered personal data. Re-identification attacks or attribute exposures become almost impossible.
[0025] In one example, a federated learning (FL) mechanism can be used to train local models in user devices, where the original user data stays in the local devices and cannot be accessed by a remote server. Differential privacy (DP) can be incorporated into the FL process to avoid leaking personal information during training.
[0026] In the context of this system, FL can provide a decentralized learning mechanism, which is formed by a network of user devices that define the federated learning framework. By distributing training tasks to local user devices under the coordination of a central server, FL can achieve computational efficiency and privacy advantages.
[0027] According to one example, a differentially private global model can be iteratively improved based on multiple differentially private local models received at a service, which are received from a network of user devices that define the federated learning framework. In each iteration, the service can, for example, randomly select a certain number of local user device clients and distribute the current differentially private global model to these local user device clients. Each client can train the model with its local data and return the (trained) local model update to the service, which can be performed, for example, on a remote server. On the server side, the local updates can be aggregated and the global model can be updated using the average of these local updates. The server can broadcast the updated global model to the local devices to initialize the next global round. Since only model parameters are exchanged during the training process, FL enables model training without collecting the original local data.
[0028] According to one example, a differentially private global model defines an autoencoder. An autoencoder is a neural network for learning an efficient and compressed feature representation in an unsupervised manner. The autoencoder includes two main parts or components: an encoder Q φ and a decoder G θ . The encoder maps the original high-dimensional input x ∼ P xCompress into a low-dimensional latent feature z = Q φ (x), and the decoder maps z to the reconstructed output x' = G θ (z), x' = G θ (z) has the same shape as x. The goal of training is to find a pair of optimized encoders and decoders to minimize the distance between x and x′ = G θ (Q φ (x)), that is:
[0029]
[0030] where c(·, ·) is a metric used to characterize the difference between two vectors. In one example, the mean squared error (MSE) can be used to measure the distance between numerical input vectors, and the cross entropy (CE) can be used to measure the distance between binary input vectors.
[0031] Therefore, the service executed on the server can generate synthetic data instead of directly collecting real user data, thus solving the privacy problem in data collection. In one example, the autoencoder model can be trained under the federated learning framework, so that the user data never leaves the user device, providing strong privacy protection for the user data.
[0032] During training, the user device and / or the server can apply local differential privacy and server differential privacy to ensure that information in the local training data cannot be inferred from the local model updates or the global model parameters, which further strengthens the privacy protection of the user data. The framework of the model can be flexibly modified according to different data dimensions to adapt to high-dimensional data, and the generated synthetic data retains high reliability and high utility and can be easily scaled up as an offline dataset for subsequent AI model training tasks.
[0033] Figure 1 is a schematic representation of a method for generating high-dimensional and high-utility synthetic data according to an example. In Figure 1 's example, a cloud-based server 101 is provided. The server 101 can implement services for the user device network 103, and the user device network 103 defines a federated learning framework including k user devices. For example, the services can include providing models for use by the user devices, and these models can be adjusted so that the user devices provide personalized data and services to the users.
[0034] In Figure 1In an example, model 105 represents a starting point for a method for generating high-dimensional and high-utility synthetic data. The model 105 can be pre-tuned or initialized (124) using initialization data 126, which can be random data, data from a synthetic database, or data using data from a public database, or a combination of the above. The pre-tuned model can be further trained according to the mechanisms described herein, in which federated learning and differential privacy are utilized to update the model or model parameters. Pre-tuning can reduce the number of iterations performed before the model converges. That is, in one example, the model 105 can be initialized before being sent to the user devices 103 that form the federated learning framework. The initialized model can be trained using the server 101, and then the trained model can be fine-tuned using the FL framework. The advantage of doing so is that the model can be better initialized later, and the number of training rounds using FL can be further reduced. In Figure 1 an example, data utility assessment 125 can be used to ensure the quality of the pre-trained model using assessment metrics (e.g., attribute-based assessment methods and / or record-based assessment methods). Thus, the utility assessment 125 of the pre-training and parameter adjustment of the model can be performed.
[0035] According to one example, after an iteration within the federated learning framework, the model 105 can be updated using the global model 107. As will be explained in more detail below, in Figure 1 an example, the global model 107 results from the aggregation of local model updates generated using local data in the user devices 103.
[0036] To provide a differentially private global model 105, the global model 107 can be differentially privatized (108) at the server 101, but this may not occur in certain cases as will be described in more detail below. As described above, the differentially private global model 105 defines an autoencoder for mapping high-dimensional user data to a low-dimensional feature space. The global model 107 is iteratively optimized based on a plurality of differentially private local models 109 received from the user device network 103. As part of the improved iteration, the differentially private global model 105 is broadcast to the user device network.
[0037] The user devices to which the model is broadcast may include all or a selected subset or a random subset of the k user devices 103. To generate a differentially private local model 109 at each user device that receives the broadcast model, the broadcast model 105 is used during the improvement process. More specifically, each user device generates a local model 113 using the global model 105 received from the server 101 (which may be a differentially private global model) and local (private) data 111. The local model 113 is differentially private at each user device to form a differentially private local model 109. The updated and privatized local model is sent (115) to the server 101. As will be described in more detail below, based on a convergence threshold indicating that the differentially private global model 105 converges to a precision metric selected according to a loss function, the differentially private global model 105 (in the form now referred to as the "final model" M) can be used to generate a set of synthetic data 117. More specifically, a set of random latent features 119 can be selected as the input to the final model using a predefined distribution 121 to generate the above-mentioned set of synthetic data 117. In one embodiment, the above-mentioned set of synthetic data 117 is generated using the decoder component of an autoencoder.
[0038] According to one example, privacy monitors can be provided at the local 103 and server 101 ends of the system depicted in the reference Figure 1 The privacy monitor 150 can be used to track the cost of privatizing local updates. The privacy monitor 160 can be used to track the overall cost of privatizing the global model. One way to calculate the overall privacy cost is to sum the privacy costs of all global iterations. In another example, the overall privacy cost can be determined by a subset of the user devices 103 randomly sampled in each global iteration. Thus, the overall privacy cost can be further reduced by a factor q, where q is the sampling rate. When, for example, Gaussian noise is applied to local updates, a smaller overall privacy cost can be achieved using the Moment Accountant algorithm. Thus, in one example, the privacy monitor can be used to control privacy loss. If a certain user is sampled too frequently and the privacy budget is insufficient, the suspect user device may be excluded from the iteration.
[0039] Thus, in one example, training is performed under a federated learning framework, where the model is co-trained by a central server 101 and multiple user devices 103. Training in a federated setting ensures that the raw data 111 on the user devices 103 is never sent to the server 101, effectively protecting data privacy. In addition, the user devices 103 use local differential privacy 104 to privatize local model updates, and the server 101 can use server differential privacy 108 to privatize the global model 107 to prevent the local model parameters or global model parameters from leaking information related to the privacy training data.
[0040] Figure 2 is according to the example Figure 1 is a schematic representation of the method. During each global training iteration, the server 101 broadcasts the current global model 105(1) to all user devices. The user device 103 trains the global model with local private data 111 and obtains local model parameters(2). To prevent the local model parameters from revealing the local training data, each user device applies a local differential privacy step 104(3) and returns the private local model 115 to the server(4). After that, the server 101 aggregates 123 all the received local models and updates the global model(5). Server differential privacy 108 can be used to privatize the global model parameters to further prevent information about the local data from being inferred from the global model parameters(6). The global model 105 can be shared with the user device 103 as the start of the next global training iteration(7), and (1)-(6) can be repeated until the global model reaches the accuracy threshold according to the loss function.
[0041] In one example, the user device 103 can send a local model update (relative to the local model parameters) to the server. In this case, the user device can use local differential privacy to privatize the local model update and send the privatized local model update to the server. After that, the server can aggregate all the received local model updates, use server differential privacy to privatize the aggregated model updates, and update the global model with the private model updates.
[0042] The server 101 can use the decoder component of the autoencoder to generate synthetic data 117. In one example, random latent features Z gen are extracted 118 from a distribution 121 such as a Gaussian distribution. gen The latent features Z
[0043] are fed 119 into the decoder component of the final model M and synthetic data 117 is produced as the output of the model.
[0044] According to one example, the user device 103 may process the original classified data in the form of local privacy data 111 into a numerical form, and the numerical form data can be used to train a generative model. The server 101 defines the framework of the generative model based on the dimensions of the local training data and initializes the model, and then co-trains the model between the client and the server under the differential privacy federated framework described in this article. Once the model is trained, the decoder component can be extracted to generate synthetic data, which can be converted back to the classified form and used for data mining and building machine learning models, etc.
[0045] In one example, since the original data 111 is classified, which means the original data 111 cannot be directly processed by the model, it is converted into a numerical form. In one example, one-hot encoding can be used to encode each categorical attribute into a binary vector. Each entry (also called a meta) in the binary vector represents a unique attribute value, and the entry corresponding to the given value is set to 1 while all other entries are set to 0. Finally, the binary vectors can be concatenated into a vector as the input data for the generative model.
[0046] Figure 3 is a schematic representation of data preprocessing and data postprocessing according to the example. In Figure 3 the example, 3a depicts encoding some original data (in classified form) into binary vectors (in numerical form), and 3b depicts reversing from the predicted vectors (in numerical form) to synthetic data (in classified form).
[0047] According to one example, the generative model can be a Wasserstein autoencoder (WAE). Compared with the Variational Autoencoder, the WAE provides better data synthesis ability and is easier to train than the generative adversarial network (GAN). The WAE retains the typical encoder-decoder structure of the autoencoder, which compresses the original high-dimensional input x into low-dimensional latent space features z, and then reconstructs the latent features back into the input space x'. Other suitable autoencoders can be used.
[0048] In one example, in addition to the reconstruction cost G θ (Q φ (x)), a regularization term D z (q z , p z ) is introduced into the objective function of the WAE. This regularization term measures the latent space distribution q z and a specific predefined distribution p zThe distance between them. The goal of training is to find a set of optimal parameters for the encoder and decoder to minimize the distance between the input and output while restricting the latent distribution to follow a predefined distribution. Therefore, the final objective function can be expressed as:
[0049]
[0050] where θ is a hyperparameter used to balance these two terms.
[0051] According to an example, a WAE model with fully connected hidden layers is used, and relu activation is applied to the output of each hidden layer to improve training performance. Additionally, since the input is a binary vector, sigmoid activation can be used on the output layer, which restricts the output values within [0, 1].
[0052] Cross-entropy can be used to measure the reconstruction cost c(x, G(z)), and the maximum mean discrepancy (MMD) can be used to measure the latent space distance D z (q z , p z ), where p z follows the standard Gaussian distribution.
[0053] Local DP 104 can be applied to privatize local updates. The DP mechanism can include adding, for example, Gaussian or Laplace noise to each dimension. In an example, the noise is calibrated according to the desired privacy guarantee. For example, given the l_2 sensitivity Δ, the noise strength σ of the (∈, δ)-DP Gaussian mechanism should satisfy σ ≥ Δ / ∈ √(2 ln(1.25 / δ)). The privatized local updates can then be returned to the server. On the server side, to update the global model, all local updates are aggregated. According to the post-processing property of DP, since the local updates satisfy DP, the updated global model also satisfies DP. Since local DP may be sufficient to protect local updates and the global model, server DP (108) may not be used. However, in the case of applying weak local privacy to improve model utility, server differential privacy 108 can be applied to ensure the privacy of the global model.
[0054] In an example, the server 101 can randomly select some user devices 103 (e.g., 10% of all users or 500 users, etc.) in each iteration. Since training is performed iteratively, information related to the user will be leaked each time a user device is selected for training. The amount of information leaked is controlled by ∈, which is the privacy parameter of the differential privacy process. The overall privacy cost for all iterations can be calculated. To achieve stronger privacy protection, the server 103 can delete all received model updates before sending the new model to the users for the next training iteration.
[0055] For example, during the process of iteratively training a model, in each global round t, the server 101 may randomly select n = qN clients 103 (where N is the total number of clients 103 and q is the sampling rate), and allocate the current global model 105 (M t ) to the above-mentioned clients. Each user device client i can use the local data (111)D i to train the global model for several gradient descent steps and calculate the local update After that, the client 103 can clip the local update using the clipping bound S and add a certain amount of Gaussian noise with, for example, variance σ 2 S 2 The noise-added local update (109) is returned to (115) the server 101. On the server side, all local updates are aggregated and averaged to serve as the global update 107 as follows:
[0056]
[0057] The above formula is used to update the global model. After that, the updated global model M t+1 can be allocated to the local clients to start the next iteration. Since the sum of Gaussians is still Gaussian, the calculation of the global update can be further derived as:
[0058]
[0059] That is to say, adding Gaussian noise with variance σ 2 S 2 to the individual local updates and then calculating the sum is equivalent to adding Gaussian noise with variance nσ 2 S 2 to the sum of the local updates. Therefore, this framework satisfies differential privacy. Since the moment accountant mechanism provides a tight privacy bound for the Gaussian mechanism, given the sampling rate q, the number of global rounds T, and the failure probability δ, the moment accountant mechanism can be used to keep track of the privacy loss ∈. The moment accountant mechanism is described in, for example, "Deep Learning with Differential Privacy" by M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang, Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, 2016, pp. 308 - 318. The final model satisfies (∈, δ)-differential privacy.
[0060] Once the model (e.g., WAE) is trained, server 101 can use the decoder part to generate synthetic data. Note that, in one example, the latent space features follow a standard Gaussian distribution p z . Thus, random latent features can be sampled from p z and fed into the decoder to obtain a predicted output.
[0061] Since the predicted output (i.e., the generated synthetic data) is a numerical vector, the numerical vector can be converted back to a categorical form (e.g., as shown in (b) of Figure 3 ). That is, in one example, given a predicted vector, the predicted vector is first split into multiple segments that form short vectors, and each segment represents a categorical attribute. After that, for each short vector, the entry with the largest numerical value is selected as the attribute value. Finally, all the categorical labels can be concatenated into a vector as the final synthetic data.
[0062] As described above, the utility of the generated synthetic data can be evaluated through statistical comparison and AI training performance. According to one example, in the case of statistical comparison, the statistical characteristics between the real data and the synthetic data at different privacy levels can be compared. More specifically, univariate distributions and multivariate distributions can be evaluated through chart visualization and distance calculation.
[0063] For example, in the case of univariate distribution, the frequency of each attribute of the real data and the synthetic data can be compared. In one example, the categorical data is converted to binary form, and the average value of each dimension is calculated, which provides a measure for the frequency of a specific attribute value. Bar charts can be used to visualize the frequency comparison of different data sets. That is, by plotting the attribute values, the distribution of the synthetic data can be compared with the distribution of the real data.
[0064] Figure 4 is a schematic representation of the attribute distribution of exemplary data according to the example, which is generated by the frequency of each attribute of the real data and the generated synthetic data. Figure 4 The example of
[0065] shows the comparison between the real data and the synthetic data in the pre-training scenario. JSD The frequency distance can be quantified using, for example, Jensen-Shannon divergence (JSD), which is a symmetric smoothed version of Kullback-Leibler (KL) divergence and is a distance metric. D
[0066]
[0067] where d is the total number of attributes, p i and q i are the i-th attribute distributions of the real and synthetic data, m i =(p i +q i ) / 2, |Ω i | is the domain size of the i-th attribute, and D KL is the KL divergence.
[0068] In one example, for a multivariate distribution, the correlation matrices of the real data and the synthetic data can be compared. The correlation matrix distance (CMD) can be used to measure the distance between the correlation matrices of the real data and the synthetic data as:
[0069]
[0070] R real and R syn are the correlation matrices of the real data and the synthetic data, tr(.) is the trace of the matrix, and ||.||2 is the Frobenius norm. D CMD is also bounded by [0; 1], where 0 means that the two correlation matrices are the same. For each dataset, D CMD can be calculated at different privacy levels and the results can be compared. Similarly, D CMD between the real data and the non-private synthetic data can be calculated as a baseline.
[0071] The methods described herein enable the generation of synthetic data from a model that is trained using high-dimensional categorical data from a user device, where the high-dimensional categorical data is collected using a privacy-preserving framework for high-dimensional data collection. In combination with federated learning, differential privacy, and (generative) autoencoders, the framework is capable of generating high-utility synthetic datasets without accessing the real local data. The generated synthetic data retains statistical properties very similar to the real data and can be used to replace the real data for data mining and model training tasks. For datasets that also contain numerical variables, such numerical data can be converted to categorical data using a histogram and then the framework can be applied to it. For the collection of multimedia data such as images, the loss function can be changed to, for example, the mean squared error, and the rest of the framework can remain unchanged.
Claims
1. A computer-implemented method for generating high-dimensional high-utility synthetic data, the method comprising: Generating a differentially private global model using a global model, the differentially private global model defining an autoencoder for mapping high-dimensional user data to a low-dimensional feature space; Iteratively optimizing the global model based on a plurality of differentially private local models received from a network of user devices defining a federated learning framework by: As part of an improved iteration, broadcasting the differentially private global model to the network of user devices; the user devices being configured to generate local models based on the differentially private global model and local private data and differentially privatize the local models to form differentially private local models; Receiving updated versions of the plurality of differentially private local models from the network of user devices; And Based on a convergence threshold indicating that the differentially private global model converges to a precision metric selected according to a loss function, generating a set of synthetic data using the differentially private global model by: Selecting a set of random latent features as input to the differentially private global model using a predefined distribution, thereby generating a set of synthetic data as output of the differentially private global model.
2. The method according to claim 1, further comprising: Initializing the differentially private global model using initialization data.
3. The method according to claim 2, wherein, The initialization data includes one or more of the following: random data, data from a synthetic database, and data from a public database.
4. The method according to claim 2 or 3, further comprising: Evaluating the utility of the initialized differentially private global model using an attribute-based evaluation method.
5. The method according to claim 4, further comprising: Generating a measure of the divergence between the attribute distribution of the data generated using the initialized differentially private global model and the attribute distribution of real data.
6. The method according to claim 2 or 3, further comprising: Evaluating the utility of the initialized differentially private global model using a record-based evaluation method.
7. The method according to any one of claims 1 to 6, further comprising: Aggregating the plurality of differentially private local models received from the network of user devices, thereby generating a set of parameters; And Updating the global model using the parameters.
8. The method according to any one of claims 1 to 7, wherein, The differentially private global model is a generative autoencoder.
9. The method according to claim 8, further comprising: Using a set of random latent features as input to the decoder of the generative autoencoder.
10. The method according to any one of claims 1 to 9, wherein, The convergence threshold represents the training loss associated with the differentially private global model.
11. The method according to any one of claims 1 to 10, further comprising: Using a local privacy monitor to track the cost of privatizing local updates at user devices.
12. The method according to any one of claims 1 to 11, further comprising: Using a server privacy monitor to track the cost of privatizing the global model.
13. The method according to claim 11 or 12, wherein, The privacy monitor is used to control privacy loss.
14. The method according to any one of claims 1 to 13, further comprising: Generating random data from a predefined distribution; And The generated random data is input into a decoder component of the autoencoder, thereby generating the set of synthetic data.
15. A user equipment forming a node in a federated learning framework, the user equipment including a processor coupled to a memory, the processor being configured to: Receive a first instance of a framework from a remote service to generate synthetic data representing a user profile; Perform a modification to the first instance of the framework using local data by adjusting a set of parameters defining the first instance of the framework, thereby generating an updated instance of the framework; Differentially privatize the updated instance of the framework to form a private local framework; And Provide the private local framework to the remote service.
16. The user equipment according to claim 15, wherein, The received first instance of the framework defines a differentially private autoencoder.