Privacy-preserving composition of artificial data

By using a generative adversarial network (GAN) to generate artificial data records, the customization and privacy leakage problems of privacy protection methods in existing technologies are solved, and the application of providing representative data while protecting privacy is realized.

CN120641893APending Publication Date: 2025-09-12VISA INTERNATIONAL SERVICE ASSOCIATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380093106.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-02-02
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing privacy protection methods require customization and a large amount of domain knowledge in data processing, and there is a risk of third parties leaking private data when accessing the model, making it difficult to provide useful data services while protecting privacy.

Method used

A generative adversarial network (GAN) is used for data synthesis. Initial preprocessing is used to remove outliers and ensure differential privacy. Artificial data records are generated to train machine learning models, ensuring privacy protection in any post-processing.

Benefits of technology

The generated artificial datasets can be used for data analysis and model training without exposing the privacy of real data, providing representative data for various applications while meeting differential privacy requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120641893A_ABST
    Figure CN120641893A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to methods and systems for generating artificial data records from (potentially privacy or sensitive) data records in a privacy-preserving manner, particularly using machine learning models such as generative adversarial networks (GANs). Such artificial data records may be used in lieu of real data in data analysis applications, such as training machine learning models. These artificial data records may be generated such that they do not leak information from the data records used to generate the artificial data records (or have low or negligible probabilities). Thus, artificial data records (or any machine learning model trained to generate such artificial data records) may potentially be published or distributed without violating rules, regulations, or laws that limit the transmission of sensitive data.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Data can be analyzed for a variety of useful purposes. For example, user data collected by a streaming service can be used to recommend shows or movies to users. For a specific user, the streaming service can identify other similar users (e.g., based on common user data characteristics) and identify shows or movies watched by those similar users. These shows or movies can then be recommended to the user. Given the similarities between users, users are more likely to like the recommended shows, and therefore, the streaming service provides a useful recommendation service to users. As another example, a bank can use user transaction data to generate a model of user purchasing patterns. Such a model can be used to detect fraudulent purchases, such as purchases made using stolen credit cards. The bank can use this model to detect whether a fraudulent purchase is being made, alert the cardholder, and deactivate the stolen card. In this way, user data can be used to provide useful services to users.

[0002] In some cases, data may be considered sensitive or confidential, for example, containing information that the data subject (e.g., user) may not wish to disclose or otherwise make publicly available. Recent concerns about data privacy have led to the widespread adoption of data privacy rules and regulations. Governments and organizations now generally restrict the use, storage, and transmission of data, especially the transmission of user data across national borders. While these regulations expand and protect individual privacy rights, they limit the ability to use user data to provide useful services (e.g., the services described above).

[0003] In some cases, such rules and regulations can cause problems. As an example, a country may experience a severe viral outbreak that requires help from the larger global community. The country may collect biodata from its citizens who are infected, which can be used to research treatments or vaccines. The country's laws (or the laws of a larger economic, defense, or administrative partnership to which the country belongs) may prohibit the transmission of this sensitive biodata outside the country, thereby slowing down research and development of treatments or vaccines.

[0004] Some recent scientific literature has proposed privacy protection methods for data processing, including the creation of complex machine learning models. These methods can be used to analyze sensitive data without violating privacy. However, these solutions have several problems that make them less useful in practice. As an example, these solutions usually need to be customized for their specific use cases (e.g., financial data analysis, health data analysis, etc.), and usually require developers who implement such models to have a lot of domain knowledge. The developers of these models also need to have a strong background in correctly applying privacy techniques not only during model deployment but also during the model development process itself. If privacy techniques are not applied (or even applied incorrectly), third parties who access the model (e.g., via an application programming interface) can use their access rights to sensitive information about the private data used to train the model.

[0005] The embodiments address these and other problems individually and collectively. Summary of the Invention

[0006] Embodiments relate to methods and systems for synthesizing privacy-preserving artificial data. Embodiments may use a machine learning model, such as a generative adversarial network (GAN), to accomplish this synthesis. In some embodiments, a data synthesizer (e.g., implemented using a computer system) may perform initial preprocessing operations on potentially sensitive or private input data records to remove outliers and ensure "differential privacy" (a concept described in more detail below). This input data can be used to train a machine learning model (e.g., a GAN) in a privacy-preserving manner to generate artificial data records, which generally represent the input data records used to train the model. These artificial data records protect the privacy of these input data records.

[0007] The trained generator model can then be used to generate an artificial dataset, which can be transmitted to a client computer or published. Alternatively, the trained generator model itself can be published or transmitted. In an embodiment, the privacy guarantees can be strong enough that privacy is preserved even under "arbitrary post-processing". That is, the client computer (or its operator) can process the data as it sees fit without risking the privacy of any sensitive data records used to train the generator model. For example, the client computer can use the artificial data to train a machine learning model to perform some form of classification (e.g., classify credit card transactions as normal or fraudulent). When performing arbitrary data analysis on the artificial dataset or using the trained generator to generate the artificial dataset, the operator of the client computer does not need to have any familiarity with privacy protection techniques, standards, etc.

[0008] In more detail, one embodiment relates to a method, performed by a computer system, for training a machine learning model to generate multiple artificial data records in a privacy-preserving manner. The computer system may retrieve multiple data records (e.g., from a database). Each data record may include multiple data values ​​corresponding to multiple data fields. Each data field may be within a category from a plurality of categories. The computer system may determine multiple noise class counts corresponding to the multiple categories. Each noise class count may indicate an estimated number of data records in the multiple data records that belong to each category from the multiple categories. A client computer may use the multiple noise class counts to identify one or more defect categories. Each defect category may include a category whose corresponding noise class count is less than a minimum count. The computer system may combine each defect category from the one or more defect categories with at least one other category from the plurality of categories. In this manner, the computer system may determine multiple combined categories. The computer system may identify one or more defective data records. Each defective data record may contain at least one defective data value corresponding to the combined category. For each defective data value contained in one or more defective data records, the computer system may replace the defective data value with a combined data value that identifies a combined category from a plurality of combined categories; the computer system may generate a plurality of condition vectors, such that each condition vector identifies one or more specific data values ​​for one or more specific data fields from a plurality of data fields. The computer system may sample a plurality of sampled data records from the plurality of data records. The plurality of sampled data records may include at least one of the one or more defective data records. Each sampled data record may include a plurality of sampled data values ​​corresponding to a plurality of data fields. The computer system may then train a machine learning model to generate a plurality of artificial data records. Each artificial data record may include a plurality of artificial data values ​​corresponding to a plurality of data fields. Based on the condition vectors, the machine learning model may generate the plurality of artificial data records, such that the machine learning model may replicate the one or more sampled data values ​​corresponding to the one or more specific data fields from the plurality of artificial data values.

[0009] Another embodiment relates to a method for training a machine learning model to generate a plurality of artificial data records that protect the privacy of sampled data values ​​contained in the plurality of sampled data records. The method may be performed by a computer system. The computer system may obtain a plurality of sampled data records, each sampled data record comprising a plurality of sampled data values. The computer system may also obtain a plurality of condition vectors, each of which may identify one or more specific data fields. The computer system may then perform an iterative training process comprising several steps described in further detail below. The computer system may determine one or more selected sampled data records from the plurality of sampled data records. Similarly, the computer system may determine one or more selected condition vectors from the plurality of condition vectors. The computer system may then identify one or more condition data values ​​from the one or more selected sampled data records, the one or more condition data values ​​corresponding to the one or more specific data fields identified by the one or more selected condition vectors. The computer system may generate the one or more artificial data records using the one or more condition data values ​​and a generator sub-model. The generator sub-model may be characterized by a plurality of generator parameters. The computer system may generate one or more comparisons using the one or more selected sampled data records, the one or more artificial data records, and a discriminator sub-model. Similar to the generator sub-model, the discriminator sub-model can be characterized by multiple discriminator parameters. The computer system can determine a generator loss value and a discriminator loss value based on one or more comparisons. The computer system can generate one or more generator update values ​​based on the generator update value. Similarly, the computer system can generate one or more initial discriminator update values ​​based on the discriminator loss value. The computer system can generate one or more discriminator noise values ​​and generate one or more noise discriminator update values ​​by combining the one or more initial discriminator update values ​​with the one or more discriminator noise values. The computer system can update the generator sub-model by updating multiple generator parameters using the one or more generator update values. The computer system can also update the discriminator sub-model by updating multiple discriminator parameters using the one or more noise discriminator update values. The computer system can determine whether a termination condition is met, and if the termination condition is met, the computer system can terminate the iterative training process; otherwise, the computer system can repeat the iterative training process until the termination condition is met.

[0010] Other embodiments relate to computer systems, non-transitory computer-readable media, and other apparatus that can be used to implement the methods described above or other methods according to embodiments.

[0011] the term

[0012] A "server computer" may refer to a computer or a cluster of computers. A server computer may be a powerful computing system, such as a mainframe. A server computer may also include a cluster of smaller computers or a group of servers operating as a unit. In one example, a server computer may include a database server coupled to a network server. A server computer may include one or more computing devices and may use any of a variety of computing structures, arrangements, and compilations to service requests from one or more client computers.

[0013] A "client computer" may refer to a computer or cluster of computers that receives some service from a server computer (or another computing system). The client computer may access this service via a communication network, such as the Internet or any other suitable communication network. The client computer may make a request, including a request for data, to the server computer. As an example, a client computer may request a video stream from a server computer associated with a movie streaming service. As another example, a client computer may request data from a database server. The client computer may include one or more computing devices and may use various computing structures, arrangements, and compilations to perform its functions, including requesting and receiving data or services from a server computer.

[0014] "Memory" may refer to any suitable device or devices that can store electronic data. Suitable memory may include non-transitory computer-readable media that stores instructions that can be executed by a processor to implement the desired method. Examples of memory include one or more memory chips, disk drives, and the like. Such memory may operate using any suitable electrical, optical, and / or magnetic operating modes.

[0015] A "processor" may refer to any suitable data computing device or devices. A processor may include one or more microprocessors working together to perform a desired function. A processor may include a CPU comprising at least one high-speed data processor sufficient to execute program components for executing user and / or system-generated requests. A CPU may be a microprocessor such as an Athlon, Duron, and / or Opteron from AMD; a PowerPC from IBM and / or Motorola; a Cell processor from IBM and Sony; a Celeron, Itanium, Pentium, Xenon, and / or XScale from Intel; and / or similar processors.

[0016] A "message" may refer to any information that can be transmitted between entities. A message may be transmitted from a "sender" to a "receiver," such as from a server computer (sender) to a client computer (receiver). A sender may refer to the originator of a message, and a recipient may refer to the recipient of a message. Most forms of digital data can be represented as messages and transmitted between a sender and a recipient over a communication network, such as the Internet.

[0017] "User" can refer to an entity that uses something for a certain purpose. An example of a user is a person who uses a "user device" (e.g., a smartphone, wearable device, laptop, tablet, desktop computer, etc.). Another example of a user is a person who uses a certain service, such as a member of an online video streaming service, a person who uses a tax preparation service, a person who receives health care from a hospital or other organization, etc. A user can be associated with "user data," which is data that describes the user or their use of something (e.g., their use of a user device or service). For example, user data corresponding to a streaming service may include a username, email address, billing address, and any data corresponding to their use of the streaming service (e.g., how often they watch videos using the streaming service, the types of videos they watch, etc.). Some user data (and data in general) may be private or potentially sensitive, and users may not want such data to become publicly available. Some user data (and data in general) may be protected by privacy rules, regulations, and / or laws that prevent its transmission.

[0018] A "dataset" may refer to a collection of related information (e.g., "data") that may include individual data elements and that can be manipulated and analyzed, for example, by a computer system. A dataset may include one or more "data records," which typically correspond to smaller sets of data about a particular event, individual, or observation. For example, a "user data record" may contain data corresponding to a user of a service (e.g., a user of an online image hosting service). A dataset, or the data contained therein, may be derived from a "data source," such as a database or data stream.

[0019] "Tabular data" may refer to a data set or collection of data records that may be represented in a "data table," such as an ordered list of rows and columns of "cells." A data table and / or data records may contain any number of "data values," individual elements, or observations of data. Data values ​​may correspond to "data fields," which are tags that indicate the type or meaning of a particular data value. For example, a data record may contain a "name" data field and an "age" data field, which may correspond to data values ​​such as "John Doe" and "59." A "numeric data value" may refer to a data value represented by a number. A "normalized numerical data value" may refer to a numerical data value that has been normalized to a certain defined range. A "categorical data value" may refer to a data value that represents a "category," which is a classification or division of things based on shared characteristics.

[0020] An "artificial data record" or "synthetic data record" may refer to a data record that does not correspond to a real event, individual, or observation. As an example, while a user data record may correspond to a real user of an image hosting service, an artificial data record may correspond to an artificial user of the image hosting service. Artificial data records may be generated based on real data records and may be used in many of the same contexts as real data records.

[0021] A "machine learning model" may refer to a file, program, software executable, instruction set, etc. that has been "trained" to recognize patterns or make predictions. For example, a machine learning model may take as input transaction data records and classify each transaction data record as corresponding to a legitimate transaction or a fraudulent transaction. As another example, a machine learning model may take as input weather data and predict whether it will rain later this week. A machine learning model may be trained using "training data" (e.g., to identify patterns in the training data) and then applied when using this training for its intended purpose. A machine learning model may be defined by "model parameters," which may include numerical values ​​that define how the machine learning model performs its function. Training a machine learning model may include an iterative process for determining a set of model parameters that achieve optimal performance for the model.

[0022] "Noise" can refer to irregular, random, or pseudo-random data that can be added to a signal or data to obscure it. Noise can be intentionally added to data for a purpose, for example, visual noise can be added to an image for artistic reasons. Noise can also exist naturally for some signals. For example, Johnson-Nyquist noise (thermal noise) includes electronic noise generated by thermal agitation of charge carriers in an electrical conductor. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1A system block diagram outlining an exemplary use case of some embodiments of the present disclosure is shown.

[0024] Figure 2 Some examples of data sets, data records, and condition vectors are shown according to some embodiments.

[0025] Figure 3 A system block diagram of an exemplary artificial data synthesizer is shown in accordance with some embodiments.

[0026] Figure 4 A flow chart illustrating an exemplary method corresponding to synthesizing artificial data records according to some embodiments is shown.

[0027] Figure 5 A flow chart corresponding to an exemplary data pre-processing method according to some embodiments is shown.

[0028] Figure 6 A graph showing an exemplary multimodal distribution that can be used to assign categories to normalized data values.

[0029] Figure 7 A diagram is shown detailing an exemplary method of category combination according to some embodiments.

[0030] Figure 8 A diagram is shown detailing an exemplary method of assigning a single category to data records associated with combined categories in accordance with some embodiments.

[0031] Figures 9A to 9B A flowchart illustrating an exemplary method for training a machine learning model to generate artificial data records in a privacy-preserving manner, according to some embodiments.

[0032] Figure 10 An exemplary computer system is shown in accordance with some embodiments. DETAILED DESCRIPTION

[0033] As described above, embodiments relate to methods and systems for synthesizing artificial data records in a privacy-preserving manner. Briefly, a computer system or other device can instantiate and train a machine learning model to generate these artificial data records. As an example, this machine learning model can include a generative adversarial network (GAN), an autoencoder (e.g., a variational autoencoder), a combination of the two, or any other suitable machine learning model. The machine learning model can be trained using (potentially sensitive or privacy-preserving) data records to generate artificial data records. After training, a trained "generator model" (or "generator sub-model") that can include a portion of the machine learning model (e.g., a portion of a GAN) can be used to generate artificial data records.

[0034] Figure 1A system block diagram generally illustrating a use case of an embodiment of the present disclosure is shown. An artificial data generating entity 102 may possess a real data set 106 containing (potentially sensitive or private) data records. The real data set 106 may include, for example, private medical records corresponding to an individual. The real data set 106 may be subject to privacy rules or regulations that prevent the transmission or release of these data records. These data records may be potentially useful to an artificial data consuming entity 104, which may include, for example, a public health organization or pharmaceutical company that wants to use private medical records to research the cure or treatment of a disease. However, due to the aforementioned rules or regulations, the artificial data generating entity 102 may not be able to provide the real data set 106 to the artificial data consuming entity 104.

[0035] Alternatively, the artificial data generating entity 102 may use an artificial data synthesizer 108, which may include a machine learning model that is instantiated, trained, and executed by a computer system (e.g., a server computer or any other suitable device) that is owned and / or operated by the artificial data generating entity 102. This machine learning model may include, for example, a generative adversarial network (GAN), an autoencoder (e.g., a variational autoencoder), a combination of these two models, or any other suitable machine learning model. Using the real dataset 106 as training data, the artificial data synthesizer 108 may be trained to generate an artificial dataset 110 that is generally representative of the real dataset 106 but protects the privacy of the real dataset 106.

[0036] Instead of sharing the real data set 106 with the artificial data consuming entity 104, the artificial data generating entity may transmit the artificial data set 110 to the artificial data consuming entity 104, or alternatively publish the artificial data set 110 in a manner that enables the artificial data consuming entity 104 to access the artificial data set 110. Alternatively, the artificial data generating entity 102 may transmit the trained artificial data synthesizer 108 itself to the artificial data consuming entity 104, so that the artificial data consuming entity 104 can optionally generate its own artificial data set 110 using a set of data generation parameters 114.

[0037] As an example, the real data set 106 may correspond to user data records corresponding to users of an online streaming service that relies on advertising revenue. These data records may correspond to various users belonging to various demographics. The artificial data usage entity 104 may include an advertising company contracted by a company to promote products to women aged 35 to 45. This advertising company may use data generation parameters 114 to instruct the artificial data synthesizer 108 to generate an artificial data set 110 corresponding to (artificial) women aged 35 to 45. The advertising company can then review this artificial data set 110 to determine which shows and movies those artificial women "watch" in order to determine when to promote products.

[0038] The artificial data generating entity 102 and the artificial data using entity 104 may each own and / or operate their own respective computer systems, which may enable the two entities to communicate via a communication network (not shown) (e.g., a cellular communication network or the Internet). However, it should be understood that such a communication network may take any suitable form and may include any one and / or combination of the following: direct interconnection; the Internet; a local area network (LAN); a metropolitan area network (MAN); an operating mission as a node on the Internet (OMNI); a secure custom connection; a wide area network (WAN); a wireless network (e.g., using protocols such as, but not limited to, Wireless Application Protocol (WAP), I-mode, etc.); etc. Figure 1 Messages between computers and devices in the system (containing, for example, an artificial data set 110 or an artificial data synthesizer 108) may be transmitted using a communication protocol such as, but not limited to: File Transfer Protocol (FTP); Hypertext Transfer Protocol (HTTP); Secure Hypertext Transfer Protocol (HTTPS); Secure Sockets Layer (SSL), ISO (e.g., ISO 8583), and the like.

[0039] After receiving or generating the artificial dataset 110, the artificial data-using entity 104 can use the artificial dataset 110 as needed without risking exposure to private data from the real dataset 106. This can include performing any data analysis process 112 (e.g., determining frequently watched programs for a particular demographic, as described above), or training a machine learning model during a model training process 116. For example, the real dataset 106 (and the artificial dataset 110) can correspond to genetic data. A health organization may wish to train a machine learning model (using the artificial dataset 110) to detect genetic markers for predicting diseases that may occur during an individual's lifetime. After training, the artificial data-using entity 104 can provide input data 120 (e.g., genetic samples of an infant or fetus) to the trained model 118 to generate output results 122 (e.g., estimates of the likelihood of different diseases).

[0040] As described in more detail below, there are generally some risks, such as the artificial data synthesizer 108 of the artificial data synthesizer inadvertently leaking information related to the data used to train the model, such as data records from the real data set 106. For example, a generator model for generating artificial user profiles (e.g., corresponding to a social network or streaming service) may inadvertently learn to copy private information from the training data (e.g., the user's name or address), and may therefore inadvertently leak this private information. To address this, the method according to the embodiment introduces several novel features that enable "differential privacy" training of the artificial data synthesizer 108 used to generate the artificial data set 110. Generally speaking, differential privacy refers to a specific mathematical definition of privacy related to the risk of data exposure. These novel features can be better understood in the context of differential privacy, machine learning, and other related concepts. Therefore, it may be useful to describe such concepts before describing the method and system according to the embodiment in more detail.

[0041] I. Privacy of Artificial Data in Machine Learning

[0042] A. Artificial Data

[0043] As stated above, artificial data can be useful in many of the same contexts as "real" data. For example, a movie recommender machine learning model can be trained to recommend movies based on the preferences of an "artificial user" (represented by the artificial data) rather than using real user data. Similarly, a machine learning model for detecting fraudulent credit card transactions can be trained using artificial transaction data rather than real transaction data. Artificial data is useful if the artificial dataset adequately represents any corresponding real dataset. That is, while a particular artificial data record preferably does not contain any data corresponding to a real data record (thereby protecting privacy), the entire artificial dataset preferably accurately represents all data records collectively. This enables the artificial dataset to be used for further data processing (e.g., training a machine learning model to identify risk factors for a disease, performing market analysis, etc.).

[0044] However, there is often an implicit trade-off between privacy and the “representativeness” of artificial data. Artificial data that is effective in protecting privacy is often not very representative. As an example, a machine learning model can generate random artificial data that is completely uncorrelated with any real data used to train the machine learning model. Due to its uncorrelated randomness, this artificial data cannot leak any information from the real data. However, it is also completely unrepresentative of the real data used to train the model. On the other hand, highly representative artificial data is often not very good at protecting privacy. As an example, a machine learning model can generate artificial data that is an exact copy of the real data used to train the machine learning model. Such artificial data will perfectly represent the real data used to generate it and will be very useful for any further analysis. However, this artificial data does little to protect the privacy of the real data.

[0045] Privacy-focused machine learning models can be designed with this trade-off in mind. One strategy for overcoming this trade-off is to determine an acceptable level of privacy or privacy risk, and then train the machine learning model to generate artificial data that is as representative as possible while still meeting the determined privacy level or privacy risk. To this end, techniques for quantifying privacy or privacy risk, such as differential privacy, can be useful.

[0046] B. Data Records and Tabular Data

[0047] Embodiments of the present disclosure are well suited for generating tabular data, particularly sparse tabular data. Tabular data generally refers to data that can be represented in a tabular form, such as an organized array of rows and columns of "fields," "cells," or "data values." An individual row or column (or even a collection of rows, columns, cells, data values, etc.) may be referred to as a "data record." Sparse data generally refers to data records in which most data values ​​are equal to zero, e.g., non-zero data values ​​are rare. Sparse data may occur when data records cover a large number of data values, some of which may not apply to all individuals or objects corresponding to those data records. For example, a data record corresponding to personal property may have data fields for car ownership, boat ownership, airplane ownership, and the like. Because most individuals do not own boats or airplanes, these data fields may typically have corresponding data values ​​of zero. Some data representation techniques, such as "one hot encoding," may also produce sparse data.

[0048] Figure 2An exemplary tabular data set 202, exemplary data records 212, and two exemplary formulas illustrating condition vectors 214 and 216 may be helpful in understanding embodiments of the present disclosure. Tabular data set 202 may include data corresponding to users of an internet service. Tabular data set 202 may be organized so that each column corresponds to an individual data record, and each row corresponds to a specific data field within those data record columns. For example, exemplary data field 204 may correspond to the user's age, data usage metric, and service plan category. Tabular data set 202 may include numerical data values ​​206, such as the user's actual age (e.g., 37), and normalized numerical data values ​​208, such as the user's data usage normalized between values ​​0 and 1. Normalized numerical data value 208 (e.g., 0.7) may indicate that the user has used 70% of their allotted data for the month, or is in the 70th percentile in data usage. Additionally, data values ​​may be categorized. For example, category data value 210 may indicate the name or tier of the user's service plan, which may be from a limited set of possible service plans.

[0049] In some cases, data values ​​may identify or be within a category. Category data value 210 "directly" identifies a service plan category (gold) among a possibly finite number of possible categories (e.g., gold, silver, bronze). However, data values ​​may also "indirectly" identify a category, for example, based on a mapping between numerical data values ​​(or normalized numerical data values) and categories. For example, numerical data value 206 (age = 37) may correspond to (or be within) a category such as "adult," and normalized numerical data value 208 (data usage = 0.7) may correspond to (or be within) a category such as "high usage." Categories may be determined based on data values ​​using any suitable means or technique. One particular technique for determining categories or assigning categories to data values ​​is to use Gaussian mixture modeling, as described below with reference to FIG. Figure 6 Further description.

[0050] Throughout this disclosure, instance categories are typically "semantic" categories, such as "children," "teenagers," "young adults," "adults," "middle-aged people," "elderly people," and the like. However, it will be understood that these exemplary categories are chosen because they are generally more easily understood by humans, and it is relatively easy for humans to determine (for example) how to assign a numerical data value such as age to one of these particular categories. However, embodiments of the present disclosure may be practiced with any form of categories and any means of identifying categories, and are not limited to such semantic categories. For example, a range of normalized data values, such as 0.1 to 0.5, may correspond to a particular category, while different random numbers of normalized data values ​​(e.g., 0.51 to 0.57) may correspond to different categories. These categories do not require names or labels to exist.

[0051] Exemplary data records 212 may include columns of data from the tabular dataset 202 corresponding to a particular user (e.g., Duke). Such data records may be sampled from the tabular dataset and used as training data to train a machine learning model (e.g., from Figure 1 108) to generate artificial data records. As described in more detail below, condition vectors (e.g., condition vectors 214 and 216) can be used for this purpose. Condition vector 214 can be used to indicate specific data fields and data values ​​to be reproduced when generating artificial data records during training. The condition vectors can indicate these data fields in a number of ways, and Figure 2 The two examples provided in are intended only as non-limiting examples. As an example, condition vector 214 can include a binary vector in which each element has a value of 0 or 1. A value of 0 can indicate that a corresponding data value in a data record (e.g., data record 212) can be ignored during artificial data generation. A value of 1 can indicate that a corresponding data value in a data record should be copied during artificial data generation in order to train a model to generate an artificial data record representing a data record containing the particular data value. Exemplary condition vector 216 includes a list of "instructions" indicating whether a corresponding data value can be ignored ("N / A") or copied ("Copy") during artificial data generation.

[0052] It should be understood that Figure 2 Only one example of tabular data is described and is intended only for the purpose of illustrating and introducing concepts or terminology that may be used throughout this disclosure. Other forms of tabular data or data records may be used to practice or implement embodiments of the present disclosure. For example, instead of representing data records as columns and data fields as rows, a tabular data set may represent data records as rows and data fields as columns. Tabular data sets do not need to be two-dimensional (e.g., Figure 2 ), but can be any number of dimensions. Furthermore, individual data values ​​need not be numerical or categorical, as Figure 2 For example, a data value may include any form of data, such as data representing an image or video, a pointer to another data table or data value, another data table itself, etc.

[0053] C. Differential Privacy

[0054] Differential privacy refers to a rigorous mathematical definition of privacy, which is outlined in extensive detail below. More information on differential privacy can be found in [1]. In the context of differential privacy, the method or process The privacy of a data set can be characterized by one or two privacy parameters: ε (epsilon) and δ (delta). These privacy parameters generally relate to the privacy of information from a particular data record in a method or process. The probability of leakage during the period. In general, smaller values ​​of ε and δ lead to greater privacy guarantees at the expense of reduced accuracy or representativeness, while large values ​​of ε and δ lead to the opposite.

[0055] Differential privacy can be particularly useful because the privacy of a process can be defined independently of the dataset on which the process is running. In other words, the (ε, δ) pair in particular generally has the same meaning regardless of whether the process is used to analyze healthcare data, train a machine learning model to generate artificial user accounts, etc. However, different ε and δ values ​​may be appropriate or desired in different contexts. For very sensitive data (e.g., the name, residential address, and social security number of a real individual), very small ε and δ values ​​may be required because the consequences of leaking such information can be significant. For less sensitive data (e.g., the number of hours of movies streamed in the past week), larger ε and δ values ​​may be acceptable.

[0056] In a broad sense, if we look at the process If it is not possible to determine whether a particular data record is included in the data set that was input to the process, then the method or process (which treats a dataset as input) is differentially private. Differential privacy can be defined based on two hypothetical "adjacent" datasets d and d', one containing a specific data record and one not containing the specific data record. If the output and similar or similarly distributed, then the method or process is differentially private. The more similar (or similarly distributed) the two outputs are, the higher the privacy. If the two outputs are identical or identically distributed, it may be impossible to distinguish which dataset d or d' was used as the process Therefore, it may not be possible to determine whether a particular data record is an input to a process input, thereby protecting the privacy of that particular data record (or, for example, an individual associated with that data record).

[0057] More technically, having a domain and scope Method (ε,δ)-differential privacy is satisfied if for any two adjacent inputs And for any output subset It believes that:

[0058] Based on this formula, it can be shown that as ε and / or δ increase, it becomes easier to satisfy the inequality, regardless of and On the other hand, for ε = 0 and δ = 0, the inequality requires a more stringent condition Therefore, lower ε and δ values ​​generally correspond to more stringent privacy requirements, while higher ε and δ values ​​generally correspond to less stringent privacy requirements.

[0059] In the context of the embodiments, the hypothetical method or process Similar to the process used to train machine learning models for generating artificial data records. Output A trained model (or even artificial data records generated using the trained model) may be included, and the input d may include the (potentially sensitive or private) data records used to train the machine learning model. The parameters ε and δ may be chosen so that the risk of the trained model leaking information about any given training data record is acceptably low.

[0060] 1. Implementing differential privacy

[0061] In general, differential privacy can be implemented by adding noise (e.g., random numbers) to a process that is not originally differentially private (and which can be deterministic). If the noise is large enough, it may be impossible to determine the "deterministic output" of the process based on the noisy output, thereby protecting differential privacy. The effect of protecting privacy by adding noise is illustrated by the following example. In this example, an individual can query a "privacy" database to determine the average income of 10 people (e.g., $50,000), including someone named Alice. Although this statistic alone is not enough to determine the income of any individual (including Alice), the individual can query the database multiple times to learn Alice's income, thereby violating Alice's privacy. For example, an individual can query the database to determine the average income of 9 people (everyone except Alice, e.g., $45,000), and then use the difference between the two results to determine Alice's income ($95,000), thereby violating Alice's privacy.

[0062] However, if enough noise is added to the average income statistics, it may no longer be possible to determine Alice's income using this technique, thereby protecting Alice's privacy. If, for example, random noise between -$5000 and $5000 is added to each of these statistics, Alice's calculated income could be anywhere between $0 and $190,000, which doesn't provide much information about Alice's actual income. At the same time, adding noise between -$5000 and $5000 only distorts the average income statistics by at most about 11.2%, meaning that the statistics are still fairly representative of the actual average income.

[0063] Much like how noise is added to achieve differential privacy during this database query example, it is also possible to add noise to achieve differential privacy when training a machine learning model. This process is outlined in considerable detail below and described in more detail in [2]. However, before describing differentially private machine learning, it may be useful to describe some machine learning concepts in more detail.

[0064] D. Machine Learning Model

[0065] This section describes some machine learning concepts at a high level and is intended to orient the reader, as well as introduce some terminology that may be used throughout this disclosure (e.g., "model parameters," "noise discriminator update values," etc.). However, it is generally assumed that those skilled in the art are already familiar with these concepts to some extent (excluding concepts related to novel aspects of the embodiments). As an example, it is assumed that those skilled in the art understand the meaning of statements such as "backpropagation can be used to update the weights of a neural network" or "gradients can be clipped before updating model parameters," without requiring a detailed explanation of how backpropagation or gradient clipping is performed.

[0066] As a high-level overview, a machine learning model can be characterized by "model parameters," which, in part, determine how the machine learning model performs its function. For example, in an artificial neural network, model parameters can include neural network weights. Two identical machine learning models, aside from their model parameters, are likely to produce different outputs. For example, two artificial data generators with different model parameters can produce different artificial data records. Broadly speaking, training a machine learning model can involve an iterative training process for refining or optimizing the model parameters that define the machine learning model.

[0067] In each training round, a "model update value" can be determined based on the performance of the model, and these model update values ​​can be used to update the model parameters. For example, for a neural network, these model update values ​​may include gradients for updating the weights of the neural network. These model update values ​​can be widely determined by evaluating the current performance of the model during each training round. A "loss value" or "error value" corresponding to the difference between the expected or ideal performance of the model and its actual performance can be calculated. For example, a binary classifier machine learning model can learn to classify data as belonging to one of two categories (e.g., legitimate and fraudulent). If the training data for the binary classifier is labeled, the loss or error value can be determined based on the difference between the classification of the machine learning model and the actual classification given by the label. If the binary classifier correctly labels the training values, the loss value may be small or zero. If the binary classifier incorrectly labels the training values, the loss value may be large. In general, for a classifier, the more similar the classification and labeling are, the lower the loss value may be.

[0068] Such loss values ​​can be used to generate model update values, which can be used to update the machine learning model parameters. In general, for the purposes of this explanation, it is assumed that large model update values ​​can result in large changes in the machine learning model parameters, while small model update values ​​can result in small changes in the machine learning model parameters. If the loss value is small or zero, it can indicate that the model parameters are generally effective for whatever task the machine learning model is learning to perform. Therefore, the model update value can be small. If the loss value is large, it can indicate that the machine learning model is performing poorly at its intended task and that large changes to the model parameters may be required. Therefore, the model update value can be large.

[0069] A "termination condition" can define a point at which training is complete and the model can be validated and / or deployed for its intended purpose (e.g., generating artificial data records). Some machine learning systems are configured to train for a specific number of training rounds or epochs, at which point training is complete. Some other machine learning systems are configured to train until the model parameters "converge," that is, the model parameters no longer change (or change only slightly) in consecutive training rounds. The machine learning system can periodically check whether the termination condition has been met. If the termination condition has not been met, the machine learning system can continue training, otherwise the machine learning system can terminate the training process. Ideally, once training is complete, the trained model parameters enable the machine learning model to effectively perform its intended task.

[0070] 1. Generative Adversarial Networks

[0071] Some embodiments of the present disclosure may use a generative adversarial network (GAN) to generate artificial data records. One advantage of GANs over other data generation models is that using GANs generally does not require any expert knowledge of the underlying training data. GANs can be better understood by referring to [3], but are summarized below to orient the reader. GANs generally include two sub-models: a generator sub-model and a discriminator sub-model. For a GAN, "model parameters" may refer to both a set of "generator parameters" that define the generator sub-model and a set of "discriminator parameters" that define the discriminator sub-model. Broadly speaking, the role of the generator may be to generate artificial data. The role of the discriminator may be to distinguish between artificial data generated by the generator and training data. Typically, training a GAN involves training these two sub-models approximately simultaneously. The generator loss value and the discriminator loss value may be used to determine a generator update value (for updating the generator parameters to improve the generator performance) and a discriminator update value (for updating the discriminator parameters to improve the discriminator performance).

[0072] The generator loss value and the discriminator loss value can be based on the performance of the generator and the discriminator at their respective tasks. If the generator is able to successfully "fool" the discriminator by generating artificial data that the discriminator cannot identify as artificial, the generator may produce a small or zero generator loss value. However, the discriminator may produce a high loss value due to failure to identify artificial data. Conversely, if the generator cannot fool the discriminator, the generator may produce a large generator loss, and the discriminator may produce a small or zero loss due to successful identification of artificial data.

[0073] In this way, the discriminator puts pressure on the generator to generate more convincing artificial data. As the generator improves at generating artificial data, it puts pressure on the discriminator to better distinguish real data from artificial data. This "arms race" ultimately ends with a trained generator that effectively generates convincing or representative artificial data. This trained generator can then be used to generate artificial datasets that can be used for some purpose (e.g., for analysis without violating privacy rules, regulations, or legislation).

[0074] 2. Differential Privacy and Machine Learning

[0075] As described above, differential privacy is typically achieved by adding noise to a method or process. To achieve differential privacy in a machine learning model, noise can be added to the model update values ​​during each training round. For example, for a neural network, the gradient can be calculated using stochastic gradient descent. The gradient can then be clipped, and noise can be added to the clipped gradient. This technique is described in more detail in [2] and is proven to be differentially private.

[0076] While added noise typically reduces the overall accuracy or performance of the model, it has the benefit of improving the privacy of the trained model. Such noise limits the influence of individual training data records on the model parameters, which in turn reduces the likelihood that information related to those individual training data records will be leaked by the trained model. Gradient clipping works in a similar manner to achieve differential privacy. Generally speaking, for machine learning models based on neural networks, gradient clipping involves setting a maximum limit on any given gradient used to update the weights of the neural network. Limiting the gradient has the effect of reducing the influence of any particular training data set on the model parameters, thereby protecting the privacy of the training data (or the individual or entity corresponding to the training data).

[0077] In many GAN architectures, the generator receives no input at all, or receives a random or pseudo-random "seed" for generating artificial data. Therefore, the generator does not "directly" risk exposing private data during the training process because it generally does not have access to the training data. However, the discriminator does use the training data to distinguish artificial data (generated by the generator) from real training data. Since the generator loss value (and therefore the generator update value and generator parameters) is based on the performance of the discriminator, the generator may inadvertently violate privacy via the discriminator. An embodiment can solve this problem by adding noise to the discriminator update value (e.g., the discriminator gradient) and optionally clipping the discriminator update value. Although noise can optionally be added to the generator update value, this is not necessary because the discriminator generally has access to the (potentially sensitive or private) training data (rather than the generator), and therefore adding noise to the discriminator update value is sufficient to achieve differential privacy.

[0078] E. Minority Data Representation

[0079] A general challenge for artificial data generation systems is to accurately represent the corresponding real dataset, including accurately representing minority data records. Some differentially private GAN systems (such as those described in [4]) have difficulty with minority data representation and sparse data representation. However, embodiments use "conditioning vectors" (described in more detail below) to provide better minority data representation in artificial data records.

[0080] A minority data record generally refers to a data record that has data characteristics that are rare or otherwise inconsistent with the "average" data record in the dataset. Typically, artificial data generators often do a good job of generating artificial data that represents average data. This is because machine learning models are typically evaluated using a loss value associated with the difference between the expected or ideal result (e.g., the real training data record) and the result produced by the generator system (e.g., the artificial data record). Generating artificial data records that are similar to the average data record is often effective in minimizing such loss values. As a result, machine learning models often learn this behavior unintentionally.

[0081] However, minority data records are (by definition) different from majority data records, and therefore different from average data records. Since many such machine learning models typically perform poorly in generating artificial data records corresponding to minority data records, since such data records are rare in nature, they typically do not have a large impact on model update values ​​and trained model parameters, and therefore, such models cannot learn how to generate minority data records.

[0082] This can be problematic in applications that involve detecting a small number of data examples. An example is detecting credit card fraud. Because most credit card transactions are legitimate, fraudulent credit card transactions comprise a very small minority. However, data analysts are much more interested in identifying fraudulent transactions than legitimate credit card transactions. Fraudulent credit card transactions require some corrections, e.g., the transaction should be canceled or reversed, the card should be deactivated, etc. This is in contrast to legitimate credit card transactions, which typically do not require such corrections and generally do not need to be detected. Due to the rarity of fraudulent credit card transactions, a generator trained to produce artificial credit card transaction data may inadvertently produce an artificial dataset that does not contain fraudulent transactions. This is problematic because such an artificial dataset may not be suitable for any further analysis or processing (e.g., training a machine learning model to detect fraudulent credit card transactions).

[0083] 1. Using Conditional Vectors to Improve Minority Data Representation

[0084] One technique for improving the representation of minority data is to use “conditional vectors” (also called “mask vectors”) in model training. The use of conditional vectors is described in more detail in [5], which proposes “Conditional Tabular GAN” (or “CTGAN”), a GAN system that uses conditional vectors to improve the representation of minority data. However, while CTGAN uses conditional vectors to improve the representation of minority data, it does not guarantee differential privacy (unlike the embodiments of this disclosure).

[0085] The use of condition vectors is generally summarized as follows. In general, the condition vector defines the conditions applied to the artificial data generated by the generator. This can involve, for example, requiring the generator to generate artificial data records with specific characteristics or specific data features, or alternatively "encouraging" or "penalizing" (e.g., based on a small or large loss value) the generator to generate artificial data records with or without those characteristics or data features. In this way, training can be better controlled. In a system where data records are sampled completely randomly for training, the sampling rate of minority data records is proportional to the minority group within the overall training dataset, and therefore the generator can only generate a minority of artificial data records proportional to this minority group. But by using condition vectors, it is possible to control the frequency with which the generator generates artificial data records belonging to a specific classification or category during training. If 10% of the condition vectors specify that the generator should generate artificial data records corresponding to minority data, then the generator can generate artificial data records at that 10% rate, rather than based on the actual proportion of minority data records in the training dataset. Therefore, the use of condition vectors can produce higher quality artificial data records, which are well suited for sparse data and for preserving minority data representations.

[0086] 2. Conditional Vector Privacy Issues

[0087] Embodiments of the present disclosure provide some techniques for providing differential privacy when using conditional vectors. As described above, conditional vectors can be used to force or incentivize a machine learning model to generate artificial data records with specific characteristics, such as minority data characteristics. In this way, the machine learning model can learn to generate artificial data records that represent the entire data set, not just the majority data. As an example, users of a streaming service may generally skew towards younger users, however, there may be a small number of older users (e.g., 90 years old or older). A machine learning model (not using conditional vectors) may inadvertently learn to generate artificial user data records corresponding to younger users, and never learn to generate artificial user data records corresponding to older users. However, by using conditional vectors, the machine learning model can be forced during training to learn how to generate artificial data records corresponding to older users, and therefore learn how to better represent the entire input data set.

[0088] However, conditional vectors present unique challenges for achieving differential privacy. In the context of machine learning, differential privacy is related to the frequency with which specific data records are sampled and used during training. If a data record (or the data value contained in that data record) is used more frequently during training, there is a greater privacy risk. This is not a problem for majority data records because there are a large number of majority data records, and therefore the probability of sampling any specific data record is low. However, because conditional vectors encourage machine learning models to generate artificial data corresponding to minority data records, they can increase the probability that minority data records are used for training. Because there are typically fewer minority data records, the probability of sampling any given minority data record increases. For example, if there are only ten users aged 90 or older, there is a 10% chance that any given user will be sampled when sampling from that subset of users. Therefore, there is a greater risk that the trained model learns private information corresponding to a specific minority data record and inadvertently leaks private information in the generated artificial dataset.

[0089] F. Differential Privacy Condition Vector

[0090] Embodiments of the present disclosure involve novel techniques that can be used to address the privacy issues described above. One such technique is "class combination." As described above, by using conditional vectors to "encourage" a machine learning model to learn to generate accurate artificial data records corresponding to a small number of data, the machine learning model has a greater chance of sampling or using data from any particular data record, thereby risking the privacy of that data record. Class combination can be used to reduce the probability of sampling or using any particular data record in training, and thus can reduce privacy risks and enable embodiments to guarantee differential privacy.

[0091] In short, an artificial data synthesizer (e.g., a computer system) can count the number of training data records corresponding to the identified categories. For example, the artificial data synthesizer can count the number of training data records corresponding to "young" users, "middle-aged" users, "old" users, and so on of the streaming service. Categories with "too few" corresponding data records (e.g., less than a predetermined limit) may correspond to a minority of data and may pose a greater privacy risk. If the artificial data synthesizer determines that a category is "defective" (e.g., contains less than a minimum number of corresponding data records), the artificial data synthesizer may merge the category with one or more other categories. For example, if there are too few "old" users of the streaming service, the artificial data synthesizer may merge the "old" and "middle-aged" categories into a single category.

[0092] Thus, when a machine learning model is trained using a conditional vector, rather than encouraging the machine learning model to generate artificial data records corresponding to, for example, "elderly" users, the conditional vector can instruct the machine learning model to generate artificial data records corresponding to users in the combined "elderly and middle-aged" category. Because this combined category is larger than the "elderly" category, the probability of sampling data from any particular data record is reduced, and thus the privacy risk is reduced.

[0093] 1. Differential Privacy and Counting

[0094] As described above, based on the method or process The output reveals that specific data records are included as part of the process Differential privacy is a strong mathematical privacy guarantee based on the probability that 123,281,392 users of the service are present in the input dataset. In general, differential privacy is "stricter" than human interpretations of what privacy means. As a result, a process may fail to provide differential privacy in ways that may not be intuitive. One such example is counting. When human users think of privacy leaks, they typically think of their personally identifiable information (e.g., name, social security number, email address, etc.) being exposed. One would not typically think of a count of users of the service, e.g., 123,281,392, as somehow including a privacy leak. However, such counts typically violate the rules of differential privacy, as explained below.

[0095] Consider two adjacent data sets d and d' that are identical except that one contains a particular data record and the other does not. To count the number of data records in each data set d and d' would produce two different counts because each data set contains a different number of data records. Differential privacy will not be satisfied because the outputs of the process applied to the two data sets are not similar or similarly distributed. In this way, counting categories in order to determine whether those categories should be combined (as described above) risks violating differential privacy.

[0096] Therefore, to ensure differential privacy, embodiments of the present disclosure may use noisy class counts to assess whether a class is defective (e.g., contains too few data records). These noisy class counts may include the sum of a class count (e.g., the actual number of data records belonging to a particular class) and a class noise value (e.g., a random number). Due to the added class noise, it may not be possible to determine whether a particular data record is included in the count for a particular class, and thus these class counts no longer violate differential privacy. This is generally similar to the database income example provided above, where adding noise to the average income of some number of individuals protects the privacy of the particular individual (e.g., Alice) whose income is used to calculate the average income statistic.

[0097] II. System Diagram

[0098] Figure 3 A diagram illustrating an artificial data synthesizer 302 and a data source 304 according to some embodiments of the present disclosure. The artificial data synthesizer 302 may include several components, including a generative adversarial network (GAN). A generator sub-model 318, a discriminator sub-model 322, a generator optimizer 334, and a discriminator optimizer 336 may be components of this GAN. The artificial data synthesizer 302 may further include a data processor 308 and a data sampler 312, which may be used to process or pre-process data records (retrieved from the data source 304) for training the GAN. Once the GAN is trained, the artificial data synthesizer 302 may use the generator sub-model 318 to generate artificial data records.

[0099] Figure 3 The components of the artificial data synthesizer 302 described in the accompanying drawings are primarily intended to explain the functionality of the artificial data synthesizer and method according to an embodiment and are not intended to be a limiting depiction of the form of the artificial data synthesizer 302. As an example, although Figure 3Separate data processor 308 and data sampler 312 are depicted, but data processor 308 and data sampler 312 may comprise a single component. Artificial data synthesizer 302 may comprise a computer system or may be implemented by a computer system. For example, artificial data synthesizer 302 may comprise a software application or executable executed by a computer system. Each component of the artificial data synthesizer may comprise a physical device (e.g., data processor 308 and data sampler 312 may comprise separate devices connected via an interface) or may comprise a software module. In some embodiments, artificial data synthesizer 302 may be implemented using a monolithic software application executed by a computer system.

[0100] In summary, the artificial data synthesizer 302 can retrieve data records from the data source 304 (in Figure 3 306). This data source 304 may comprise, for example, a database, a data stream, or any other suitable data source. Raw data 306 may have several generally undesirable characteristics. For example, raw data 306 may include duplicate data records, erroneous data records, data records that do not conform to a specific data format, abnormal data records, etc. Artificial data synthesizer 302 may use data processor 308 to process raw data 306 to address these undesirable characteristics, thereby generating processed data 310. This processed data 310 may be sampled by data sampler 312 and used to train a GAN to generate artificial data records.

[0101] References Figures 4 to 7 Specific data processing (or data preprocessing; the terms are used largely interchangeably herein) operations are described in more detail. These data processing operations may include data preprocessing steps, such as data validation, data cleaning, removal of outliers, etc., as well as specific data processing steps that can generate privacy-preserving artificial data records. As a brief overview, these steps may include (1) identifying and removing non-sparse data records, (2) normalizing numerical data values, (3) assigning categories to the normalized numerical data values, (4) counting the number of data records corresponding to each category, (5) identifying any defect categories, (6) combining defect categories, and (7) updating the data record to identify the combined categories. These steps are described in more detail further below. Some of these steps are described and motivated in Section I above. For example, defect categories may be combined to reduce the probability of sampling any particular data record or data value during training, thereby reducing privacy risks.

[0102] Data sampler 312 can sample data records 316 from processed data 310 to serve as training data. This training data can be used to train the GAN to generate artificial data records. Data sampler 312 can also generate condition vectors 314. These condition vectors 314 can be used to encourage generator sub-model 318 to generate artificial data records 326 with certain characteristics or data values. For example, data records contained in processed data 310 may correspond to users of a streaming service. Such user data records may include a data field corresponding to the user's age, and users may be categorized by the data value corresponding to this data field. Some users may be characterized as "young people," while other users may be categorized as "adults," "middle-aged," "elderly," and so on. Condition vectors 314 can be used to cause generator sub-model 318 to generate artificial data records 326 corresponding to each of these categories, thereby training generator sub-model 318 to generate artificial data records 326 that better represent the entire processed data 310.

[0103] In some cases, condition vector 314 may identify specific data fields corresponding to sampled data record 316 that generator sub-model 318 should replicate when generating artificial data record 326. For example, if sampled data record 316 has a data field indicating that the corresponding user is a "senior citizen," generator sub-model 318 may generate artificial data record 326 that also contains a data field indicating that the "artificial user" corresponding to artificial data record 326 is a "senior citizen." This may be useful if there is a very small percentage of elderly users.

[0104] During training, these condition vectors 314 and sampled data records 316 can be divided into batches. Over the course of several training rounds, the generator sub-model 318 can use this data and generator input noise 342 (e.g., a random seed value sampled from a distribution unrelated to the processed data 310) to generate artificial data records 326 corresponding to each training round. These artificial data records 326 and any corresponding sampled data records 316 can be provided to the discriminator sub-model 322 without indicating which data records are artificial and which data records are sampled.

[0105] The discriminator sub-model 322 can attempt to identify the artificial data record 326 by comparing the artificial data record 326 with the sampled data records 316 in the batch. Based on this comparison, loss values ​​328 can be determined, including a generator loss value 330 and a discriminator loss value 332. As described in Section I, these loss values ​​328 can be based on the ability of the discriminator sub-model 322 to identify the artificial data record 326. For example, if the discriminator sub-model 322 correctly identifies the artificial data record 326 with a high degree of confidence, the discriminator loss value 332 can be small, while the generator loss value 330 can be large.

[0106] The generator loss value 330 and the discriminator loss value 332 can be provided to a generator optimizer 334 and a discriminator optimizer 336, respectively. The generator optimizer 334 can use the generator loss value 330 to determine one or more generator update values ​​338, which can be used to update the generator parameters 320, which can characterize the generator sub-model 318. As an example, the generator sub-model 318 can be implemented using a generator artificial neural network, and the plurality of generator parameters can include a plurality of generator weights corresponding to the generator artificial neural network. In this case, the generator optimizer 334 can use stochastic gradient descent to determine the generator update values ​​338 including gradients. These gradients can be used to update the generator weights, for example, using backpropagation.

[0107] In a similar manner, the discriminator optimizer 336 can use the discriminator loss value 330 to determine a noise discriminator update value 340, which can be used to update the discriminator parameters 324 representing the discriminator sub-model 322. However, the discriminator optimizer 336 can perform some additional operations to ensure differential privacy. As an example, the discriminator optimizer 336 can generate initial discriminator update values ​​(e.g., noise-free gradient discriminator update values) and then clip these initial discriminator update values. The discriminator optimizer 336 can then add noise to the initial discriminator update values ​​to generate the noise discriminator update value 340, which can then be used to update the discriminator parameters 324. As an example, the discriminator sub-model 322 can be implemented using a discriminator artificial neural network, and the plurality of discriminator parameters 324 can include a plurality of discriminator weights corresponding to the discriminator artificial neural network. In this case, the noise discriminator update value 340 can include one or more noise discriminator gradients, which can be used to update the discriminator weights, for example, using backpropagation.

[0108] This training process can be repeated over several training rounds or periods. In each training round, a new condition vector 314 and a new sampled data record 316 can be used to generate artificial data records 326, loss values ​​328, and model update values, thereby producing updated generator parameters 320 and discriminator parameters 324. In this way, training improves the ability of the generator sub-model 318 to generate convincing or representative artificial data records 326 and improves the ability of the discriminator sub-model 322 to identify artificial data records 326. This training process can be repeated until a termination condition is met. For example, the termination condition can specify a specific number of training rounds (e.g., 10,000), and once the number of training rounds has been performed, the training process can be completed. The artificial data synthesizer 302 can periodically check to see if the termination condition has been met. If the termination condition has not been met, the artificial data synthesizer 302 can repeat the iterative training process; otherwise, the artificial data synthesizer 302 can terminate the iterative training process.

[0109] Once trained, the trained generator sub-model 318 can be used to generate a privacy-preserving artificial dataset that can be published or transmitted to a client computer, for example. Alternatively, the generator sub-model 318 itself can be published or transmitted to a client computer so that an entity (e.g., from Figure 1 The artificial data using entity 104) can generate artificial data sets as it sees fit. Figure 3 An artificial data synthesizer 302 comprising a GAN is depicted, but other model architectures are possible, such as autoencoders, variational autoencoders (VAEs), or transformations or combinations thereof.

[0110] III. Preprocessing Model Training Data

[0111] Figure 4 A flowchart illustrating a method for training a machine learning model to generate a plurality of artificial data records (sometimes referred to as an "artificial dataset") in a privacy-preserving manner. This method can be implemented by an artificial data synthesizer (e.g., from Figure 3 The artificial data synthesizer 302) is executed by a computer system.

[0112] A. Retrieval and Initial Data Processing

[0113] At step 402, a computer system may retrieve a plurality of data records from a data source (e.g., a database or a data stream) and perform any initial data processing operations. Each data record may include a plurality of data values ​​corresponding to a plurality of data fields. Each data value may identify a category within a plurality of categories or within a category within a plurality of categories. As an example, a data record corresponding to a restaurant may have a "popularity" data value, such as 0.9, indicating that it is in the 90th percentile of restaurant popularity within a given location. This popularity data value may be a category (e.g., "very popular") within a plurality of categories (e.g., "not popular," "mildly popular," "popular," "very popular," etc.) or otherwise indicate the category.

[0114] Initial processing can use data processor components (e.g. Figure 3 The data processor 308 is used to complete the process. It may include various processing functions, which are described in detail in the following sections. Figure 5 Describe in more detail.

[0115] At step 502, the computer system may perform various data pre-processing operations on the plurality of data records. These may include, for example, "data validation" operations, such as operations for verifying that the data record is valid (e.g., conforms to a particular format or contains more or less than a particular amount of data (e.g., greater than 1 KB, less than 1 GB, etc.), and "data cleaning" or "data scrubbing" operations, which may involve removing incomplete, inaccurate, incorrect, or erroneous data records from the plurality of data records prior to any further pre-processing or training of a machine learning model. Additionally, data records corresponding to identifiable outliers may also be removed from the plurality of data records. These examples are intended to be illustrative and are not intended to provide an exhaustive list of every operation that may be performed on the data records prior to further processing.

[0116] At step 504, the computer system may identify and remove non-sparse data records from the plurality of data records. These non-sparse data records may include outliers or may have increased privacy risks. Figure 4 For each data record in the training set (retrieved at step 402), the computer system can determine whether the data record has more than a maximum number of non-zero data values. Then, for each data record containing more than the maximum number of non-zero data values, the computer system can remove the data record from the plurality of data records, preventing these anomalous data records from being used for later training.

[0117] In the implementation reference Figure 4 and Figure 5Prior to the described training method, this maximum number of non-zero data values ​​can be predetermined. Alternatively, a privacy analysis can be performed to determine the maximum number of non-zero data values ​​corresponding to a particular set of privacy parameters (ε, δ). For example, for lower (ε, δ) values ​​corresponding to more stringent privacy requirements, the maximum number of non-zero data values ​​can be lower than for higher (ε, δ) values. The relationship between the privacy parameters (ε, δ) and hyperparameters of the machine learning process (including, for example, the maximum number of non-zero data values) can be complex and, in some cases, cannot be expressed by a simple closed formula. In this case, the privacy analysis can enable the computer system (or, for example, a data analyst operating the computer system) to determine the maximum number of non-zero data values ​​that achieves a desired level of differential privacy.

[0118] As described above, embodiments of the present disclosure can use conditional vectors to protect minority data representations and thus generate more representative artificial data. However, as described above, using conditional vectors to improve minority data representations may lead to additional privacy risks because this technique increases the rate at which minority data values ​​can be sampled during training. An embodiment addresses this problem by combining minority classes with other classes, which can reduce the probability of sampling any particular data record or data value during training. To do this, a class can be determined for a particular data value to determine which data values ​​and data records correspond to the minority class. The computer system can perform a two-step process (steps 506 and 508) to determine a class or assign a class to a data value.

[0119] At step 506, the computer system may normalize any non-normalized numerical data values ​​in the plurality of data records. For each data record in the plurality of data records, the computer system may normalize one or more data values ​​between 0 and 1, including the end value (or any other appropriate range), thereby generating one or more normalized data values. As an example, a data record corresponding to a golfer may contain data values ​​corresponding to his or her driving distance (measured in yards), driving accuracy (percentage), and average ball speed (measured in meters per second). The numerical driving distance data value and the average ball speed data value may be normalized to a range of 0 to 1. Driving accuracy may not require normalization because percentages are typically already normalized data values. Normalized data values ​​may be more easily assigned categories (e.g., in step 508 described below) than non-normalized data values ​​due to their defined ranges.

[0120] At step 508, the computer system may assign a normalized category to each of the normalized numerical data values. For example, a normalized numerical data value corresponding to a golfer's driving distance may be assigned a "low" driving distance category, a "medium" driving distance category, or a "long" driving distance category. These normalized categories may be included in a plurality of categories already determined or identified by the computer system. For example, these categories may be included in already determined categories such as "amateur," "semi-professional," and "professional," or any other determined categories.

[0121] In more detail, the computer system can determine multiple normalized categories for each normalized numerical data value in the one or more normalized data values ​​based on corresponding probability distributions in one or more probability distributions. Each probability distribution can correspond to a different normalized numerical data value in the one or more normalized numerical data values. In some cases, these probability distributions can include multimodal Gaussian mixture models. Examples of such distributions are described in Figure 6 Such a probability distribution may include a predetermined number (m) of equally weighted modes.

[0122] Figure 6 Three such modes (mode 1 604, mode 2 606, and mode 3 608) are shown distributed within a normalized range corresponding to the normalized data values. Each mode can correspond to a Gaussian distribution where, for example, the mean is equal to its corresponding mode and the standard deviation is equal to the inverse of the number of modes (1 / m). Each mode can also correspond to a class such that for each normalized class in a plurality of normalized classes, there is a corresponding mode in a plurality of equally weighted classes. In other words, the number of normalized classes can be equal to the number of equally weighted modes. For example, in Figure 6 , the normalized data values ​​corresponding to this probability distribution can be assigned to one of three categories because there are three modes 604 to 608. For example, a golfer's normalized driving distance can be assigned to a category such as low, medium, or high based on its value.

[0123] The computer system may use any suitable method to assign normalized data values ​​to normalized categories using such a probability distribution. For example, the computer system may determine the distance between a particular normalized data value and each of a plurality of equally weighted patterns and then assign the normalized data value to the category corresponding to the closest pattern. For example, in Figure 6 , normalized data values ​​close to 0.5 may be assigned to category 2 606 , while normalized numerical data values ​​corresponding to 0.9 may be assigned to category 3 608 .

[0124] Using a Gaussian mixture model with an equal weighting pattern may not perfectly represent the actual distribution of categories in the data records. For example, for a streaming service, most users may be "light users," corresponding to low "hourly ratings" data values ​​and low normalized hourly ratings data values. However, a probability distribution with an equal weighting pattern implicitly indicates that the distributions of "light users," "medium users," and "heavy users" are roughly equal. More accurate Gaussian mixture model techniques can be used to generate probability distributions that more closely reflect the actual distribution of categories in the data records. However, such techniques depend on the actual distribution of data values ​​and therefore introduce another means for leaking sensitive data. For example, knowing the relative proportion of data records corresponding to each category can enable an individual to identify a specific data record based in part on its category. However, using an equal weighting pattern is independent of the actual distribution of the data and therefore does not leak any information about the distribution of data values ​​in the data records, thereby protecting privacy.

[0125] At this point, the computer system can now count and group the categories (steps 404 to 412) in order to train the machine learning model in a privacy-preserving manner (step 418). As described above in Section I, if any category is defective, i.e., corresponds to too few data records, it may be sampled too frequently during training and may be at risk of exposing the private data contained in those data records. By counting and grouping defective categories, the sampling probability can be reduced, thereby improving privacy.

[0126] B. Category Counting and Merging

[0127] return Figure 4 At step 404, the computer system may determine a plurality of noise class counts corresponding to the plurality of classes. Each noise class count may indicate an estimated (e.g., approximate) number of data records in the plurality of data records (retrieved at step 402) belonging to each of the plurality of classes. For example, if the data records correspond to patient health information, the computer system may determine an estimated or approximate number of data records corresponding to "elderly" patients, "hypotensive" patients, patients with valid health insurance, and the like. Each noise class count may include the sum of a class count (of the plurality of class counts) and a class noise value in the one or more class noise values. For example, the same class noise value may be added to each class count, in which case the one or more class noise values ​​may include a single noise value. Alternatively, different class noise values ​​may be added to each class count, in which case the one or more class noise values ​​may include multiple class noise values.

[0128] The noise value of each category can be obtained by the category noise mean and category noise standard deviation σ count Definitions. The class noise mean and class noise standard deviation may correspond to a probability distribution that may be used to determine the class noise value. For example, each class noise value may be sampled from a normally distributed Gaussian distribution (sometimes referred to as a "first Gaussian distribution") having a mean equal to the class noise mean and a standard deviation equal to the class noise standard deviation.

[0129] As described above, the process of (noise-free) counting may violate the definition of differential privacy because two data sets (one including a particular data record and one not including the data record) may result in different data record counts. Therefore, in theory, noise-free class counts can enable an individual to determine whether a particular data record is included in a particular class. Therefore, class noise values ​​can be added to class noise counts to determine noise class counts, which indicate the estimated (or approximate) number of data records corresponding to each class and thus protect privacy. In general, a larger class noise standard deviation results in a greater variety of class noise values ​​that can be added to the class counts, and thus provides higher privacy than a smaller class noise standard deviation.

[0130] Thus, the class noise mean and the class noise standard deviation can be determined based on one or more class noise parameters, which can include one or more target privacy parameters related to specific privacy requirements for artificial data generation. The target privacy parameters can correspond to the desired privacy level and can include the epsilon (ε) privacy parameter and delta (δ) for characterizing differential privacy for training of machine learning systems. The class noise parameters can also include a minimum count L (used to identify whether a class is defective, i.e., corresponding to too few data records), a maximum number of non-zero data values ​​X, and a maximum number of non-zero data values ​​X. max (for removing non-sparse data records, as described above), a safety margin a, and the total number V of data values ​​in a given data record.

[0131] The relationship between the class noise mean, the class noise standard deviation, and the class noise parameter may not have a closed form or otherwise accessible parameter relationship. In some cases, a "privacy analysis" may be performed by a computer system or by a data analyst operating a computer system to determine the class noise mean and the class noise standard deviation based on the class noise parameter. For example, a "worst case" privacy analysis may be performed based on The "worst case value" to perform, that is, the maximum number of non-zero data values ​​X maxdivided by the total number of data values ​​in a given data record V, multiplied by one divided by the minimum count L minus the safety margin a. If later model training (e.g., at step 418) uses batches b of size b>1, the worst-case value can instead be given by Represents. Privacy analysis can also be used to determine the number of training rounds to perform during model training. Further information on privacy analysis and how to perform it can be found in references [6] and [7].

[0132] Any of these worst-case values ​​can relate to the probability of sampling a particular data value contained in a particular data record during training, which is further proportional to the privacy risk, as defined by the (ε, δ) differential privacy definition provided in Section 1. Thus, the class noise mean and class noise standard deviation can be determined based on the worst-case values. For example, to accommodate large worst-case values ​​(indicating a greater privacy risk), a large class noise standard deviation can be determined, while for smaller worst-case values ​​(indicating a lower privacy risk), a smaller class noise standard deviation can be determined.

[0133] At step 406, after determining the noise class count, the computer system may identify defective classes based on the class noise count and the minimum count. Each defective class may include classes whose corresponding noise class count is less than the minimum count. For example, if the minimum class count is "1000" and a class (e.g., "Popular Restaurants" corresponding to data records for restaurants) includes only 485 data records, then the class may be identified as a defective class. The computer system may parse the retrieved data records and, whenever the computer system encounters a data record corresponding to each class, increment the class count corresponding to the class.

[0134] As described above, the probability of sampling any given data record or data value during training can be proportional to the number of data records in a given class. Therefore, defective classes pose a greater privacy risk because they correspond to fewer data records. Defective classes can be combined (e.g., in step 408) to address this privacy risk and provide differential privacy. Similar to the class noise standard deviation, the minimum count can be determined in whole or in part through a privacy analysis, which can involve determining the minimum count based on, for example, specific (ε, δ) privacy parameters.

[0135] At step 408, the computer system may combine each defect category from the one or more categories (e.g., identified in step 406) with at least one other category from the plurality of categories, thereby determining a plurality of combined categories. Typically, the combined categories preferably include a plurality of data records greater than a minimum count. However, categories may be combined in any suitable manner. For example, a defect category may be combined with other defect categories to produce a combined category that is free of defects (i.e., containing more data records than the minimum count). Alternatively, a defect category may be combined with a non-defect category to achieve the same result. Defect categories may be combined with similar categories. For example, for medical data records, if the category "very old" is a defect category, this category may be combined with a similar category "elderly" to create a combined "elderly / very old" category. While such a category combination may be logical or may produce more representative artificial data, there is no strict requirement to combine categories in this manner. Alternatively, the "very old" category may be combined with the "newborn" category if both categories are defective and if combining the categories would produce a non-defective "very old / newborn" category.

[0136] At step 410, the computer system may identify one or more defective data records. Each defective data record may contain at least one defective data value that may correspond to the combined categories. For example, if the "very old" category was found to be defective and the combined "old / very old" category was generated at step 408, the computer system may identify defective data records containing data values ​​corresponding to the "old" category or the "very old" category. The computer system may do this by iterating through the retrieved data records and their corresponding data values ​​to identify these defective data records and defective data values.

[0137] At step 412, the computer system may replace the defective data values ​​in the defective data record with the combined data value. For each defective data value contained in one or more defective data records, the computer system may replace the defective data value with a combined data value identifying a combined category of the plurality of categories. For example, if the computer system identifies a defective healthy data record containing a defective data value identifying the "very old" category (which has been combined into the "old / very old" category), the computer system may replace the defective data value with a data value identifying the combined "old / very old" category instead of the "old" category. This combined data value may also include a noise category count corresponding to each of the combined categories.

[0138] refer to Figure 7The process of steps 404 to 412 can be better understood by illustrating an exemplary data record 702. This data record 702 shows three data fields corresponding to age, height, and blood pressure, and three categories corresponding to these three data fields (i.e., very old, short, and very low blood pressure). The computer system can determine a noise category count corresponding to each of these categories (e.g., in Figure 4 404).

[0139] Figure 7 Three such noise class counts are shown. The "elderly" noise class count 704 includes approximately 751 data records. The "short" noise class count 706 includes approximately 3212 data records. The "very low blood pressure" noise class count 708 includes approximately 653 data records.

[0140] The computer system may compare each of these noise category counts 704 to 708 to a minimum count 710 (ie, 1000) to determine the noise category counts. Figure 4 At step 406, the computer system identifies whether any of these categories are deficient. Based on this comparison, the computer system may determine that the "very old" category and the "very low blood pressure" category are deficient (and therefore any data values ​​contained in data record 702 indicating these categories are deficient data values), while the "short" category is not deficient. The computer system may combine these deficient categories with other categories (e.g., in Figure 4 At step 408). Figure 7 In the example of , the computer system can combine the "very old" category with the "elderly" category to create a combined "elderly / very old" category. Similarly, the computer system can combine the "very low blood pressure" category with the "low blood pressure" category to create a combined "low / very low blood pressure category."

[0141] The computer system may then (for example, in Figure 4 At step 410 of ), any defective data records in the data set are identified, including data record 702, which includes data values ​​identifying two different defect classes. The computer system may then replace the defective data values ​​in these defective data records with the combined data values. These combined data values ​​may identify the combined classes and may additionally include noise class counts corresponding to the classes in the combined classes. For example, in Figure 7 , the updated data record 712 has an "age" data value of "Old 997 / Very Old 751," which indicates the combined "Old / Very Old" category, and noise category counts for both the "Old" and "Very Old" categories.

[0142] Return Reference Figure 4At step 414, the computer system may generate a plurality of condition vectors for use during training. Each condition vector may identify one or more specific data fields. These data fields may be used to determine the data values ​​that the generator sub-model should copy or reproduce during training. Again, briefly referring to Figure 2 , the value "1" in the third position of condition vector 214 can indicate that the generator should copy the data value corresponding to the "data usage" field during training. Thus, this exemplary condition vector identifies this data field. In some embodiments, each condition vector can include the same number of elements as each data record.

[0143] The condition vectors can be generated in any suitable manner, including randomly or pseudo-randomly. In some cases, it may be preferable to generate the condition vectors so that they identify data fields with equal probability. For example, if a plurality of data records each includes ten data fields, the probability of any particular data field being identified by the generated condition vector may be equal (approximately 10%). Alternatively, it may be preferable to generate the condition vectors so that certain data fields are "prioritized" relative to other data fields. This may be the case if, for example, a particular data field is associated with a minority of data records more than other data fields. In this case, if the condition vector identifies this data field more frequently than other data fields, better minority data representation may be achieved.

[0144] At step 416, the computer system may sample a plurality of sampled data records from the plurality of data records. The sampled data records may include at least one of the one or more defective data records (e.g., identified at step 410). Each sampled data record may include a plurality of sampled data values ​​corresponding to a plurality of data fields. The sampled data records may include data records used for machine learning model training (e.g., at step 418). In some embodiments, all data records retrieved from the data source (excluding those that are filtered or removed, for example, due to non-sparseness) may be sampled and used as sampled data records for training. Additionally, at step 416, the computer system may process the sampled data records, particularly if any sampled data record contains a sampled data value that identifies the combined class.

[0145] C. Processing sampled data records before training

[0146] The sampled data records containing data values ​​identifying the combined categories can be updated to identify a single category. This can be used in conjunction with the conditional vector. If the conditional vector indicates a data field that includes two categories (e.g., "Old 997 / Very Old 751"), it may be difficult to use the conditional vector to identify a single category or date value to reproduce in training. Therefore, the computer system can update the data value identifying the combined categories (e.g., "Old 997 / Very Old 751") to identify a single category, such as "Old" or "Very Old."

[0147] At step 418, prior to the step of training the machine learning model to generate the plurality of artificial data records, the computer system may identify one or more sampled data values ​​from the plurality of sampled data records. Each of the one or more identified sampled data values ​​may correspond to a corresponding combined category in one or more corresponding combined categories. The computer system may do this by iterating over the data values ​​in each sampled data record and identifying whether those data values ​​correspond to a combined category. Such data values ​​may include a string, flag, or other indicator indicating that they correspond to a combined category, or may be in a form indicating that they correspond to a combined category, e.g., a string such as "elderly / very old" may define two categories ("elderly" and "very old") based on the position of backslashes.

[0148] For each identified sampled data value, the computer system can determine two or more categories that are combined to create each of the combined categories. For example, for a string data value such as "Old 997 / Very Old 751," the computer system can determine that the two categories are "Old" and "Very Old" based on the structure of the string. The computer system can then randomly select a random category from the two or more categories, for example, by randomly selecting "Old" or "Very Old" from a given instance. The computer system can then generate a replacement sampled data value that identifies the random category and replace the identified sampled data value with the replacement sampled data value. In this way, each sampled data record can now identify a single category for each data field, rather than any combined categories.

[0149] refer to Figure 8 This process can be better understood by illustrating an exemplary sampled data record 802 corresponding to health data, wherein the data fields correspond to age, height, and blood pressure. The age data field (and the data value corresponding to this data field) corresponds to the combined "Old 997 / Very Old 751" category. Similarly, the blood pressure data field (and the data value corresponding to this data field) corresponds to the combined "Low 1300 / Very Low 653" category.

[0150] This sampled data record can be updated so that the two data values ​​corresponding to age and blood pressure identify a single category, rather than a combined category. There are four possible combinations of identified categories. Two such possible combinations are shown in updated sampled data record 804 and updated sampled data record 806. In updated sampled data record 804, the category "Old" has been randomly selected to replace the combined category "Old 997 / Very Old 751", and the category "Low" has been randomly selected to replace the combined category "Low 1300 / Very Low 653". Similarly, in updated sampled data record 806, the category "Very Old" has been randomly selected to replace the combined category "Old 997 / Very Old 751", and the category "Very Old" has been selected to replace the combined category "Low 1300 / Very Low 653".

[0151] Although Figure 8 Individual data values ​​are not depicted in , but such data values ​​can indicate their corresponding categories. For example, if a normalized data range of 0.0 to 0.2 is assigned to the "very low blood pressure" category (e.g., using a multimodal distribution as described above). The data value corresponding to the blood pressure field can be replaced with a replacement data value corresponding to any number within this range (e.g., 0.1) selected by any appropriate means (e.g., the average data value within this range, a random data value within this range, etc.). Alternatively, the replacement data value can include a string or other identifier that identifies the corresponding category (e.g., "low blood pressure").

[0152] In some embodiments, a random class may be selected using weighted random sampling using any noise class count indicated by the combined class. For example, for the combined "Old 997 / Very Old 751" class, the probability of randomly selecting the "Old" class may be equal to The probability of randomly selecting a very old category can be equal to The computer system may, for example, uniformly sample random numbers in the range of 1 to (997+751). If the sampled random number is 997 or less, the computer system may randomly select the "elderly" category. If the sampled random number is 998 or greater, the computer system may randomly select the "very old" category.

[0153] return Figure 4 At step 418, the computer system may train a machine learning model using the plurality of sampled data records and the plurality of condition vectors to generate a plurality of artificial data records. Each artificial data record may include a plurality of artificial data values ​​corresponding to a plurality of data fields. The plurality of data fields may include one or more data fields identified by the condition vectors. The machine learning model may replicate one or more sampled data values ​​corresponding to one or more specific data fields in the plurality of artificial data values ​​based on the plurality of condition vectors.

[0154] In slightly more accessible terms, if a particular condition vector (used during a particular training round) identifies a data field, such as the "height" data field in a medical data record, the machine learning model can replicate the "height" value corresponding to the particular sampled data record (used during that particular training round) in multiple artificial data records. In this way, the machine learning model can learn to generate artificial data records that represent the entire sampled data record. In this context, "replicate" generally means to create with the intention of replicating. The machine learning model may not necessarily be able to exactly replicate one or more sampled data values ​​identified by the condition data vector (especially in early training rounds). Even after training is complete, the machine learning model may still not be able to exactly replicate such values. For example, if the condition vector identifies a sampled data value of "0.7", the machine learning model may "replicate" such data value as "0.689" in the artificial data record.

[0155] As described above, the machine learning model may include an autoencoder (e.g., a variational autoencoder), a generative adversarial network, or a combination thereof. In some embodiments, the machine learning model may include a generator submodel and a discriminator submodel. The generator submodel may be characterized by a plurality of generator parameters. Likewise, the discriminator submodel may be characterized by a plurality of discriminator parameters. In some embodiments, the generator submodel may be implemented using an artificial neural network (also referred to as a "generator artificial neural network"), and the generator parameters may include a plurality of generator weights corresponding to the generator artificial neural network. Likewise, the discriminator submodel may be implemented using an artificial neural network (also referred to as a "discriminator artificial neural network"), and the discriminator parameters may include a plurality of discriminator weights corresponding to the discriminator artificial neural network.

[0156] exist Figure 4 At any time during the method, the computer system can perform a "privacy analysis," such as the "worst case" privacy analysis described above with reference to class counts and merging. This privacy analysis can inform some of the steps performed by the computer system. As described above, embodiments of the present disclosure provide for differentially private machine learning model training. The "level" of privacy provided by the embodiments can be defined based on target privacy parameters, such as an epsilon (ε) privacy parameter and a delta (δ) privacy parameter. The computer system can perform this privacy analysis to ensure that the privacy of the machine learning training is consistent with these privacy parameters.

[0157] As an example, the privacy of this training process can be proportional to the amount of class noise added to the noise class counts. Larger noise can provide more privacy at the expense of lower representation of the artificial data. The computer system can perform this privacy analysis to determine how much class noise to add to the noise class counts in order to achieve differential privacy consistent with the target privacy parameters. As another example, the privacy provided by the machine learning model generally decreases with each training round or epoch. However, more training rounds generally result in more accurate or representative records of the artificial data. Therefore, the computer system can perform this privacy analysis to determine the number of training rounds or epochs to perform during step 418.

[0158] In some embodiments, training the machine learning model may include an iterative training process comprising a certain number of training rounds or epochs. This iterative training process may be repeated until a termination condition has been met. Figures 9A to 9B An exemplary training process is described.

[0159] IV. Model Training

[0160] Figures 9A to 9B An exemplary method for training a machine learning model to generate a plurality of artificial data records is described. This method can protect the privacy of sampled data values ​​contained in the plurality of sampled data records used during training, for example, by providing (ε, δ) differential privacy. Before performing this training process, a computer system can obtain a plurality of sampled data records. Each sampled data record can include a plurality of sampled data values. Similarly, the computer system can obtain a plurality of condition vectors. Each condition vector can identify one or more specific data fields. The computer system can use the above (e.g., reference Figure 4 ) to obtain these sampled data records and condition vectors. However, the computer system may also obtain these sampled data records and condition vectors via some other means. For example, the computer system may receive the sampled data records and condition vectors from another computer system, from a database of preprocessed sampled data records and condition vectors, or from any other source.

[0161] This training may include an iterative process that may include multiple training rounds and / or training epochs.

[0162] At step 902, the computer system may determine one or more selected sampled data records from a plurality of sampled data records. These selected sampled data records may include sampled data records used in a particular training round. For example, if there are 10,000 training rounds and the batch size of each training round is 100, the computer system may select 100 selected sampled data records for use in the particular training round. Alternatively, if the batch size is one, the computer system may select a single selected sampled data record for use in the particular training round.

[0163] At step 904, the computer system may determine one or more selected condition vectors. Similar to the selected sampled data records, these selected condition vectors may be used for a particular training run and may depend on the batch size. In some embodiments, for a particular training run, there may be as many selected condition vectors as there are selected sampled data records.

[0164] At step 906, the computer system may identify one or more conditional data values ​​from the one or more selected sampled data records. These one or more conditional data values ​​may correspond to one or more specific data fields identified by the one or more conditional vectors. Figure 2 As an example, condition vector 214 identifies "Data Usage" data field 204 (as well as other data fields). If data record 212 is the selected sampled data record, the computer system can use condition vector 214 to identify a data value "0.7" corresponding to the "Data Usage" data field identified by condition vector 214. This data value "0.7" can then comprise the conditional data value.

[0165] At step 908, the computer system may generate one or more artificial data records using the one or more conditional data values ​​and the generator sub-model. As described above, the generator sub-model may be characterized by a plurality of generator parameters, such as a plurality of generator neural network weights that characterize a neural network-based generator sub-model. The generator sub-model may copy (or attempt to copy) one or more conditional data values ​​in the one or more artificial data records. The number of artificial data records generated by the generator sub-model may be proportional to the batch size. For example, if the batch size is one, the generator sub-model may generate a single artificial data record, while if the batch size is 100, the generator sub-model may generate 100 artificial data records.

[0166] At step 910, the computer system may generate one or more comparisons using one or more selected sampled data records, one or more artificial data records, and a discriminator sub-model. The discriminator sub-model may be characterized by a plurality of discriminator parameters, such as a plurality of discriminator neural network weights representing a neural network-based discriminator sub-model. These comparisons may include classification outputs generated by the discriminator for one or more artificial data records, or for one or more pairs of artificial data records and the selected sampled data records. For example, for a particular artificial data record, the discriminator sub-model may generate a comparison such as "artificial, 80%," indicating that the discriminator sub-model classified the artificial data record as artificial with 80% confidence. As another example, for two data records "A" and "B," one of which is an artificial data record and the other is a selected sampled data record, the discriminator sub-model may generate a comparison such as "B, artificial, 65%," indicating that, of the two provided data records "A" and "B," the discriminator predicted that "B" is an artificial data record with 65% confidence.

[0167] At step 912, the computer system may determine a plurality of loss values. This plurality of loss values ​​may include a generator loss value and a discriminator loss value. The computer system may determine the plurality of loss values ​​based on one or more comparisons between one or more artificial data records generated during training (e.g., at step 908) and one or more sampled data records in the plurality of sampled data records. These loss values ​​may generally be used to evaluate the performance of the generator sub-model and the discriminator sub-model, which may be used to update the parameters of the generator sub-model and the discriminator sub-model in order to improve their performance. Thus, these loss values ​​may be proportional to the difference between the ideal or expected performance of the generator sub-model and the discriminator sub-model. For example, if the discriminator predicts (as indicated by a comparison in one or more comparisons) that the artificial data value is the artificial data value with a high confidence level (e.g., 99%), then the discriminator is generally successful in its intended function of discriminating between the artificial data record and the sampled data record. Thus, the discriminator loss value may be low (indicating that little change is required in the discriminator parameters).

[0168] Alternatively, if the discriminator predicts that the artificial data value is the real data value with high confidence, then the discriminator not only misidentified the artificial data record, but its misidentification was also very credible. Therefore, the discriminator loss value can be high (indicating that a large change in the discriminator parameters is required). Similar reasoning can be applied to the generator loss value, that is, if the generator generates artificial data records that successfully deceive the discriminator, then the generator loss value can be low, otherwise the generator loss value can be high. For batches greater than one, the generator loss value and the discriminator loss value can be based on the average of the generator and discriminator performance over all one or more sampled data records and one or more artificial data records.

[0169] The computer system can now determine a plurality of model update values ​​that can be used to update the machine learning model. These can include a plurality of generator update values ​​that can be used to update generator parameters and, thereby, update the generator submodel. Similarly, these model update values ​​can include a plurality of noise discriminator update values ​​that can be used to update discriminator parameters and, thereby, update the discriminator submodel.

[0170] At step 914, the computer system may generate one or more generator update values ​​based on the generator loss values. The computer system may use a generator optimizer component or software routine (e.g., Figure 3 ) to generate one or more generator update values. This generator optimizer can implement any suitable optimization method, such as stochastic gradient descent. In such cases, the one or more generator update values ​​may include one or more generator gradients or one or more values ​​derived from one or more generator gradients. Broadly speaking, a computer system can use a generator optimizer to determine which changes in generator model parameters result in the largest immediate reduction in generator loss value (e.g., determined based on the gradient of the generator loss value), and the generator model update value can reflect, indicate, or otherwise be used to implement the said change in the generator parameters.

[0171] At step 916, the computer system may generate one or more initial discriminator update values ​​based on the discriminator loss value. The computer system may use a discriminator optimizer component or software module (e.g., Figure 3 ) to generate one or more discriminator update values. The discriminator optimizer can implement any suitable optimization method, such as stochastic gradient descent. In such cases, the one or more initial discriminator values ​​can include one or more discriminator gradients or one or more values ​​derived from one or more discriminator gradients. Broadly speaking, a computer system can use the discriminator optimizer to determine which changes in discriminator model parameters result in the largest immediate decrease in the initial discriminator loss value (e.g., determined based on the gradient of the initial discriminator loss value), and the discriminator model update value can reflect, indicate, or otherwise be used to implement the change in the discriminator parameters.

[0172] At step 918, the computer system may generate one or more discriminator noise values. These discriminator noise values ​​may include random or pseudo-random numbers sampled from a Gaussian distribution (sometimes referred to as a "second Gaussian distribution") to distinguish them from the Gaussian distribution used for the sampled class noise values ​​(as described above with reference to Figure 4To generate one or more discriminator noise values, the computer system may determine a discriminator standard deviation. The second Gaussian distribution may have a mean of zero and a standard deviation equal to the discriminator standard deviation. The discriminator standard deviation may be based (in whole or in part) on the specific privacy requirements of the system, including those indicated by a pair of (ε, δ) differential privacy parameters. For example, the computer system may determine a larger standard deviation for more stringent privacy requirements and a smaller standard deviation for less stringent privacy requirements.

[0173] At step 920, the computer system may generate one or more noisy discriminator update values ​​(sometimes more generally referred to as "discriminator update values") by combining the one or more initial discriminator update values ​​with the one or more discriminator noise values. This may be accomplished by calculating one or more sums of the one or more initial discriminator update values ​​and the one or more discriminator noise values, and the one or more noisy discriminator update values ​​may include these sums. As described above (see, e.g., Section ID), adding noise to these discriminator model update values ​​may help achieve differential privacy.

[0174] Once model update values ​​(e.g., one or more generator update values ​​and one or more discriminator update values) have been determined, the computer system can update multiple model parameters (e.g., multiple generator parameters and multiple discriminator parameters) based on these model update values.

[0175] refer to Figure 9B , at step 922, the computer system may update the generator sub-model by updating multiple generator parameters using one or more generator update values.

[0176] At step 924, the computer system may update the discriminator sub-model by updating the plurality of discriminator parameters using one or more discriminator update values. This update process may depend on the specific properties of the generator and discriminator sub-models, their model parameters, and update values. As a non-limiting example, for generator and discriminator sub-models based on an artificial neural network architecture (e.g., as in a GAN), techniques such as backpropagation may be used to update the generator and discriminator model parameters based on the generator update values ​​and the discriminator update values.

[0177] Optionally, at step 926, the computer system may perform a privacy analysis of the model training. In non-private machine learning applications, the training phase is typically performed for a set number of training rounds, or until the model parameters have converged, e.g., do not change much (or at all) in successive training rounds. However, as described above, the privacy of a machine learning model is proportional to the probability that a particular data value or data record is sampled during training. The more training rounds that are performed, the greater the probability that a given data record or data value is sampled, and therefore, any privacy guaranteed by the machine learning model typically degrades with each successive training round (see, e.g., [2] for more details).

[0178] Therefore, a privacy analysis can be performed to determine how much of the "privacy budget" has typically been used during training. To perform this privacy analysis, the computer system can determine one or more privacy parameters corresponding to the current state of the machine learning model. These one or more privacy parameters can include an epsilon privacy parameter and a delta privacy parameter that can characterize the differential privacy of the machine learning model. The computer system can compare the one or more privacy parameters to one or more target privacy parameters, which can include a target epsilon privacy parameter and a target delta privacy parameter. If the epsilon privacy parameter and the delta privacy parameter equal or exceed their respective target privacy parameters, this can indicate that further training may violate any differential privacy requirements placed on the system.

[0179] At step 928, the computer system may determine whether a termination condition has been met. The termination condition may define a condition under which training is complete. For example, some machine learning model training programs involve training the model for a predetermined number of training rounds or epochs. In such cases, determining whether the termination condition has been met may include determining whether the current number of training rounds or the current number of training rounds is greater than or equal to a predefined number of training rounds or a predefined number of training epochs.

[0180] As another example, if the computer system performs a privacy analysis at step 926, the computer system may compare one or more privacy parameters (e.g., epsilon and delta privacy parameters) to one or more target privacy parameters (e.g., target epsilon and target delta privacy parameters) and determine that a termination condition has been met if the one or more privacy parameters are greater than or equal to the one or more target privacy parameters.

[0181] If the termination condition has not been met, the computer system can proceed to step 930 and repeat the iterative training process. The computer system can return to step 902 and select new sampled data records for subsequent training rounds. The computer system can repeat steps 902 to 928 until the termination condition has been met. Otherwise, if the termination condition has been met, the computer system can proceed to step 932 and terminate the iterative training process. The generator submodel can now be used to generate representative differentially private artificial data records.

[0182] V. After training

[0183] After training a machine learning model to generate a plurality of artificial data records, the machine learning model can be referred to as a "trained machine learning model." A component of the machine learning model, such as a generator sub-model, can be referred to as a "trained generator." Any artificial data records generated by such a trained generator can protect the privacy of the sampled data records used to train the machine learning model based on any privacy parameters used during this training process. Thus, the trained generator or artificial data records generated by the trained generator can be securely used, for example, by an artificial data-consuming entity, such as Figure 1 Described in.

[0184] Optionally, at step 934, the computer system may publish the trained generator (e.g., on a publicly accessible website or database). Alternatively, the computer system may transmit the trained generator to a client computer. The client computer may then use the trained generator to generate an artificial data set comprising a plurality of output artificial data records. In some cases, the client computer may generate and use its own condition vector (which may be different from and independent of the condition vector used during model training) to encourage the trained generator to generate artificial data records with specific characteristics. For example, if a medical research organization is interested in statistical analysis of health characteristics of elderly individuals, the medical research organization may use the condition vector to cause the trained generator to generate artificial data records corresponding to the elderly individuals.

[0185] Alternatively, at step 936, the computer system can use the trained machine learning model (e.g., a trained generator) to generate an artificial data set comprising a plurality of output artificial data records. The computer system can then transmit this artificial data set to the client computer at step 938. The client computer can then use this artificial data set as needed. For example, a client associated with the client computer can use the artificial data set to train a machine learning model to perform some useful function on the data set or perform statistical analysis, such as described in Section 1. These artificial data records protect the privacy of any sampled data records used to train the machine learning model, regardless of the nature of the post-processing performed by the client computer.

[0186] VI. Computer Systems

[0187] Any computer system mentioned herein may utilize any suitable number of subsystems. Figure 10 An example of such a subsystem in computer system 1000 is shown in FIG. In some embodiments, the computer system includes a single computer device, wherein a subsystem may be a component of the computer device. In other embodiments, the computer system may include multiple computer devices with internal components, each of which is a subsystem. The computer system may include desktop and laptop computers, tablet computers, mobile phones, and other mobile devices.

[0188] Figure 10 The subsystems shown in FIG1002 are interconnected by a system bus 1012. Additional subsystems are shown, such as a printer 1008, a keyboard 1018, a storage device 1020, a monitor 1024 (e.g., a display screen, such as an LED) coupled to a display adapter 1014, and the like. Peripheral devices and I / O devices coupled to the input / output (I / O) controller 1002 can be connected to the computer system through various means known in the art, such as input / output (I / O) ports 1016 (e.g., USB, ). For example, I / O port 1016 or external interface 1022 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 1000 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus 1012 allows central processing unit 1006 to communicate with each subsystem and control the execution of multiple instructions from system memory 1004 or storage device 1020 (e.g., a fixed disk such as a hard drive or optical disk), as well as the exchange of information between subsystems. System memory 1004 and / or storage device 1020 can embody computer-readable media. Another subsystem is data collection device 1010, such as a camera, microphone, accelerometer, etc. Any data mentioned herein can be output from one component to another and can be output to a user.

[0189] The computer system may include multiple identical components or subsystems, which are connected together, for example, via external interfaces 1022, via internal interfaces, or via removable storage devices that can be connected and removed from one component to another. In some embodiments, the computer systems, subsystems, or devices may communicate via a network. In such cases, one computer may be considered a client and another computer may be considered a server, where each computer may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0190] Any computer system mentioned herein can use any suitable number of subsystems. In some embodiments, the computer system includes a single computer device, wherein the subsystem can be a component of the computer device. In other embodiments, the computer system can include multiple computer devices with internal components, each of which is a subsystem.

[0191] A computer system may include multiple components or subsystems connected together, for example, by external interfaces or by internal interfaces. In some embodiments, the computer systems, subsystems, or devices may communicate over a network. In such cases, one computer may be considered a client and another computer may be considered a server, where each computer may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0192] It should be understood that any embodiment of the present invention can be implemented in the form of control logic using hardware (e.g., an application specific integrated circuit or a field programmable gate array) and / or using computer software, wherein a general-purpose programmable processor is modular or integrated. As used herein, a processor includes a single-core processor, a multi-core processor on the same integrated chip, or a plurality of processing units on a single circuit board or networked. Based on the present disclosure and the teachings provided herein, those of ordinary skill in the art will know and understand other ways and / or methods of implementing embodiments of the present invention using hardware and combinations of hardware and software.

[0193] Any of the software components or functions described in this application can be implemented as software code to be executed by a processor using any suitable computer language such as, for example, Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python using, for example, conventional or object-oriented techniques. The software code can be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission, suitable media including random access memory (RAM), read-only memory (ROM), magnetic media (such as a hard drive or floppy disk), or optical media (such as a compact disc (CD) or digital versatile disc (DVD)), flash memory, etc. The computer-readable medium can be any combination of such storage or transmission devices.

[0194] Such program can also use the carrier signal that is suitable for transmitting via the wired network, optical network and / or wireless network that meet multiple protocols including the Internet to encode and send.Therefore, the computer-readable medium according to an embodiment of the present invention can use the data signal with such program encoding to create.The computer-readable medium with program code encoding can be encapsulated together with compatible devices or (for example, downloading via the Internet) is provided separately with other devices.Any such computer-readable medium can reside on or in a single computer product (for example, hard disk drive, CD or whole computer system), and can be present on or in the different computer products in a system or network.A computer system may include a monitor, printer or other suitable displays for providing any result mentioned herein to the user.

[0195] Any method described herein can be performed completely or in part with a computer system comprising one or more processors that can be configured to perform these steps. Therefore, an embodiment may relate to a computer system that is configured to perform the steps of any method described herein, and may have different components that perform the corresponding steps or corresponding step groups. Although presented in numbered steps, the method steps herein can also be performed simultaneously or in different orders. In addition, parts of these steps can be used together with parts of other steps of other methods. In addition, all or part of the steps can be optional. In addition, any step of any method can be performed with modules, circuits or other means for performing these steps.

[0196] Without departing from the spirit and scope of the embodiments of the present invention, the specific details of the specific embodiments can be combined in any suitable manner. However, other embodiments of the present invention may relate to specific embodiments associated with each individual aspect, or specific combinations of these individual aspects. The above description of exemplary embodiments of the present invention has been presented for the purpose of illustration and description. It is not intended to be exhaustive or to limit the present invention to the precise form described, and many modifications and variations are possible in light of the teachings above. These embodiments are selected and described in order to best explain the principles of the present invention and their practical application, so that those skilled in the art can best utilize the present invention in various embodiments and make various modifications suitable for the intended specific use.

[0197] The above description is illustrative and not restrictive. After reading this disclosure, many variations of the present invention will become apparent to those skilled in the art. Therefore, the scope of the present invention should not be determined with reference to the above description, but should be determined with reference to the pending claims and their full scope or equivalents.

[0198] One or more features from any embodiment may be combined with one or more features of any other embodiment without departing from the scope of the present invention.

[0199] Unless expressly indicated to the contrary, the recitation of "a," "an," or "the" is intended to mean "one or more." Unless expressly indicated to the contrary, the use of "or" is intended to mean an inclusive or rather than an exclusive or.

[0200] All patents, patent applications, publications, and descriptions mentioned herein are incorporated by reference in their entirety for all purposes. No admission is made that they are prior art.

[0201] VII. References

[0202] [1]Dwork, Cynthia and Aaron Roth. "The Algorithmic Foundations ofDifferential Privacy." Foundations and Trends in Theoretical ComputerScience9, Issues 3-4 (2014): 211-407.

[0203] [2]Abadi, Martin, Andy Chu, Ian Goodfellow, H. Brendan McMahan, Ilya Mironov, Kunal Talwar and Li Zhang. “Deep Learning with Differential Privacy.” In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 308–318. 2016

[0204] [3]Goodfellow, Ian, Jean Pouget - Abadie, Mehdi Mirza, Bing Xu, David Warde - Farley, Sherjil Ozair, Aaron Courville and Yoshua Bengio. “Generative Adversarial Nets.” In Advances in Neural Information Processing Systems, pp. 2672–2680. 2014.

[0205] [4]Xie, Liyang, Kaixiang Lin, Shu Wang, Fei Wang and Jiayu Zhou. “Differentially Private Generative Adversarial Network.” arXiv prepritn arXiv:1802.06739 (2018).

[0206] [5]Xu, Lei, Maria Skoularidou, Alfredo Cuesa - Infante and Kalyan Veeramachaneni. “Modeling Tabular Data Using Conditional GAN.” In Advances in Neural Information Processing Systems, pp. 7335–7345. 2019.

[0207] [6]Li, Qiongxiu, Jaron Skovsted Gundersen, Katrine Tjell, Rafal Wisniewski, Mads Christensen, “Privacy-Preserving DistributedExpectation Maximization for Gaussian Mixture Model Using SubspacePerturbation,” in ICASSP 2022-2022IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Singapore, Singapore, 2022, pp. 4263–4267. Digital Object Identifier: 10.1109 / ICASSP43922.2022.9746144

[0208] [7] Shashanka, Madhusudana, “A Privacy Preserving Framework for Gaussian Mixture Models,” 2010 IEEE International Conference on Data Mining Workshops, Sydney, New South Wales, Australia, 2010, pp. 499–506, DIO: 10.1109 / ICDMW.2010.109

Claims

1. A method, performed by a computer system, for training a machine learning model to generate a plurality of artificial data records in a privacy-preserving manner, the method comprising: Retrieving a plurality of data records, each data record comprising a plurality of data values ​​corresponding to a plurality of data fields, each data value being within a category of a plurality of categories; determining a plurality of noise class counts corresponding to the plurality of classes, each noise count indicating a number of data records in the plurality of data records belonging to each class in the plurality of classes; identifying one or more defect categories, each defect category including a category whose corresponding noise category count is less than a minimum count; combining each defect class of the one or more defect classes with at least one other class of the plurality of classes to thereby determine a plurality of combined classes; identifying one or more defective data records from the plurality of data records, each defective data record containing at least one defective data value corresponding to the combined class; for each defective data value contained in the one or more defective data records, replacing the defective data value with a combined data value identifying a combined category of the plurality of combined categories; generating a plurality of condition vectors, each condition vector identifying one or more specific data fields of the plurality of data fields for replication; sampling a plurality of sampled data records from the plurality of data records, wherein the plurality of sampled data records includes at least one of the one or more defective data records, each sampled data record including a plurality of sampled data values ​​corresponding to the plurality of data fields; as well as The machine learning model is trained using the multiple sampled data records to generate the multiple artificial data records, each artificial data record including multiple artificial data values ​​corresponding to the multiple data fields, wherein in each artificial data record in the multiple artificial data records, the machine learning model replicates one or more sampled data values ​​corresponding to the one or more specific data fields in the multiple artificial data values ​​of a specific sampled data record according to the multiple condition vectors, wherein the machine learning model is trained based on a comparison between the multiple artificial data records and the multiple sampled data records.

2. The method of claim 1 , wherein the machine learning model comprises a trained machine learning model after the step of training the machine learning model to generate the plurality of artificial data records, and wherein the method further comprises: generating an artificial dataset comprising a plurality of output artificial data records using the trained machine learning model; as well as The artificial data set is transmitted to a client computer.

3. The method of claim 1 , further comprising, before the step of training the machine learning model to generate the plurality of artificial data records: identifying one or more identified sampled data values, each identified sampled data value corresponding to a corresponding combined category of the one or more corresponding combined categories; and For each identified sampled data value: determining two or more categories that are combined to create the corresponding combined category, randomly selecting a random category from the two or more categories, generating a sampled data value with replacement identifying the random class, and The identified sampled data value is replaced with the replacement sampled data value.

4. The method of claim 1, wherein each noise class count comprises a sum of a class count in a plurality of class counts and a class noise value in one or more class noise values, wherein each class noise value is defined by a class noise mean and a class noise standard deviation.

5. The method according to claim 4, further comprising: generating the one or more class noise values ​​by sampling the one or more class noise values ​​from a first Gaussian distribution having the class noise mean and the class noise standard deviation; as well as The class noise mean and the class noise standard deviation are determined based on one or more class noise parameters including one or more target privacy parameters.

6. The method of claim 5 , wherein the one or more target privacy parameters include an epsilon privacy parameter and a delta privacy parameter, and wherein the one or more class noise parameters include the epsilon privacy parameter, the delta privacy parameter, the minimum count, the maximum number of non-zero data values, a safety margin, and the total number of data values ​​in a data record.

7. The method of claim 1, wherein the machine learning model comprises an autoencoder, a generative adversarial network, or a combination of the autoencoder and the generative adversarial network.

8. The method of claim 1 , further comprising, before determining the plurality of noise category counts: for each data record in the plurality of data records, determining whether the data record contains more than a maximum number of non-zero data values; and For each data record containing more than the maximum number of non-zero data values, the data record is removed from the plurality of data records.

9. The method of claim 1 , further comprising, before determining the plurality of noise category counts: for each data record in the plurality of data records, normalizing one or more data values ​​between 0 and 1, inclusive, thereby generating one or more normalized data values; for each of the one or more normalized data values, determining a plurality of normalized categories based on a corresponding probability distribution of one or more probability distributions; as well as The plurality of normalized categories are included in the plurality of categories.

10. A method according to claim 9, wherein each of the one or more probability distributions comprises a multimodal distribution having a predetermined number of equally weighted modes, wherein the number of normalized categories is equal to the predetermined number of equally weighted modes, such that for each normalized category in the plurality of normalized categories, there is a corresponding mode in the predetermined number of equally weighted modes.

11. The method of claim 1 , wherein the machine learning model is characterized by a plurality of model parameters; and Training the machine learning model includes an iterative training process, and the iterative training process includes: determining a plurality of loss values ​​based on the comparison between the plurality of artificial data records and the plurality of sampled data records, determining a plurality of model update values ​​based on the plurality of loss values, updating the plurality of model parameters based on the plurality of model update values, Determine whether the termination conditions have been met, and If the termination condition is met, the iterative training process is terminated; otherwise, the iterative training process is repeated until the termination condition is met.

12. The method according to claim 11, wherein: The machine learning model includes a generative adversarial network, wherein the generative adversarial network includes a generator sub-model and a discriminator sub-model; The plurality of model parameters include a plurality of generator parameters characterizing the generator sub-model and a plurality of discriminator parameters characterizing the discriminator sub-model; The plurality of loss values ​​include a generator loss value and a discriminator loss value; The plurality of model update values ​​includes one or more generator update values ​​and one or more discriminator update values; Determining a plurality of model update values ​​based on the plurality of loss values ​​includes: determining the one or more generator update values ​​based on the generator loss values, and determining the one or more discriminator update values ​​based on the discriminator loss values; and Updating the plurality of model parameters based on the plurality of model update values ​​includes updating the plurality of generator parameters using the one or more generator update values ​​and updating the plurality of discriminator parameters using the one or more discriminator update values.

13. The method of claim 12, wherein after training the machine learning model, the generator sub-model comprises a trained generator, and wherein the method further comprises: The trained generator is transmitted to a client computer, wherein the client computer uses the trained generator to generate an artificial data set comprising a plurality of output artificial data records.

14. The method of claim 12 , wherein the one or more discriminator update values ​​comprise one or more noise discriminator update values, the one or more noise discriminator update values ​​comprising a sum of one or more initial discriminator update values ​​and one or more discriminator noise values, and wherein determining the one or more discriminator update values ​​based on the discriminator loss value comprises: determining the one or more initial discriminator update values ​​based on the discriminator loss value; Determine the discriminator standard deviation; generating the one or more discriminator noise values ​​by sampling from a second Gaussian distribution having a mean of zero and a standard deviation equal to the discriminator standard deviation; as well as The one or more noise discriminator update values ​​are determined by calculating one or more sums of the one or more initial discriminator update values ​​and the one or more discriminator noise values.

15. The method of claim 12, wherein: The generator sub-model is implemented using a generator artificial neural network; The plurality of generator parameters includes a plurality of generator weights corresponding to the generator artificial neural network; The one or more generator update values ​​include one or more generator gradients or one or more values ​​derived from the one or more generator gradients; The discriminator sub-model is implemented using a discriminator artificial neural network; The plurality of discriminator parameters includes a plurality of discriminator weights corresponding to the discriminator artificial neural network; and The one or more discriminator update values ​​include one or more discriminator gradients or one or more values ​​derived from the one or more discriminator gradients.

16. The method of claim 11 , wherein the iterative training process comprises a plurality of training rounds or a plurality of training epochs, and wherein determining whether the termination condition has been met comprises determining whether a current number of training rounds or a current number of training epochs is greater than or equal to the number of training rounds or the number of training epochs.

17. The method of claim 11, wherein determining whether the termination condition has been satisfied comprises: determining one or more privacy parameters corresponding to a current state of the machine learning model; as well as The one or more privacy parameters are compared to one or more target privacy parameters, wherein the termination condition has been met if the one or more privacy parameters are greater than or equal to the one or more target privacy parameters.

18. The method of claim 17, wherein the one or more privacy parameters include an epsilon privacy parameter and a delta privacy parameter, and wherein the one or more target privacy parameters include a target epsilon privacy parameter and a target delta privacy parameter.

19. A method for training a machine learning model to generate a plurality of artificial data records that protect the privacy of sampled data values ​​contained in a plurality of sampled data records, the method being performed by a computer system and comprising: Obtaining the plurality of sampled data records, each sampled data record comprising a plurality of sampled data values; Obtaining a plurality of condition vectors, each condition vector identifying one or more specific data fields; as well as Perform an iterative training process, the iterative training process comprising: determining one or more selected sampled data records from the plurality of sampled data records, determining one or more selected conditional vectors from the plurality of conditional vectors, identifying one or more conditional data values ​​from the one or more selected sampled data records, the one or more conditional data values ​​corresponding to one or more specific data fields identified by the one or more selected conditional vectors, generating one or more artificial data records using the one or more conditional data values ​​and a generator sub-model, wherein the generator sub-model is characterized by a plurality of generator parameters, generating one or more comparisons using the one or more selected sampled data records, the one or more artificial data records, and a discriminator sub-model, wherein the discriminator sub-model is characterized by a plurality of discriminator parameters, determining a generator loss value and a discriminator loss value based on the one or more comparisons, generating one or more generator update values ​​based on the generator loss value, generating one or more initial discriminator update values ​​based on the discriminator loss value, generate one or more discriminator noise values, generating one or more noise discriminator update values ​​by combining the one or more initial discriminator update values ​​with the one or more discriminator noise values, updating the generator submodel by updating the plurality of generator parameters using the one or more generator update values, updating the discriminator sub-model by updating the plurality of discriminator parameters using the one or more noise discriminator update values, Determine whether the termination conditions have been met, and If the termination condition is met, the iterative training process is terminated; otherwise, the iterative training process is repeated until the termination condition is met.

20. A computer system, comprising: processor; as well as A non-transitory computer-readable medium coupled to the processor, the non-transitory computer-readable medium comprising code executable by the processor to implement the method of any one of claims 1 to 19.