System and method for generating synthetic data using a masked transformer
A transformer-based architecture addresses the challenges of generating robust, scalable, and privacy-preserving synthetic tabular data by embedding, masking, and predicting data values, resulting in high-quality synthetic data that maintains original data properties and correlations.
Patent Information
- Application Number
- JP2024565948
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-09
- Filing Date
- 2023-05-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing deep tabular generators face challenges in generating robust, scalable, and privacy-preserving synthetic data, particularly when handling missing data in fields like healthcare, finance, and social science.
A transformer-based modeling architecture that includes an embedding model for data embedding, a masking model for handling missing data, and a transformer model for predicting actual data values, enabling the generation of synthetic tabular data that is robust, scalable, and privacy-preserving.
The proposed solution effectively generates high-quality synthetic data that maintains the statistical properties and correlations of the original data, while ensuring privacy and handling missing data efficiently, outperforming state-of-the-art methods in various benchmarks.
Smart Images

Figure 2025516532000001_ABST
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 339,878, entitled "Modeling and Generation of Network Traffic For Network Security", filed on May 9, 2022, which is hereby incorporated by reference in its entirety.
Background Art
[0002] The general technical field of the exemplary embodiments is synthetic data generation, and more particularly, a transformer - based model for generating synthetic tabular data.
[0003] In recent years, generative models have received significant attention in the field of deep learning due to their ability to synthesize high - quality data and learn the structure underlying complex datasets. Such models have been successful in applications to various data types, including images, text, and tabular data.
[0004] The development of an effective synthetic tabular data generator is significant for a number of reasons, including privacy preservation, data augmentation, model interpretability, and anomaly detection. Conventional research in this domain has produced numerous generative models, also known as deep tabular generators, including Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Autoregressive Transformers, and Diffusion models. These existing deep tabular generator models are attempting to address the challenges associated with tabular data generation, but there is still room for exploration and improvement, particularly in terms of robustness, scalability, and privacy preservation, especially when handling missing data. These characteristics are particularly important in fields such as healthcare, finance, and social science where tabular data is prevalent. The heterogeneous nature of tabular data is characterized by its diverse data types, distributions, and relationships, presenting different challenges not found in other domains.
[0005] Therefore, there is a need in the art for a synthetic tabular data generator that can effectively generate robust synthetic data, is scalable to the underlying training data, and preserves privacy. SUMMARY OF THE INVENTION
[0006] In a first exemplary embodiment, a transformer-based modeling architecture for training a generator to generate synthetic data includes at least one embedding model for embedding input data, where the input data includes a plurality of fields F storing real data values i,j and the at least one embedding model constructs l embedding matrices, one for each of the plurality of fields F i,j of the data rows F in the embedded input data seti A set of unmasked fields for each
Number
Number
Number
Number
Number
Number
[0007] In a second exemplary embodiment, the process for training a transformer-based generator model to generate synthetic data is: embedding the input data by at least one embedding model, where the input data includes a plurality of fields F storing actual data values i,j and the at least one embedding model constructs l embedding matrices, one for each of the plurality of fields F i,j ; by a masking model, for each row F of data in the embedded input data set i a set of masked fields for each
Number
Number
Number
Number
Number
[0008] In a third exemplary embodiment, a transformer-based model for generating synthetic data is: at least one embedding model for embedding an input data set, where the input data set has a predetermined format and includes a plurality of fields F indicating real values i,j Including, and the at least one embedding model constructs l embedding matrices, one for each of the plurality of fields F i,j Receiving a first embedded input data set, where each real-world value for each of the plurality of fields F is masked with a mask token, and until all mask tokens in the first embedded input data set contain synthetic data values, the masked fields over a plurality of iterations i,j Receiving, and
Number
Brief Description of the Drawings
[0009] The exemplary embodiments will be more fully understood from the detailed description given hereinbelow and the accompanying drawings, in which like elements are represented by like reference numerals, which are given by way of illustration only and, therefore, do not limit the exemplary embodiments herein.
[0010]
Figure 1a
Figure 1b
[0011]
Figure 2a
Figure 2b
[0012]
Figure 3a
Figure 3b
Figure 3c
Figure 3d
Figure 3e
Figure 3f
Figure 3g
Figure 3h
Figure 3i
Figure 3j
Figure 3k
Figure 3l
[0013]
Figure 4a
Figure 4b
Figure 4c
Figure 4d
Figure 4e
Figure 4f
[0014]
Figure 5a
Figure 5b
[0015]
Figure 6a
Figure 6b
[0016] **Definitions** Field: A column of values within a dataset that stores values of all the same type.
[0017] Categorical field: A field within the data that takes values from a set of fixed values and has no clear order between categories, e.g., gender (male, female).
[0018] Numeric field: A field that takes values from real numbers or integer values. For example, housing price (400k, 225k...).
[0019] Embedding: A table of values that enables the conversion of each categorical value into a vector. Each unique categorical value is mapped to a unique vector. Male -> [0.5, -0.1...], Female -> [0.8, 0.2...].
[0020] Quantizer: A block responsible for converting real or integer values into a set of fixed values. The inventors do this so that they can use the embedding before feeding the values into the transformer. In this paper, the inventors use K-Means, but they have tried many different quantizers in their design. The exemplary quantizer will map all values to something in [0.0, 0.1, 0.2, 0.3...].
[0021] Ordered embedding: A special embedding unique to the inventors' design that enables the order of values to be taken into account in the embedding.
[0022] Masking: Replacing a value with a mask token that indicates to the model that the value is missing. The model then learns to predict the missing value.
[0023] Transformer: A general neural network that specifically attempts to learn the relationships between input entries.
[0024] Dynamic linear layer: Typical models use fixed linear layers, but the inventors use dynamic linear layers generated on the fly instead. It is generated using the same procedure as the ordered embedding.
[0025] Synthetic data: A dataset of generated data from a model that is statistically similar to real data.
[0026] The Transformer, designed for natural language processing tasks, has achieved great success in the NLP (Natural Language Processing) domain and has brought significant progress in various applications. Due to their powerful ability to model complex dependencies and generalize across applications, researchers have been driven to extend the Transformer to other data types such as images and audio. The initial description of the Transformer concept can be found in Vaswani et al., Attention Is All You Need, arXiv:1706.03762v5 [cs.CL] 6 Dec 2017, which is incorporated herein by reference.
[0027] In the embodiments described herein, the Transformer is used as a synthetic tabular data generator, and their cross-domain applicability is further extended. Specifically, the embodiments implement a Masked Transformer (MT) as a tabular data generator that achieves state-of-the-art performance across rich datasets.
[0028] The embodiments described herein implement an MT architecture design, referred to throughout this specification as TabMT. TabMT is general enough to function across many tasks and scenarios. A high-level diagram of the TabMT model training process 1 is shown in FIG. 1a. FIG. 1a includes a set of complete real datasets 5 related to the domain of interest. In the specific example discussed herein, the domain dataset supports tabular data, which includes both categorical and numerical data fields populated with relevant sample data. Each field in the dataset is a column of values that all store values of the same type. As an example, a categorical field is a field within the data that takes values from a set of fixed values and has no clear order between the categories, such as, for example, gender (male, female), eye color, true / false, etc. On the other hand, a numerical field is a field that takes values from real or integer values, such as, for example, housing prices (400k, 225k...), a person's height and weight, temperature, etc. The datasets of the sample data are both batched (10).
[0029] Next, the numerical data from the batch is fed to a quantizer block (e.g., K-Means) 15 to convert real or integer values to a set of fixed values (e.g., map all values to something in [0.0, 0.1, 0.2, 0.3…]), and then an ordered embedding 20 is applied. This ordered embedding makes it possible for the order of the values to be taken into account in the embedding. The categorical data from the batch receives an embedding 25, where a table of values is used to convert each unique categorical value to a unique vector (e.g., the male "value" is mapped to the vector [0.5, -0.1...], and the female "value" is mapped to the vector [0.8, 0.2...]).
[0030] Next, the embedded batch 30, which contains both the embedded numerical data and the embedded categorical data, undergoes masking to produce a masked batch 35, where the values within the embedded batch data are replaced with mask tokens, which indicates to the model that the masked values are missing. The model then learns to predict the missing values. As an illustration, if the original value is 0.5, this would first be mapped to a vector if it were not masked. If it is masked, it is replaced with a vector corresponding to the missing value. The model would then need to predict that 0.5 was the original value based on the other unmasked values. Specifically to the preferred embodiment by the inventors, by masking a random percentage of tokens between 0% and 100%, the TabMT architecture can also be used to generate synthetic data (specifically tabular data) as further described herein below.
[0031] As part of the prediction process, the model uses a transformer 40, e.g., a neural network (NN), to learn the relationships between the input entries within the masked batch and maps the input data to output data using a dynamically generated linear layer 45 (instead of a fixed linear layer). The dynamic linear layer 45 is generated using the same procedure as the positional embedding.
[0032] Merely by way of example, FIG. 1b provides an example of an input training data set DS storing fields F1 to F5 before masking, and after random masking of a specific field in the input training data set DS M and provides an example after random masking of a specific field in the input training data set DS M wherein, in the input training data set DS m F2 mis masked for training the model. An exemplary dataset DS includes both categorical (F1) and numerical (F2 - F5) fields. This random masking of the dataset is repeated multiple times during training, and the training is based on other known field values in the set, e.g., F1, F3, and F5, to predict the masked field values, e.g., F2 m and F4 m to train the model to predict.
[0033] Referring to FIGS. 2a and 2b, the trained TabMT model can not only predict missing field values, but it can also generate synthetic data that is statistically similar to real data (e.g., DS). In FIG. 2a, the process for generating synthetic data 100 using the TabMT model is similar in most respects to the training process. In process 100, the numerical fields of the masked data batch 110 (see FIG. 2b) are processed by a quantizer 115, receive an ordered embedding 120, and the categorical data fields receive an embedding 125 to generate an embedded batch 130. The transformer 140 generates randomly predicted values for one or more masked fields according to its training, uses a dynamic linear layer 145 to map the randomly predicted values to output data, returns a randomly predicted dataset containing one or more predicted values, and returns to the start of the generation process as the next masked batch, and the process is repeated until one sample dataset is fully generated.
[0034] FIG. 2b provides a simplified diagram of the masking during the generation of the synthetic dataset SDS according to the process in FIG. 2a. In this example, in the initial masked batch dataset DS m1 all field values are masked (F1 m , F2 m , F3 m , F4 mand F5 m )。After iteration through the TabMT model process of Figure 2a, F2 s A composite value of 18 for is generated, and the following masked batch dataset DS m2 in which the field values are masked (F1 m , F2 s , F3 m , F4 m and F5 m ), and a composite value of 4.1 for the next field, e.g., F4 s is generated, and this continues until a complete synthetic dataset SDS is generated, as illustrated in Figure 2b.
[0035] TabMT uses Transformer-based modeling. In the prior art, Transformer-based language models have been used to learn human language via token prediction. Given the context of either the tokens that are temporally prior or a random subset of the tokens within a window, the model is tasked with predicting the remaining tokens using either an autoregressive model (see Borisov et al., Language Models Are Realistic Tabular Data Generators, arXiv:2210.06280v1 [cs.LG] 12 Oct 2022, which is incorporated herein by reference) or a masked model. Across a wide range of tasks and domains, both of these paradigms have shown success. Masked language models have conventionally been used for their embeddings, although some papers have explored masked generation of content, e.g., image generation, as described in Chang et al., Maskgit: Masked generative image transformer, In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 11315-11325 (2022). In this embodiment, TabMT demonstrates the effectiveness of Transformer-based modeling for modeling and generating tabular data.
[0036] This embodiment utilizes a variant of the pre-training of deep bidirectional transformers for language understanding (「BERT」), which establishes a masking procedure for training a bidirectional transformer. Given an N×l dataset F, for each row F i of the transformer, there is a set of unmasked fields
Number
Number
Number
Number
[0037] The BERT masking procedure generates a powerful embedding model but not a powerful generator. To understand this reason, we can focus on the distribution of the masked set. As a result of repeated Bernoulli trials during masking, the size of the masked set for each row
Number
Number
[0038] Conventional autoregressive generators generate fields sequentially from F i,0 ...F i,l‐1 However, unlike language, tabular data does not have an inherent order. By generating fields in a fixed order, another inconsistency between training and inference is introduced. During training,
Number
Number
Number
Number
[0039] The transformer model will typically have an N×d input embedding matrix E, where N is the number of unique input tokens and d is the transformer dimension. Since tabular data is heterogeneous, the inventors instead construct l embedding matrices, one for each field. Each embedding matrix will have a different number N of unique tokens.
[0040] For categorical fields, the inventors use a standard embedding matrix initialized with a normal distribution. For continuous, e.g., numerical, fields, the inventors construct an N×d ordered embedding O from its N×d unordered embedding matrix E and two d-dimensional endpoint vectors a and b.
[0041] To construct each ordered embedding matrix O, the inventors first use K-Means to cluster the values of the continuous field. The inventors consider the maximum number of clusters as a hyperparameter. Let the N-dimensional vector of the ordered cluster centers be v. The inventors use min-max normalization to construct an N-dimensional vector of ratio r as follows.
Equation
[0042] Relying too heavily on unordered embeddings can cancel out the advantages of ordered embeddings by the inventors, because information is not effectively shared between nearby values. To address this, the inventors bias TabMT to rely on order as much as possible. For continuous fields, the inventors zero-init the unordered embedding matrix E. In contrast, the endpoint vectors a and b use a normal distribution with a magnitude of 0.05. Additionally, the inventors include a learned temperature for the output, which allows the model to sharpen the predictive distribution as needed. The predictive distribution of each field
Number
Number
[0043] Therefore, the TabMT process differs from BERT in the following very important aspects: TabMT uses a separate embedding matrix for each field; to handle continuous entries, TabMT uses ordered embeddings and a quantizer, and uses a dynamic linear layer to pair with the ordered embeddings; TabMT masks randomly between 0% and 100% of the data, while BERT always masks 15% of its tokens; TabMT generates fields in a random order, while BERT cannot generate data and is not used to generate data. Additionally, BERT is designed for human language and does not predict (or generate) tabular data.
[0044] The TabMT structure is particularly well-suited for generating tabular data for several reasons. First, TabMT considers patterns between fields bidirectionally. The lack of order in tabular data means that bidirectional learning is likely to produce better understanding and embeddings within the model.
[0045] Second, the "prompt" to the tabular generator is likely not sequential. The masking procedure of TabMT enables any prompt to the model during generation. This is unique as most other generators have very limited conditional capabilities.
[0046] Third, missing data is far more common in tabular data than in other domains. TabMT can learn by setting their masking probabilities to 1 even when values are missing. Other generators require separately complementing the data until they can generate high-quality clean samples.
[0047] The following section presents a comprehensive evaluation of the effectiveness of TabMT across a wide range of tabular datasets. The analysis involves a thorough comparison with state-of-the-art methods that cover almost all generative model families. To ensure a robust evaluation, the inventors evaluate across several dimensions and metrics.
[0048] Regarding the data quality and privacy experiments described below, the inventors used the same dataset as that used in the recent TabDDPM paper by Kotelnlkov et al., TabDDPM: Modeling Tabular Data with Diffusion Models, arXiv:2209.15421v1 [cs.LG] 30 Sep 2022, which is hereby incorporated by reference in its entirety. As listed in Table 1, the sizes of these 15 datasets range from approximately 400 samples to approximately 150,000 samples. They include continuous features, categorical features, and integer features. The datasets range from 6 to 50 columns.
Table 1
[0049] For the scaling experiment, the CIDDS-001 dataset consisting of Netflow traffic from a simulated small business network is used. Netflow consists of 12 attributes (Table 2), and the inventors post-process these into 16 attributes (Table 3).
Table 2
Table 3
[0050] This dataset is extremely large, having over 30 million rows and a field cardinality of tens of thousands. All of the other listed datasets have a cardinality of less than 50. Unlike other benchmarks by the inventors, the inventors do not intentionally quantize continuous variables here in order to further test the scaling of the models by the inventors. In other words, all unique values are treated as separate classifications in the prediction process by the inventors.
[0051] For comparison, four prior arts are used, one from each of the main families of deep generative models.
[0052] TVAE is one of the first deep tabular generative techniques, introducing two models, GAN and VAE, as described by Lei Xu et al. in Modeling Tabular Data using Conditional GAN, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), which is incorporated herein by reference. As discussed below, TabMT is compared against VAE because it is the most powerful VAE tabular generator that the inventors are aware of.
[0053] CTABGAN+ described by Zhao et al. in Ctab-gan+: Enhancing tabular data synthesis, arXiv preprint arXiv:2204.00401 (2022) is considered to be the state of the art for GaN-based tabular data synthesis against which TabMT is compared herein, which is incorporated herein by reference.
[0054] TabDDPM, referred to above, has achieved the most powerful results so far, adapting diffusion models to tabular data. The results of TabMT are compared as explained below.
[0055] Finally, RealTabFormer, described in Solatorio et al., Realtabformer: Generating realistic relational and tabular data using transformers, arXiv preprint arXiv:2302.02041, (2023), is a concurrent study on adapting autoregressive transformers to tabular and relational data, which is hereby incorporated by reference. This method is most similar to TabMT; however, it uses an autoregressive transformer that demonstrates worse results than the masked transformers by the inventors.
[0056] The Catboost variant of ML efficiency is used to evaluate the quality of the synthetic data by the inventors. This metric trains a Catboost model on synthetic data instead of a weak ensemble. The Catboost model can pick up finer-grained patterns in data that weak classifiers cannot utilize. This is an overall metric that takes into account both the diversity and quality of the samples. For a fair comparison, the inventors use standard hyperparameters that tune the budget of 50 trials. The complete search space by the inventors can be seen in Table 4. For evaluation, the inventors trained the Catboost model 10 times on 5 samples of synthetic data to generate scores and standard deviations for the test set.
Table 4
[0057] The MLE scores are presented in Table 5; it should be noted that the inventors have outperformed or matched the state of the art for 11 out of 15 datasets. To gain a qualitative understanding of the data quality, the inventors visualize the distribution of correlation errors for datasets AD, BU, CA, CAR, DI, and KI, as shown in FIGS. 3a - 3l. Specifically, FIGS. 3a and 3b visualize the correlation errors between TabMT, TabDDPM, and CTabGAN+ for dataset AB; FIGS. 3c - 3d visualize the correlation errors between TabMT, TabDDPM, and CTabGAN+ for dataset BU; FIGS. 3e - 3f visualize the correlation errors between TabMT, TabDDPM, and CTabGAN+ for dataset CA; FIGS. 3g - 3h visualize the correlation errors between TabMT, TabDDPM, and CTabGAN+ for dataset CAR; FIGS. 3i - 3j visualize the correlation errors between TabMT, TabDDPM, and CTabGAN+ for dataset DI; FIGS. 3k - 3l visualize the correlation errors between TabMT, TabDDPM, and CTabGAN+ for dataset KI.
Table 5
[0058] To calculate the correlation error, the inventors first calculate the correlation r between each pair of fields i,j. i,j To calculate the correlation with categorical columns, the inventors convert them to one - hot vectors. The inventors then calculate the correlation
Equation
Equation
[0059] Maintaining the privacy of the original data is an important application for synthetic tabular data. Machine learning is expanding its applications across a wide range of fields to generate valuable insights. At the same time, both the regulations that need to be considered and the concerns regarding privacy are increasing rapidly. As demonstrated above, the data generated by TabMT is of sufficiently high quality to generate strong classifiers. Next, the inventors evaluate the TabMT model for privacy. This evaluation supplements the quality evaluation by the inventors and verifies that the model by the inventors is generating new data.
[0060] A high-quality non-private model can be trivially formed by directly replicating the training set. To ensure that the TabMT model is both private and high-quality, the inventors verify that the TabMT model learns the inherent structure of the data and does not simply memorize it. To evaluate privacy and novelty, the inventors adopt the Median Distance to the Closest Record (DCR) score. To calculate the DCR of the synthetic samples, the inventors find the closest data point in the real training set in terms of Euclidean distance. The inventors report the median of this value across the synthetic samples they generate. There is an inherent trade-off between privacy and quality. The higher the quality of the samples, the closer the points in the training set tend to be, and vice versa. While models such as CTabGAN+ and TabDDPM have a fixed trade-off between privacy and quality after training, TabMT can dynamically trade off between quality and privacy using temperature scaling. By using temperature scaling to move along the Pareto curve of the TabMT model, the inventors can controllably tune the privacy for each application. By scaling the temperature of the field higher, the values become more diverse and more private, but they also become less faithful to the true correlations in the data. The trade-off between quality and privacy here forms the Pareto front for TabMT for each dataset. The inventors use a separate temperature for each column and perform a small random search to find the Pareto front. In Table 6, the inventors compare the DCR and corresponding MLE score of TabMT with those of TabDDPM. The inventors can always obtain a higher DCR score and, in most cases, a higher MLE score as well. Figures 4a - 4f show the Pareto front of TabMT across several datasets including AD (Figure 4a); FB (Figure 4b); CAR (Figure 4c); MI (Figure 4d); DI (Figure 4e) and KI (Figure 4f).
Table 6
[0061] Real - world data often contains many missing values, which can make training difficult. If a row has missing values, one has to either drop that row or find a way to impute it. Other techniques such as RealTabformer or TabDDPM cannot inherently handle real - world missing data and have to either use different imputation techniques or drop the corresponding rows. The novel masking procedure described in the embodiments herein enables TabMT to inherently handle any missing data. To demonstrate this, the inventors randomly dropped 25% of the values from the dataset, ensuring that almost all rows have permanently missing data. Nevertheless, the model by the inventors can still generate and train on synthetic rows with no missing values, which facilitates training on real - world data. Table 7 shows the accuracy by the inventors when training using the missing data for datasets AD and KI.
Table 7
[0062] The provisional patent application for which this application claims the benefit of priority describes a specific use case where the TabMT model for synthetic data generation is applicable and very beneficial: network data. As discussed in this provisional application, having the ability to synthesize network data such as metadata for network traffic flows is useful for troubleshooting, detecting security incidents, planning, and billing.
[0063] In a computer network, network data is, in most cases, encapsulated within network packets that provide a load in the network. Network traffic is a main component for network traffic measurement, network traffic control, and simulation. Network traffic flows can be measured to understand what hosts are talking about on the network using details of the traffic's address, volume, and type. Data from this model can be used to train downstream models across a variety of tasks. The inventors can fine-tune for specific network data or condition on portions of the data by the inventors, enabling the data to be extended beyond what is available. As a result, a better downstream model is obtained than would otherwise be possible. This model can be used to directly perform anomaly detection because it can be used to measure the probability of new, unseen traffic, making it possible to discover rare traffic. The inventors can also generate a baseline for anomalous traffic for which other models are directly trained or for normal traffic that is generated. Furthermore, the embedding of the model by the inventors is very beneficial and can be directly used in the same way by other models.
[0064] Flow metadata describes the characteristics of packets in a flow. Since this information about the data (metadata) is at a higher level than individual packets, it takes up less space and is easier to analyze and scale better. Flow metadata can include aggregated totals, such as total data and packets, and environmental data such as time and interface. The types of flow metadata are: (1) generally invariant, including flow properties such as characteristics like IP protocol and address and process properties like interface / port, and (2) dynamic, including flow characteristics that can be measured across all packets in a flow. These flow characteristics can include the average / maximum / minimum / sum / difference of metrics such as packets, packet size, start time, end time, etc.
[0065] The network traffic of an organization is high in volume, which presents an initial barrier to measurement and monitoring. And the second barrier is data privacy, which restricts access to the organization's network data to protect the organization from, for example, cyber security threats and theft of business confidential information. Without access to accurately represent network data flows, i.e., data representing diverse real-world datasets, the ability to evaluate such data is necessarily limited. The TabMT model described herein can generate synthetic bidirectional flow data. TabMT is capable of generating the data according to the specification of its attributes when the data is generated. TabMT is capable of generating malicious or abnormal traffic based on new and evolving threats such as APT-29.
[0066] Netflow data is a specific type of tabular data that captures network communication events and is commonly used for network traffic analysis, intrusion detection, and cyber security applications. Netflow datasets have complex rules between fields and a large number of possible values for each field and are typically extremely large. Generating realistic synthetic netflow data is crucial for developing and testing network monitoring tools and security algorithms. The TabMT model functions well when scaling to very large datasets such as Netflow datasets. Table 8 shows the structure of a typical netflow. Netflow data includes both categorical and continuous attributes, and the range and number of unique values between these attributes vary dramatically. [Table 8]
[0067] The inventors use the CIDDS-001 dataset as a benchmark dataset by the inventors. These results, along with those discussed above, demonstrate that TabMT is sufficiently sample-efficient to learn using only a few hundred samples and, at the same time, maintains sufficient generality to scale to over 30 million samples.
[0068] For this example, the Netflow data is preprocessed such that it is categorical rather than ordinal or continuous. The date field is split into several fields representing the timestamp, i.e., the day of the week, hour, minute, second, and millisecond parts. In this study, the inventors assume that the traffic distribution is stationary across weeks and thus only track the day of the week. The inventors quantize the Byte, Duration, and Packet fields. However, for this dataset, the inventors are able to do this in a reversible manner. Despite these fields being unbounded, only a finite number of values are actually encountered within the dataset, and thus the inventors quantize according to the values encountered. For example, in the Byte field across 33 million flows in the dataset, only approximately 180,000 unique values are encountered in the dataset. Therefore, the inventors convert this ordinal field into a categorical field with a cardinality of 180,000. Subsequently, when generating data, the model outputs the index over the values rather than directly outputting these values. The structure of the post-processed netflow can be seen in Table 9. [Table 9]
[0069] The inventors trained three model sizes on the CIDDS-001 dataset and calculated metrics for the resulting samples. The model topologies are outlined in Table 10. Since ML efficiency and DCR are very costly to calculate for a dataset of this scale, the inventors instead adapt Precision and Recall to the tabular domain. Precision represents the proportion of data that is likely to be generated from the reference distribution by the inventors. Recall, on the contrary, is the percentage of data that is expected to be generated by the model by the inventors.
Table 10
[0070] The original definitions of these metrics rely on a visual model for generating the embeddings used. The inventors discovered that the TabMT masking procedure generates strong embeddings for each sample and chose to use the embeddings generated by TabMT-S, the minimal model by the inventors, and concatenate them to create one embedding per flow. Specifically, the inventors averaged the embeddings across fields to generate embeddings of dim64 per flow. The inventors used a fixed neighborhood size of k = 3. The inventors included additional diversity metrics since the average set converges across all properties of the generated data. Diversity is calculated for each field as the ratio of the number of unique values generated by the inventors' model to the proportion of unique values present in the reference set of data for that field. To calculate the reference values and the values for the inventors' model, the inventors split their validation set in half, treated one half as the reference set, and the other half as the baseline set to achieve the ceiling for the inventors' metrics. The performance results in Table 11 demonstrate strong scaling and performance across three model sizes. It can be seen that both the precision and recall are very close to those of the validation set. Both sample diversity and quality scale as the model size increases.
Table 11
[0071] The present inventors compared these with the prior art level NetflowGAN (NFGAN) described in Markus Ring et al., Flow-based network traffic generation using generative adversarial networks, Computers & Security, 82:156-172, 2019 (hereinafter, "Ring"), which is incorporated herein by reference. NFGAN was specifically tuned for the CIDDS-001 dataset. As described in Ring, it is trained in two phases. First, IP2Vec is trained to generate Netflow embeddings. These embeddings are then used as targets for the generator during GAN training. The results from NFGAN are shown in Table 11. It can be seen that although NFGAN obtains a fairly high precision rate, the recall rate and diversity are poor. This is because the model suffers from mode collapse where it generates samples from only a small proportion of the entire distribution.
[0072] Figures 5a and 5b show a comparison between the fake (Figure 5a) and real (Figure 5b) manifolds of the data. The present inventors used PaCMAP (Pairwise Controlled Manifold Approximation) to embed and visualize the data, embedding half of the validation set and an equal-sized sample of the data generated by the largest model by the present inventors for this visualization. It can be seen that the manifolds are very similar, indicating a good capture of the data distribution. The manifolds are projected from the embeddings generated by the smallest model by the present inventors.
[0073] Figures 6a and 6b are histograms over connections represented in the data. Specifically, the inventors plot the frequency of network flows between pairs of hosts on the network by the inventors. Again here, a similar distribution of connections can be seen in the fake data (Figure 6a) and the real data (Figure 6b).
[0074] Netflow has both correlations between fields and complex invariants between fields. The inventors can measure the violation rate of these invariants to understand how well the models by the inventors detect patterns in the data. The inventors measured against the 7 invariants proposed in Ring. Since the inventors construct an embedding for each field, the models by the inventors cannot violate the check 5 (*). These tests check the structural rules reflected in Netflow, such as that two public IP addresses cannot talk to each other.
[0075] As shown in Table 12, TabMT generates substantially more diverse data while achieving a 20x improvement in the median violation probability over NFGAN.
Table 12
[0076] Therefore, a masked transformer architecture, namely TabMT, which is a novel architecture for generating high-quality synthetic tabular data, is described herein. As discussed above with respect to a comprehensive set of benchmarks, TabMT has been shown to generate higher quality data than state-of-the-art tabular generators including GANs, VAEs, autoregressive transformers, and diffusion models. Furthermore, TabMT is able to do this while functioning under a broader set of conditions such as missing data, and is able to function under any privacy budget; it has the ability to arbitrarily trade off privacy and quality through temperature scaling. TabMT is also scalable from very small tabular databases to very large tabular databases. The masking procedure of TabMT enables it to effectively handle missing data, thus enhancing privacy and model applicability in real-world use cases.
[0077] Due to these superior features, TabMT has broad applicability across all industries that rely on tabular data, including but not limited to the financial field for fraud detection and economic prediction, the medical and insurance fields for client and event research, the social media field for user behavior, the streaming service field for recommendation behavior and settings, and the marketing and advertising campaign field for customer behavior and response. Tabular data is one of the most common and important data modalities. A vast amount of data, such as clinical trial records, financial data, census results, etc., are all represented in tabular format. The ability to use synthetic datasets, i.e., datasets where confidential attributes and Personally Identifiable Information (PII) are not disclosed, is crucial for maintaining compliance with privacy regulations while still providing data for analysis, sharing, and experimentation. As an example, given the concerns regarding patient privacy in the medical industry, TabMT can generate synthetic data, such as synthetic patient medical record data, synthetic genomic datasets, for use in research projects. Additionally, synthetic tabular data can provide data diversity and generate data for rare cases where it is difficult to obtain data from real data while still representing realistic probabilities.
Claims
1. A transformer-based modeling architecture for training a generator to generate synthetic data, At least one embedding model for embedding input data, where the input data includes a plurality of fields F storing actual data values i,j and the at least one embedding model constructs l embedding matrices, one for each of the plurality of fields F i,j respectively; Row F of the data in the embedded input data set i Set of unmasked fields for each 【Number 28】 and a set of masked fields 【Number 29】 is generated, and the training distribution 【30 Numbers】 A masking model for replacing the actual data values of each field in i the masked fields in 【Number 31】 of the masked set 【Number 32】 is uniform; and a transformer model 【Number 33】 for predicting the real data value for each mask token in each of the masked fields over a plurality of iterations.
2. The masking model masks a random number of fields F i,j The transformer-based modeling architecture according to claim 1, which masks
3. The masking model masks the field F i,j with a probability p~U(0, 1) and samples p for each row F i The transformer-based modeling architecture according to claim 2.
4. A first embedding model for embedding categorical data in the input data; and A second embedding model for embedding numerical data in the input data, where the second embedding model includes a quantizer and generates an ordered embedding, The transformer-based modeling architecture according to claim 1, further comprising.
5. The transformer-based modeling architecture according to claim 4, further comprising a dynamic linear layer for receiving output data from the transformer model.
6. The transformer-based modeling architecture according to claim 5, wherein the dynamic linear layer is paired with the ordered embedding.
7. The transformer-based modeling architecture according to claim 1, wherein the input data is tabular data.
8. The transformer-based modeling architecture according to claim 1, wherein the transformer model is a neural network.
9. A process for training a transformer-based generator model to generate synthetic data, At least one embedding model embeds input data, where the input data includes a plurality of fields F storing actual data values i,j and the at least one embedding model constructs l embedding matrices, one for each of the plurality of fields F i,j ; By the masking model, the rows F of data in the embedded input dataset i each with a set of masked fields 【Number 34】 is generated, and the training distribution 【Number 35】 replacing the actual data value of each field in with a mask token, where each row F of the data i the masked fields in 【Number 36】 of the masked set 【No. 37】 is uniform; and Predicting the real data value for each mask token in each of the masked fields 【Number 38】 over a plurality of iterations by a transformer model A process comprising.
10. The masking model masks a random number of fields F i,j The process according to claim 9, which masks
11. The masking model masks field F i,j with probability p ~ U(0, 1), and samples p for each row F i The process according to claim 10
12. Embedding categorical data in the input data by a first embedding model; and Embedding numerical data in the input data by a second embedding model, where the step of embedding numerical data includes quantizing the numerical data and generating an ordered embedding, The process according to claim 9, further comprising.
13. The process according to claim 12, further comprising the step of outputting data from the transformer model to a dynamic linear layer.
14. The process according to claim 13, further comprising the step of pairing the dynamic linear layer with the ordered embedding.
15. The process according to claim 9, wherein the input data is tabular data.
16. The process according to claim 9, wherein the transformer model is a neural network.
17. A transformer-based model for generating synthetic data, At least one embedding model for embedding an input data set, where the input data set has a predetermined format and includes a plurality of fields F indicating actual data values i,j and the at least one embedding model constructs l embedding matrices, one for each of the plurality of fields F i,j ; Receiving a first embedded input data set, wherein each real-world value for each of the plurality of fields F i,j is masked with a mask token, and iterating over the masked fields for a plurality of iterations until all mask tokens in the first embedded input data set contain synthetic data values 【Number 39】 a transformer model for generating synthetic data values equal to the real data values for the mask tokens in one or more of comprising a transformer-based model.
18. A first embedding model for embedding categorical data in the input data set; and A second embedding model for embedding numerical data in the input data set, wherein the second embedding model includes a quantizer and generates an ordered embedding. The transformer-based model according to claim 17, further comprising.
19. The transformer-based model according to claim 18, further comprising a dynamic linear layer for receiving output data from the transformer model.
20. The transformer-based model according to claim 19, wherein the dynamic linear layer is paired with the ordered embedding.
21. The transformer-based model according to claim 17, wherein the input data set stores tabular data.
22. The transformer-based model according to claim 17, wherein the transformer model is a neural network.
Citation Information
Patent Citations
Machine learning-assisted polypeptide analysis
JP2022521686A
Classification of erroneous cell data
US20220051126A1