Table data generation method and device based on language model adversarial data enhancement
Through the adversarial data enhancement method based on language model, the problem of insufficient data quality in the prior art is solved, and more real and high-quality data is generated, which improves the performance and data security of the machine learning model.
Patent Information
- Application Number
- CN202510023574.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-16
AI Technical Summary
The synthetic table data generated by existing data augmentation technologies are of limited quality and it is difficult to meet the needs of high-quality data.
Adversarial data augmentation method based on language model is adopted, and multiple iterative training is carried out by building discriminators and generators, optimizing the parameters of generators and discriminators to generate more realistic and high-quality synthetic table data.
It improves the authenticity and quality of synthetic tabular data, enables machine learning models to be better learned and generalized, and solves data security and privacy issues.
Smart Images

Figure CN120012733A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for generating tabular data, and in particular to a method and device for generating tabular data based on language model adversarial data enhancement. Background Art
[0002] Data is an indispensable and important part of the development of artificial intelligence, and its quality directly affects the performance ceiling of machine learning models. Tabular data is a common and important data type in various fields, especially in healthcare and finance. However, obtaining high-quality data is a complex and arduous task, especially when it comes to large-scale data sets. With the increasing public concern about data privacy issues and the continuous improvement of government regulations, relevant regulations require companies to protect consumer personal information and emphasize that consumers have the right not to disclose personal information. In addition, various relevant laws and regulations of various countries also regulate the collection, storage and processing of personal data to varying degrees to protect personal privacy and data security.
[0003] In order to solve data security and privacy issues, data enhancement technology has gradually become an effective method. Data enhancement technology can generate synthetic data sets with similar feature distributions but without real personal information, thereby avoiding the leakage of sensitive information in machine learning tasks. Data enhancement technology is especially important when processing tabular data. It can help improve the diversity and quality of data, allowing machine learning models to learn and generalize better. Therefore, data enhancement technology is of great significance in the current development of artificial intelligence, providing effective protection for data security and model performance.
[0004] However, the synthetic data generated by existing data augmentation techniques still suffers from limited quality. Summary of the invention
[0005] The present invention is made to solve the above-mentioned problem, and its purpose is to provide a method and device for generating tabular data based on language model adversarial data enhancement.
[0006] The present invention provides a table data generation method based on language model adversarial data enhancement, which is used to generate synthetic table data according to existing table data, and has such characteristics, comprising the following steps: step S1, constructing a discriminator and a generator, and training the discriminator and the generator according to an original training data set constructed from a plurality of training table data, and obtaining a trained generator as a table data generation model; step S2, preprocessing the existing table data to obtain preprocessed data; step S3, inputting the preprocessed data into the table data generation model to obtain synthetic table data, wherein step S1 includes the following sub-steps: step S1-1, constructing Discriminator and generator, and use the original training data set as the training data set; step S1-2, generate training synthetic data corresponding to the training data set through the generator, and optimize the generator according to the training synthetic data and the training data set; step S1-3, determine whether the preset termination condition is met, if so, obtain the table data generation model, if not, enter step S1-4; step S1-4, filter the enhanced data from all the training synthetic data through the discriminator, and optimize the discriminator according to the training synthetic data and the original training data set; step S1-5, merge the original training data set and the enhanced data as the training data set, and execute step S1-2.
[0007] The method for generating tabular data based on language model adversarial data enhancement provided by the present invention may also have the following feature: wherein, in step S1-4, the authenticity prediction probability corresponding to each training synthetic data is generated by the discriminator, and the training synthetic data whose authenticity prediction probability is greater than a preset threshold is used as enhanced data.
[0008] The method for generating tabular data based on language model adversarial data enhancement provided by the present invention may also have the following features: wherein, in step S1-2, the parameters of the tabular data generation model are updated by maximizing the likelihood function in combination with the optimization algorithm, and the calculation expression of the likelihood function is: Where θ is the parameter of the table data generation model before optimization, and N is the token sequence length.
[0009] In the method for generating tabular data based on language model adversarial data enhancement provided by the present invention, it may also have the following characteristics: wherein, in step S1-4, the parameters of the discriminator are optimized by a binary cross entropy loss function, and the calculation expression of the binary cross entropy loss function is: Where N is the total number of tabular data in the training synthetic data and the original training data set, y(i) is the true label value of the i-th tabular data, is the predicted label value of the discriminator for the i-th tabular data. The true label value of each tabular data in the training synthetic data is set to 1, and the true label value of each tabular data in the original training data set is set to 0.
[0010] In the method for generating tabular data based on language model adversarial data enhancement provided by the present invention, it may also have the following features: wherein, preprocessing includes jsonizing the values of each feature in the existing tabular data, and sorting and sampling all the features in random order to obtain preprocessed data; in step S3, the trained tabular data generation model outputs json text according to the preprocessed data, and decodes the json text to obtain synthetic tabular data.
[0011] The method for generating tabular data based on language model adversarial data enhancement provided by the present invention may also have the following feature: wherein the generator is an auto-regression model.
[0012] The present invention also provides a tabular data generation device based on language model adversarial data enhancement, which is used to generate synthetic tabular data based on existing tabular data and has the following characteristics, including: a preprocessing module, which is used to preprocess the existing tabular data to obtain preprocessed data; a generation module, which includes a tabular data generation model, which is used to input the preprocessed data into the tabular data generation model to obtain synthetic tabular data, wherein the construction and training of the tabular data generation model include the following steps: step S1-1, constructing a discriminator and a generator, and using the original training data set as the training data set; step S1-2, generating training synthetic data corresponding to the training data set through the generator, and optimizing the generator according to the training synthetic data and the training data set; step S1-3, judging whether the preset termination condition is met, if so, obtaining the tabular data generation model, if not, entering step S1-4; step S1-4, filtering all the training synthetic data through the discriminator to obtain enhanced data, and optimizing the discriminator according to the training synthetic data and the original training data set; step S1-5, merging the original training data set and the enhanced data as the training data set, and executing step S1-2.
[0013] Functions and Effects of the Invention
[0014] According to the method and device for generating tabular data based on language model adversarial data enhancement involved in the present invention, on the one hand, the generation effect of the subsequent tabular data generation model is improved through the preprocessing of jsonization and random sorting; on the other hand, the generator and the discriminator are iterated through the adversarial network, and in each iteration, the enhanced data is screened from the training synthetic data generated by the generator through the discriminator as part of the training data for the next iteration, thereby optimizing the performance of the final tabular data generation model. Therefore, the method and device for generating tabular data based on language model adversarial data enhancement of the present invention can generate more realistic and high-quality synthetic tabular data. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a block diagram of a table data generating device in an embodiment of the present invention;
[0016] Figure 2 is a schematic diagram of the process of constructing and training a table data generation model in an embodiment of the present invention;
[0017] Figure 3 It is a flowchart of a method for generating tabular data based on language model adversarial data enhancement in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and the accompanying drawings specifically illustrate the method and device for generating table data based on language model countermeasures against data enhancement of the present invention.
[0019] The present embodiment provides a table data generation device based on language model adversarial data enhancement, hereinafter referred to as a table data generation device, which is used to generate synthetic table data according to existing table data.
[0020] Figure 1 It is a block diagram of a table data generating device in an embodiment of the present invention.
[0021] like Figure 1 As shown, the table data generating device 100 includes a preprocessing module 11 and a generating module 12 .
[0022] The preprocessing module 11 is used to preprocess the existing table data to obtain preprocessed data.
[0023] The preprocessing includes converting the values of each feature in the existing table data into json, and sorting and sampling all the features in a random order to obtain preprocessed data.
[0024] In this embodiment, the existing table data contains multiple features, each of which includes a corresponding feature name and feature value. i The transformed text obtained by jsonizing the jth feature in is expressed as:
[0025] t j (D i )=[′f j ′:D i,j ],
[0026] Where f j is the feature name corresponding to the jth feature, D i,j is the eigenvalue corresponding to the jth feature.
[0027] Then, the existing table data D i The expression of the non-random JSON prompt words obtained through jsonification is:
[0028] Json(D i )=[{t i (D i ),…,t j (D i )}],j∈{1,…,m},
[0029] Where m is the existing table data D i The total number of features in .
[0030] Since json data is an unordered key-value pair, the converted texts after jsonization can be arranged in any order, and the disorder is conducive to the subsequent data synthesis reasoning of any feature conditions. Then, the expression of preprocessing data obtained by converting the existing table data C into json and sorting and sampling all features in disorder is:
[0031]
[0032] Where k l is the lth feature after the disorder, and l is less than the total number of features of the existing table data C, that is, in this embodiment, only part of the features are extracted as preprocessing data through sampling, is the transformed text corresponding to the lth feature after shuffling, and ' is the left quotation mark in the JSON text.
[0033] The generating module 12 includes a table data generating model, which is used to input the pre-processed data into the table data generating model to obtain the synthesized table data. The generator is an auto-regression model.
[0034] Among them, the trained table data generation model outputs JSON text according to the preprocessed data, and the generation module 12 then decodes the JSON text to obtain the synthesized table data.
[0035] Figure 2 It is a flowchart of the construction and training of the table data generation model in an embodiment of the present invention.
[0036] like Figure 2 As shown in the figure, the construction and training of the tabular data generation model includes the following steps:
[0037] Step S1-1, constructing a discriminator and a generator, and using the original training data set constructed from a plurality of training table data as the training data set. In this embodiment, the training table data is converted into json and sorted in random order in the above preprocessing to construct the original training data set.
[0038] Step S1-2, generating training synthetic data corresponding to the training data set through a generator, and optimizing the generator according to the training synthetic data and the training data set.
[0039] In this embodiment, the tabular data generation model based on the autoregressive model can predict the next tag according to the generated text sequence, that is, the expression of the tabular data generation model generating text through conditional probability is:
[0040]
[0041] Where x1, x2, …, x n is a token sequence of json text that is input to the tabular data generation model.
[0042] Then in step S1-2, the parameters of the table data generation model are updated by maximizing the likelihood function in combination with the optimization algorithm. The calculation expression of the likelihood function is:
[0043]
[0044] Where θ is the parameter of the table data generation model before optimization, and N is the token sequence length.
[0045] The optimization algorithm used in this embodiment is the gradient descent algorithm.
[0046] Step S1-3, determine whether the preset termination condition is met, if yes, obtain the table data generation model, if no, proceed to step S1-4. In this embodiment, the preset termination condition is whether the iteration round is greater than the maximum iteration round, and in other embodiments, it can be determined whether to terminate the iteration based on the model performance.
[0047] Step S1-4, the enhanced data is obtained by filtering all the training synthetic data through the discriminator, and the discriminator is optimized according to the training synthetic data and the original training data set.
[0048] Among them, the authenticity prediction probability corresponding to each training synthetic data is generated by the discriminator, and then the training synthetic data with the authenticity prediction probability greater than the preset threshold is used as the enhanced data.
[0049] Among them, the parameters of the discriminator are optimized by the binary cross entropy loss function, and the calculation expression of the binary cross entropy loss function is:
[0050]
[0051] Where N is the total number of tabular data in the training synthetic data and the original training data set, y(i) is the true label value of the i-th tabular data, is the predicted label value of the discriminator for the i-th table data.
[0052] In this embodiment, the true label value of each table data in the training synthetic data is set to 1, and the true label value of each table data in the original training data set is set to 0. The expression is:
[0053]
[0054] Where D origin is the original training data set, D synthesis For training synthetic data, x (i) For table data i, y )i( is the true label value corresponding to the table data i.
[0055] Step S1-5, merging the original training data set and the enhanced data as the training data set, and executing step S1-2.
[0056] In step S1-5 of this embodiment, the expression for updating the training data set is:
[0057] D enhance =D high ∪D real ,
[0058] Where D enhance is the updated training dataset, D high To enhance the data, D real is the original training data set.
[0059] The following describes the process of using the table data generation device 100 to perform a table data generation method based on language model adversarial data enhancement in conjunction with the accompanying drawings.
[0060] Figure 3 It is a flowchart of a method for generating tabular data based on language model adversarial data enhancement in an embodiment of the present invention.
[0061] like Figure 3 As shown, the method for generating tabular data based on language model adversarial data enhancement includes the following steps:
[0062] Step S1, constructing a discriminator and a generator, and training the discriminator and the generator according to an original training data set constructed from a plurality of training tabular data, to obtain a trained generator as a tabular data generation model.
[0063] Step S2, using the preprocessing module 11 to preprocess the existing table data to obtain preprocessed data.
[0064] Step S3, using the generation module 12 to input the pre-processed data into the table data generation model to obtain the synthesized table data.
[0065] In this embodiment, the existing Loan dataset and Insurance dataset are used as training table data to construct the original training dataset, and the table data generation method that collaboratively enhances the small model and the language model, namely Ours method, is compared with the method that uses the preprocessing method of this embodiment but does not use the discriminator to filter the enhanced data for training, namely JSON method.
[0066] The autoregression models are all distilgpt2 models, and the discriminator is a random forest model. In the Ours method, 5 rounds of training are performed on the Loan dataset, and 20 rounds of training are performed on the Insurance dataset. In the JSON method, 25 rounds of training are performed on the Loan dataset, and 100 rounds of training are performed on the Insurance dataset.
[0067] With the above settings, the two methods are trained and tested on two data sets, and the ACC, AUC, and MAPE indicators of the two methods are shown in the following table:
[0068]
[0069] In the above table, the first column is each data set, the second column is each indicator, the third column is the indicator value obtained by verifying the original data in each data set, and the fourth and fifth columns are the indicator values corresponding to the synthetic tabular data generated by the model trained by the JSON method and the Ours method, respectively. It can be seen that the tabular data generation method based on language model adversarial data enhancement can further improve the authenticity of synthetic data and improve the efficiency of model training compared to the simple JSON method.
[0070] Functions and Effects of the Embodiments
[0071] According to the table data generation method and device based on language model adversarial data enhancement involved in this embodiment, on the one hand, the generation effect of the subsequent table data generation model is improved through jsonization and random sorting preprocessing; on the other hand, the generator and discriminator are iterated through the adversarial network, and in each iteration, the enhanced data is selected from the training synthetic data generated by the generator through the discriminator as part of the training data for the next iteration, thereby optimizing the performance of the final table data generation model. In short, this method can generate more realistic and high-quality synthetic table data.
[0072] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A method for generating tabular data based on language model adversarial data enhancement, used to generate synthetic tabular data based on existing tabular data, characterized in that: The following steps are involved: Step S1, constructing a discriminator and a generator, and training the discriminator and the generator according to an original training data set constructed from a plurality of training table data, to obtain a trained generator as a table data generation model; Step S2, preprocessing the existing table data to obtain preprocessed data; Step S3, inputting the pre-processed data into the table data generation model to obtain the synthesized table data, Wherein, the step S1 includes the following sub-steps: Step S1-1, constructing the discriminator and the generator, and using the original training data set as the training data set; Step S1-2, generating training synthetic data corresponding to the training data set by the generator, and optimizing the generator according to the training synthetic data and the training data set; Step S1-3, determining whether a preset termination condition is met, if so, obtaining the table data generation model, if not, proceeding to step S1-4; Step S1-4, filtering all the training synthetic data to obtain enhanced data through the discriminator, and optimizing the discriminator according to the training synthetic data and the original training data set; Step S1-5: merge the original training data set and the enhanced data as the training data set, and execute step S1-2.
2. The method for generating tabular data based on language model adversarial data enhancement according to claim 1, characterized in that: in, In step S1-4, the authenticity prediction probability corresponding to each of the training synthetic data is generated by the discriminator. The training synthetic data whose authenticity prediction probability is greater than a preset threshold is used as the enhanced data.
3. The method for generating tabular data based on language model adversarial data enhancement according to claim 1, characterized in that: in, In step S1-2, the parameters of the table data generation model are updated by maximizing the likelihood function in combination with an optimization algorithm. The calculation expression of the likelihood function is: Where θ is the parameter of the table data generation model before optimization, and N is the token sequence length.
4. The method for generating tabular data based on language model adversarial data enhancement according to claim 1, characterized in that: in, In step S1-4, the parameters of the discriminator are optimized by a binary cross entropy loss function, and the calculation expression of the binary cross entropy loss function is: Where N is the total number of tabular data in the training synthetic data and the original training data set, y(i) is the true label value of the i-th tabular data, is the predicted label value of the discriminator for the i-th table data, The true label value of each piece of the table data in the training synthetic data is set to 1, The true label value of each piece of the table data in the original training data set is set to 0.
5. The method for generating tabular data based on language model adversarial data enhancement according to claim 1, characterized in that: in, The preprocessing includes converting the values of each feature in the existing table data into json, and sorting and sampling all the features in a random order to obtain the preprocessed data. In step S3, the trained table data generation model outputs json text according to the preprocessed data. Decode the JSON text to obtain the synthetic table data.
6. The method for generating tabular data based on language model adversarial data enhancement according to claim 1, characterized in that: in, The generator is an autoregressive model.
7. A tabular data generation device based on language model adversarial data enhancement, used to generate synthetic tabular data based on existing tabular data, characterized in that: include: A preprocessing module, used for preprocessing the existing table data to obtain preprocessed data; A generation module, comprising a table data generation model, for inputting the pre-processed data into the table data generation model to obtain the synthesized table data, The construction and training of the tabular data generation model includes the following steps: Step S1-1, constructing the discriminator and the generator, and using the original training data set as the training data set; Step S1-2, generating training synthetic data corresponding to the training data set by the generator, and optimizing the generator according to the training synthetic data and the training data set; Step S1-3, determining whether a preset termination condition is met, if so, obtaining the table data generation model, if not, proceeding to step S1-4; Step S1-4, filtering all the training synthetic data to obtain enhanced data through the discriminator, and optimizing the discriminator according to the training synthetic data and the original training data set; Step S1-5: merge the original training data set and the enhanced data as the training data set, and execute step S1-2.