Table data generation method and device for collaborative enhancement of small model and language model

Through the method of synergistic enhancement of small models and language models, the problem of insufficient prediction ability of existing language models when generating tabular data is solved, and the quality and authenticity of synthetic data are improved.

CN120011427APending Publication Date: 2025-05-16FUDAN UNIVERSITY

Patent Information

Application Number
CN202510023571.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing language models lack prediction capabilities when generating tabular data, resulting in the quality of synthetic data not reaching the ideal level.

Method used

The method of synergistic enhancement of small models and language models is adopted to build small models and tabular data generation models, and use small models to optimize the output of the generated model during the training process to improve the performance of tabular data generation models.

Benefits of technology

Through collaborative enhancement, the generation effect and performance of the tabular data generation model are improved, and the generated synthetic table data is more realistic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011427A_ABST
    Figure CN120011427A_ABST
Patent Text Reader

Abstract

The invention provides a table data generation method and device for collaborative enhancement of a small model and a language model, and the method is characterized in that the method comprises the steps: S2-1, constructing a table data generation model, and taking an original training data set as a training data set; step S2-2, generating training synthetic data corresponding to the training data set through the table data generation model, and optimizing the table data generation model according to the training synthetic data and the training data set; step S2-3, judging whether a preset termination condition is met or not, if so, obtaining a trained table data generation model, and if not, entering step S2-4; and S2-4, auxiliary training data corresponding to the training synthesis data is generated through the trained small model, the original training data set and the auxiliary training data are merged to serve as a training data set, and the step S2-2 is executed. In a word, the method can generate more real synthetic table data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for generating tabular data, and in particular to a method and device for generating tabular data by collaboratively enhancing a small model and a language model. Background Art

[0002] Tabular data plays a common and important role in many fields, especially in the fields of medicine, finance, and social sciences. Considering the necessity of tabular data enhancement and rebalancing and the public's growing concern about data privacy, the development of effective tabular data generators has gradually become a trend. However, existing generative models perform well in the fields of images and text, and are able to synthesize high-quality data and learn the underlying structure of complex data sets, but these models face a series of challenges when processing tabular data. This is mainly attributed to the heterogeneity of tabular data, that is, the diversity of data types, distributions, and relationships, which brings unique complexity to the training and application of models that are not available in other fields.

[0003] Initially, the tabular data generation method based on generative adversarial networks (GANs) was widely used. TableGAN applies deep convolutional adversarial networks to tabular data, and the features of a record are converted into a matrix, which is then processed by the convolution kernel of the convolutional neural network. CTGAN designs a conditional generator trained by sampling for mode normalization to overcome non-Gaussian and multi-modal distribution problems, thereby processing unbalanced discrete columns. This method is able to generate a large amount of synthetic data that has no privacy issues and conforms to the distribution of real data.

[0004] At the same time, tabular data generation methods based on auto-encoding variational Bayesian VAE are also very popular. TVAE adopts the framework of variational autoencoder, uses two neural networks to model conditional probability distribution and posterior probability distribution respectively, and trains them using evidence lower limit loss to generate realistic synthetic data, and the synthetic results are better than CTGAN method.

[0005] In addition, the tabular data method based on the denoising diffusion probability model DDPM has also attracted much attention. TabDDPM uses polynomials and Gaussian diffusion to model the one-hot encoding of categorical variables and normalized numerical features, and is trained by minimizing the mean square error of Gaussian diffusion and the KL divergence of each polynomial diffusion term, thus solving the challenges encountered in processing tabular data containing continuous and categorical variables.

[0006] With the widespread application of language models such as Transformer in the text field, many studies have begun to explore the applicability of language models in the field of tabular data generation. However, the prediction results generated by existing language models as tabular data generators are not as effective as some traditional small machine learning models in some aspects. Therefore, tabular data generators based on language models still have insufficient effective prediction capabilities, resulting in the problem that the quality of synthetic data cannot reach the ideal level. Summary of the invention

[0007] The present invention is made to solve the above-mentioned problems, and its purpose is to provide a method and device for generating tabular data by collaboratively enhancing a small model and a language model.

[0008] The present invention provides a method for generating tabular data by collaboratively enhancing a small model and a language model, which is used for generating synthetic tabular data according to existing tabular data, and has such characteristics, comprising the following steps: step S1, constructing a small model, and training the small model according to an original training data set constructed from a plurality of training tabular data to obtain a trained small model; step S2, constructing a tabular data generation model, and training the tabular data generation model according to the original training data set and the trained small model to obtain a trained tabular data generation model; step S3, preprocessing the existing tabular data to obtain preprocessed data; step S4, inputting the preprocessed data into the trained tabular data generation model to obtain a synthetic tabular data. Grid data, wherein step S2 includes the following sub-steps: step S2-1, constructing a tabular data generation model, and using the original training data set as the training data set; step S2-2, generating training synthetic data corresponding to the training data set through the tabular data generation model, and optimizing the tabular data generation model according to the training synthetic data and the training data set; step S2-3, judging whether a preset termination condition is met, if so, obtaining a trained tabular data generation model, if not, entering step S2-4; step S2-4, generating auxiliary training data corresponding to the training synthetic data through the trained small model, and merging the original training data set and the auxiliary training data as the training data set, and executing step S2-2.

[0009] In the tabular data generation method for collaboratively enhancing the small model and the language model provided by the present invention, it can also have the following characteristics: wherein, the training synthetic data includes known features and real features, and in step S2-4, the trained small model generates predicted features based on the known features, and replaces the real features with the predicted features, to obtain updated training synthetic data as auxiliary training data.

[0010] The method for generating tabular data by collaboratively enhancing the small model and the language model provided by the present invention may also have the following features: wherein, in step S2-2, the parameters of the tabular data generation model are updated by maximizing the likelihood function in combination with the optimization algorithm, and the calculation expression of the likelihood function is: Where θ is the parameter of the table data generation model before optimization, and N is the token sequence length.

[0011] In the tabular data generation method for collaboratively enhancing the small model and the language model provided by the present invention, the following features may also be provided: wherein the small model includes a decision tree and XGBoost, and in step S1, the parameters of the small model are optimized by training, and the expression thereof is: In the formula is the parameter of the optimized small model, θ sm is the parameter of the small model before optimization, f(X o θ sm ) is the small model before optimization based on the known features X in the input training table data o Generated prediction features, Y o is the known feature X in the training table data o The corresponding real features, L() is the loss function.

[0012] The method for generating tabular data with collaborative enhancement of the small model and the language model provided by the present invention may also have the following features: wherein, preprocessing includes jsonizing the values ​​of each feature in the existing tabular data, and sorting and sampling all the features in random order to obtain preprocessed data; in step S4, the trained tabular data generation model outputs json text according to the preprocessed data, and decodes the json text to obtain synthetic tabular data.

[0013] The method for generating tabular data by collaboratively enhancing the small model and the language model provided by the present invention may also have the following feature: wherein the tabular data generation model is an auto-regression model.

[0014] The present invention also provides a tabular data generation device with collaborative enhancement of a small model and a language model, which is used to generate synthetic tabular data based on existing tabular data, and has the following characteristics, including: a preprocessing module, which is used to preprocess the existing tabular data to obtain preprocessed data; a generation module, which includes a trained tabular data generation model, which is used to input the preprocessed data into the trained tabular data generation model to obtain synthetic tabular data, wherein the construction and training of the tabular data generation model includes the following steps: step S1, constructing a small model, and training the small model based on an original training data set constructed from a plurality of training tabular data to obtain a trained small model; step S2, constructing a tabular data generation model, and training the tabular data based on the original training data set and the trained small model The grid data generation model is trained to obtain a trained tabular data generation model, and step S2 includes the following sub-steps: step S2-1, constructing a tabular data generation model, and using the original training data set as the training data set; step S2-2, generating training synthetic data corresponding to the training data set through the tabular data generation model, and optimizing the tabular data generation model according to the training synthetic data and the training data set; step S2-3, judging whether the preset termination condition is met, if so, obtaining the trained tabular data generation model, if not, entering step S2-4; step S2-4, generating auxiliary training data corresponding to the training synthetic data through the trained small model, and using the original training data set and the auxiliary training data as the training data set, and executing step S2-2.

[0015] Functions and Effects of the Invention

[0016] According to the method and device for generating tabular data with the small model and language model synergistically enhanced involved in the present invention, on the one hand, the generation effect of the subsequent tabular data generation model is improved through the preprocessing of jsonization and random sorting; on the other hand, when training the tabular data generation model, the output of the tabular data generation model is optimized by using the small model, and the optimized output is used as part of the training data for the next iterative training, thereby improving the performance of the tabular data generation model. Therefore, the method and device for generating tabular data with the small model and language model synergistically enhanced involved in the present invention can generate more realistic synthetic tabular data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a block diagram of a table data generating device in an embodiment of the present invention;

[0018] Figure 2 is a schematic diagram of the process of constructing and training a table data generation model in an embodiment of the present invention;

[0019] Figure 3 It is a flowchart of a method for generating tabular data by collaboratively enhancing a small model and a language model in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, specifically illustrate the method and device for generating table data by collaboratively enhancing the small model and the language model of the present invention.

[0021] The present embodiment provides a table data generation device that is collaboratively enhanced with a small model and a language model, hereinafter referred to as a table data generation device, which is used to generate synthetic table data based on existing table data.

[0022] Figure 1 It is a block diagram of a table data generating device in an embodiment of the present invention.

[0023] like Figure 1 As shown, the table data generating device 100 includes a preprocessing module 11 and a generating module 12 .

[0024] The preprocessing module 11 is used to preprocess the existing table data to obtain preprocessed data.

[0025] The preprocessing includes converting the values ​​of each feature in the existing table data into json, and sorting and sampling all the features in a random order to obtain preprocessed data.

[0026] In this embodiment, the existing table data contains multiple features, each of which includes a corresponding feature name and feature value. i The transformed text obtained by jsonizing the jth feature in is expressed as:

[0027] t j (D i )=[′f j ′:D i,j ],

[0028] Where f j is the feature name corresponding to the jth feature, D i,j is the eigenvalue corresponding to the jth feature.

[0029] Then, the existing table data D i The expression of the non-random JSON prompt words obtained through jsonification is:

[0030] Json(D i )=[{t i (D i ),…,t j (D i )}],j∈{1,…,m},

[0031] Where m is the existing table data di The total number of features in .

[0032] Since json data is an unordered key-value pair, the converted texts after jsonization can be arranged in any order, and the disorder is conducive to the subsequent data synthesis reasoning of any feature conditions. Then, the expression of preprocessing data obtained by converting the existing table data C into json and sorting and sampling all features in disorder is:

[0033]

[0034] Where k l is the lth feature after the disorder, and l is less than the total number of features of the existing table data C, that is, in this embodiment, only part of the features are extracted as preprocessing data through sampling, is the transformed text corresponding to the lth feature after shuffling, and ' is the left quotation mark in the JSON text.

[0035] The generation module 12 includes a trained table data generation model, which is used to input the pre-processed data into the trained table data generation model to obtain synthesized table data. The table data generation model is an auto-regression model.

[0036] Among them, the trained table data generation model outputs JSON text according to the preprocessed data, and the generation module 12 then decodes the JSON text to obtain the synthesized table data.

[0037] Figure 2 It is a flowchart of the construction and training of the table data generation model in an embodiment of the present invention.

[0038] like Figure 2 As shown, the construction and training of the tabular data generation model includes the following steps:

[0039] Step S1, construct a small model, and train the small model according to the original training data set constructed by multiple training table data to obtain a trained small model. In this embodiment, the training table data is jsonized and sorted in random order in the above preprocessing to construct the original training data set.

[0040] Among them, the small model includes a decision tree and XGBoost, that is, a decision tree or XGBoost is selected as the small model in this embodiment.

[0041] In step S1, the parameters of the small model are optimized through training, and the expression is:

[0042]

[0043] In the formula is the parameter of the optimized small model, θ smis the parameter of the small model before optimization, f(X o θ sm ) is the small model before optimization based on the known features X in the input training table data o Generated prediction features, Y o is the known feature X in the training table data o The corresponding real features, L() is the loss function. In this embodiment, the loss function is the cross entropy function. In other embodiments, the existing loss function can be set as needed or a specific function can be constructed as the loss function. In this embodiment, all features in the training table data are divided into known features and real features, so that the small model is trained to predict the corresponding real features using the known features. For example, a training table data is {'age':42,'sex':'female','bmi':32.87,'charge':5344.8}, and this embodiment focuses on improving the prediction quality of the feature 'charge', then 'age', 'sex' and 'bmi' are all used as known features, and 'charge' is used as a real feature.

[0044] Step S2, constructing a tabular data generation model, and training the tabular data generation model according to the original training data set and the trained small model to obtain a trained tabular data generation model.

[0045] Wherein, step S2 includes the following sub-steps:

[0046] Step S2-1, constructing a tabular data generation model, and using the original training data set as the training data set.

[0047] Step S2-2, generating training synthetic data corresponding to the training data set through the tabular data generation model, and optimizing the tabular data generation model according to the training synthetic data and the training data set, wherein the training synthetic data includes known features and real features.

[0048] In this embodiment, the tabular data generation model based on the autoregressive model can predict the next tag according to the generated text sequence, that is, the expression of the tabular data generation model generating text through conditional probability is:

[0049]

[0050] Where x1, x2, …, x n is a token sequence of json text that is input to the tabular data generation model.

[0051] Then in step S2-2, the parameters of the table data generation model are updated by maximizing the likelihood function in combination with the optimization algorithm. The calculation expression of the likelihood function is:

[0052]

[0053] Where θ is the parameter of the table data generation model before optimization, and N is the token sequence length.

[0054] The optimization algorithm used in this embodiment is the gradient descent algorithm.

[0055] Step S2-3, determine whether the preset termination condition is met, if so, obtain the trained table data generation model, if not, proceed to step S2-4. In this embodiment, the preset termination condition is whether the iteration round is greater than the maximum iteration round, and in other embodiments, it can be determined whether to terminate the iteration based on the model performance.

[0056] Step S2-4, generate auxiliary training data corresponding to the training synthetic data through the trained small model, and use the original training data set and the auxiliary training data as the training data set to execute step S2-2.

[0057] Among them, in step S2-4, the trained small model generates predicted features based on known features, and replaces the real features with the predicted features to obtain updated training synthetic data as auxiliary training data.

[0058] Then, the calculation expressions of auxiliary training data and updated training data set are:

[0059]

[0060] D new =D s ∪D,

[0061] Where X s To train known features in synthetic data, The prediction features generated for the trained small model, D s is the auxiliary training data, D is the original training data set, and D new is the updated training dataset.

[0062] The following describes the process of a table data generation method for collaboratively enhancing a small model and a language model using the table data generation device 100 in conjunction with the accompanying drawings.

[0063] Figure 3 It is a flowchart of a method for generating tabular data by collaboratively enhancing a small model and a language model in an embodiment of the present invention.

[0064] like Figure 3 As shown, the method for generating tabular data by collaboratively enhancing the small model and the language model includes the following steps:

[0065] Step S3, using the preprocessing module 11 to preprocess the existing table data to obtain preprocessed data.

[0066] Step S4, using the generation module 12 to input the pre-processed data into the trained table data generation model to obtain synthetic table data.

[0067] In this embodiment, the existing Loan data set is used as training table data to construct the original training data set, and the table data generation method that synergistically enhances the small model and the language model, namely the Ours method, is compared with the method that uses the preprocessing method of this embodiment but does not use the small model for training, namely the JSON method.

[0068] Among them, the autoregression model is the distilgpt2 model. In the Ours method, 5 rounds of training are performed, each round generates 1,000 data for data merging, the number of sub-rounds in each round is 5, and the small model is the random forest model. In the JSON method, the training is performed for 25 rounds.

[0069] With the above settings, the two methods are trained and tested on the Loan dataset, and the ACC and AUC indicators are shown in the following table:

[0070] index Origin JSON Ours ACC 0.992 0.915 0.917 AUC 0.998 0.643 0.761

[0071] In the above table, the first column is the various indicators, the second column is the indicator value obtained by verifying the original data in the Loan dataset, and the third and fourth columns are the indicator values ​​corresponding to the synthetic table data generated by the models trained by the JSON method and the Ours method, respectively. It can be seen that the table data generation method enhanced by the small model and the language model can further improve the authenticity of the synthetic data compared to the simple JSON method.

[0072] Functions and Effects of the Embodiments

[0073] According to the method and device for generating tabular data with collaborative enhancement of the small model and the language model involved in this embodiment, on the one hand, the generation effect of the subsequent tabular data generation model is improved through the preprocessing of jsonization and random sorting; on the other hand, the output of the tabular data generation model is optimized by using the small model when training the tabular data generation model, and the optimized output is used as part of the training data in the next iterative training, thereby improving the performance of the tabular data generation model. In short, this method can generate more realistic synthetic tabular data.

[0074] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A method for generating tabular data by collaboratively enhancing a small model and a language model, for generating synthetic tabular data based on existing tabular data, characterized in that: The following steps are involved: Step S1, constructing a small model, and training the small model according to an original training data set constructed from a plurality of training table data to obtain a trained small model; Step S2, constructing a tabular data generation model, and training the tabular data generation model according to the original training data set and the trained small model to obtain a trained tabular data generation model; Step S3, preprocessing the existing table data to obtain preprocessed data; Step S4, inputting the pre-processed data into the trained table data generation model to obtain the synthesized table data, Wherein, the step S2 includes the following sub-steps: Step S2-1, constructing the table data generation model, and using the original training data set as the training data set; Step S2-2, generating training synthetic data corresponding to the training data set through the tabular data generation model, and optimizing the tabular data generation model according to the training synthetic data and the training data set; Step S2-3, determining whether a preset termination condition is met, if so, obtaining the trained table data generation model, if not, proceeding to step S2-4; Step S2-4, generating auxiliary training data corresponding to the training synthetic data through the trained small model, and merging the original training data set and the auxiliary training data as the training data set, and executing step S2-2.

2. The method for generating tabular data by collaboratively enhancing a small model and a language model according to claim 1, characterized in that: in, The training synthetic data includes known features and real features, In step S2-4, the trained small model generates predicted features based on the known features, and replaces the real features with the predicted features to obtain the updated training synthetic data as the auxiliary training data.

3. The method for generating tabular data by collaboratively enhancing a small model and a language model according to claim 1, characterized in that: in, In step S2-2, the parameters of the table data generation model are updated by maximizing the likelihood function in combination with an optimization algorithm. The calculation expression of the likelihood function is: Where θ is the parameter of the table data generation model before optimization, and N is the token sequence length.

4. The method for generating tabular data by collaboratively enhancing a small model and a language model according to claim 1, characterized in that: in, The small model includes decision tree and XGBoost, In step S1, the parameters of the small model are optimized through training, and the expression is: In the formula is the parameter of the optimized small model, θ sm is the parameter of the small model before optimization, f(X o θ sm ) is the small model before optimization based on the known features X in the input training table data o Generated prediction features, Y o is the known feature X in the training table data o The corresponding real features, L() is the loss function.

5. The method for generating tabular data by collaboratively enhancing a small model and a language model according to claim 1, characterized in that: in, The preprocessing includes converting the values ​​of each feature in the existing table data into json, and sorting and sampling all the features in a random order to obtain the preprocessed data. In step S4, the trained table data generation model outputs json text according to the preprocessed data. Decode the JSON text to obtain the synthetic table data.

6. The method for generating tabular data by collaboratively enhancing a small model and a language model according to claim 1, characterized in that: in, The tabular data generation model is an autoregressive model.

7. A tabular data generation device with collaborative enhancement of a small model and a language model, used to generate synthetic tabular data based on existing tabular data, characterized in that: include: A preprocessing module, used for preprocessing the existing table data to obtain preprocessed data; A generation module, comprising a trained table data generation model, is used to input the pre-processed data into the trained table data generation model to obtain the synthesized table data, The construction and training of the tabular data generation model includes the following steps: Step S1, constructing a small model, and training the small model according to an original training data set constructed from a plurality of training table data to obtain a trained small model; Step S2, constructing a tabular data generation model, and training the tabular data generation model according to the original training data set and the trained small model to obtain a trained tabular data generation model, The step S2 comprises the following sub-steps: Step S2-1, constructing the table data generation model, and using the original training data set as the training data set; Step S2-2, generating training synthetic data corresponding to the training data set through the tabular data generation model, and optimizing the tabular data generation model according to the training synthetic data and the training data set; Step S2-3, determining whether a preset termination condition is met, if so, obtaining the trained table data generation model, if not, proceeding to step S2-4; Step S2-4, generating auxiliary training data corresponding to the training synthetic data through the trained small model, and using the original training data set and the auxiliary training data as the training data set, and executing step S2-2.

Citation Information

Patent Citations

  • Cross-table multi-task pre-training method and device based on language model

    CN117272149A

  • User repayment behavior prediction method and system based on table type data, and storage medium

    CN117994028A

  • Oil tea tree inherent frequency data enhancement method based on improved CTGAN network

    CN118395859A

  • Relational table-oriented data generation method and device

    CN118734817A

  • Systems and methods for generating synthetic tabular data for machine learning and other applications

    US20240330682A1

Cited By

  • Classification task-oriented data generation method based on big and small model collaboration

    CN120822037A

  • Data generation method based on size model cooperation for classification task

    CN120822037B

  • Table data synthesis method, system and equipment and storage medium

    CN121118858A

  • A table data synthesis method, system, device and storage medium

    CN121118858B