Structured data generation method and device based on language model and feature correlation

Through a structured data generation method based on the correlation between language models and features, the problem of limited computing performance in the existing technology when processing complex data sets is solved, and accurate and effective synthetic structured data is generated, which improves the generalization ability and privacy protection ability of the model.

CN120105091APending Publication Date: 2025-06-06FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510071410.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has limited computing performance when processing large dimensions and complex data sets, making it difficult to effectively generate and analyze data sets containing a large number of variables or complex relationships.

Method used

A structured data generation method based on the correlation between language models and features is adopted, and through correlation calculation, filtering and text generation, a trained language model is constructed to generate synthetic structured data.

Benefits of technology

It realizes the generation of more accurate and effective synthetic structured data, solves the problems of data scarcity and privacy protection, and improves the generalization ability of the model and its reliability in real scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105091A_ABST
    Figure CN120105091A_ABST
Patent Text Reader

Abstract

The invention provides a structured data generation method and device based on language model and feature correlation, and the method is characterized in that the method comprises the following steps: a preprocessing step: generating a preprocessing value; a correlation calculation step of obtaining a correlation value between every two of all the features; a correlation screening step: selecting pairwise features from all pairwise features as correlation feature pairs according to a preset threshold value and a correlation value; a related text generation step: for each related feature pair, generating a text describing a relationship between two corresponding features as a relationship description text; an initial text generation step: generating an initial text corresponding to each feature of the sample; a training text generation step: combining the initial text with the relation description text to obtain a corresponding training sample; a model training step: training a language model through the training sample; and a synthetic data generation step: generating synthetic structured data through the language model. In a word, the method can generate more accurate and effective synthetic structured data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a structured data generation method, and in particular to a structured data generation method and device based on language model and feature correlation. Background Art

[0002] Structured data plays an important role in the medical, engineering, and financial industries, and is of great significance in promoting business development, data governance, and security management. However, with the improvement of privacy awareness and the increasing prominence of data security issues, data involving sensitive information is usually regarded as confidential content and is difficult to be directly used for data analysis by various computer algorithms. In addition, due to the small size of some sub-industries, the number of samples available for analysis is also relatively limited.

[0003] In the face of these challenges, researchers have proposed an innovative solution: using data generation methods to expand data sets to meet the needs of artificial intelligence model training. This method can not only help solve the problem of data scarcity, but also improve the generalization ability of the model, making it more reliable and effective in real-world scenarios. This innovative idea has opened up new ways to handle sensitive data, protect privacy, and promote the development of artificial intelligence applications.

[0004] There are two main traditional methods for generating structured data based on statistical models: link function and Bayesian network. The link function connects the joint distribution function with the respective marginal distribution functions to describe the correlation between variables. Through the link function, the correlation of real data can be used to generate new data. Another method is to combine Bayesian networks with differential privacy, which can generate data while ensuring privacy security. However, although these methods have solved the problem of data generation to a certain extent, the computational performance limitations of statistical models have led to their difficulties in dealing with large-dimensional and complex data sets. This means that for data sets containing a large number of variables or complex relationships, traditional statistical models may not be able to effectively generate and analyze data. This has also prompted researchers to continuously explore new methods and technologies to better cope with the increasingly complex and huge data challenges in the real world.

[0005] In the era of deep learning, Vadim Borisov et al. first applied language models to structured data generation tasks and proposed a real table data generation model GReaT. The GReaT method converts the generated text data into structured data through regular expressions, but the use of regular expressions has its limitations, which may lead to generation failure and redundant feature generation due to matching errors of some data. To solve the privacy protection problem, Solatorio et al. proposed a method of using masks to prevent the language model from directly copying data from the original data, but did not improve the encoding method of structured data. Zhao et al. proposed a more token-saving encoding method based on the encoding method of GReaT. He et al. used the language model to generate structured data in the proposed semi-supervised learning method, but did not provide a systematic structured data generation solution based on the language model. The above-mentioned structured data generation method based on the language model does not provide a method for associating data features, so the effect of the model still has room for further improvement. Summary of the invention

[0006] The present invention is made to solve the above-mentioned problems, and its purpose is to provide a method and device for generating structured data based on language model and feature correlation.

[0007] The present invention provides a structured data generation method based on language model and feature correlation, which is used to generate corresponding synthetic structured data according to a structured data set, wherein the structured data set includes n samples, each sample includes m features, and has such features, and comprises the following steps: a preprocessing step, preprocessing the values ​​corresponding to each feature of each sample to obtain the corresponding preprocessing values; a correlation calculation step, according to all preprocessing values ​​corresponding to each feature in the structured data set, obtaining the correlation values ​​between all the features by a correlation calculation method; a correlation screening step, according to a preset threshold and the correlation value, selecting the two-by-two features from all the two-by-two features as the relevant features for each relevant feature pair, generating a text describing the relationship between the two corresponding features as a relationship description text; an initial text generating step, for each sample, converting the feature name and value of each corresponding feature into the initial text corresponding to the feature of the sample through preset conjunctions; a training text generating step, for each sample, combining all the corresponding initial texts and all the relationship description texts to obtain the corresponding training sample; a model training step, constructing a language model, and training the language model according to all the training samples to obtain a trained language model; a synthetic data generating step, generating synthetic structured data through the trained language model.

[0008] The structured data generation method based on language model and feature correlation provided by the present invention may also have the following characteristics: wherein, in the correlation calculation step, the correlation calculation method is one of the Pearson correlation coefficient calculation method, the Spearman correlation coefficient calculation method, and the Kendall correlation coefficient calculation method.

[0009] In the structured data generation method based on language model and feature correlation provided by the present invention, it can also have the following characteristics: when the correlation calculation method is the Pearson correlation coefficient calculation method, the correlation value between the kth feature and the lth feature The calculation expression is: Where E[] is the function for calculating mathematical expectation, Y k is the kth feature, Y k The mean value of Y l is the lth feature, Y l The mean of Y k The standard deviation of Y l The standard deviation of .

[0010] The structured data generation method based on language model and feature correlation provided by the present invention may also have the following feature: in the correlation screening step, any two features whose absolute values ​​of correlation values ​​are greater than a preset threshold are regarded as relevant feature pairs.

[0011] The structured data generation method based on language model and feature correlation provided by the present invention may also have the following feature: wherein, in the related text generation step, the relationship description text is a text of feature names connecting two features through a preset relationship conjunction.

[0012] The structured data generation method based on language model and feature correlation provided by the present invention may also have the following feature: wherein, in the training sample, the relationship description text corresponding to the relevant feature pair is set between the initial texts corresponding to the two features corresponding to the relevant feature pair.

[0013] The structured data generation method based on language model and feature correlation provided by the present invention may also have the following characteristics: wherein, in the preprocessing step, the text type value is one-hot encoded to obtain the corresponding one-hot encoding as the preprocessing value, and for the numerical type value, the value is used as the corresponding preprocessing value.

[0014] The present invention also provides a structured data generation device based on language model and feature correlation, which is used to generate corresponding synthetic structured data according to a structured data set, wherein the structured data set contains m samples, each sample contains m features, and has such characteristics, including: a preprocessing module, which is used to preprocess the values ​​corresponding to each feature of each sample to obtain the corresponding preprocessing values; a correlation calculation module, which is used to obtain the correlation values ​​between all the features according to all the preprocessing values ​​corresponding to each feature in the structured data set through a correlation calculation method; a correlation screening module, which is used to select two-by-two features from all the two-by-two features as relevant feature pairs according to a preset threshold and correlation value; and a correlation filtering module. The text generation module is used to generate, for each relevant feature pair, a text describing the relationship between the corresponding two features as a relationship description text; the initial text generation module is used to convert, for each sample, the feature name and value of each feature into the initial text corresponding to the feature of the sample through preset conjunctions; the training text generation module is used to combine all the corresponding initial texts and all the relationship description texts for each sample to obtain the corresponding training sample; the model training module is used to build a language model and train the language model according to all the training samples to obtain a trained language model; the synthetic data generation module is used to generate synthetic structured data through the trained language model.

[0015] Functions and Effects of the Invention

[0016] According to the method and device for generating structured data based on language model and feature correlation involved in the present invention, the correlation values ​​between various features are calculated by the correlation calculation method, and then the relevant feature pairs are screened according to the correlation values ​​and the corresponding relationship description texts are generated, and finally the language model is trained using the relationship description texts. Therefore, the method and device for generating structured data based on language model and feature correlation of the present invention can generate more accurate and effective synthetic structured data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic diagram of a structured data set in an embodiment of the present invention;

[0018] Figure 2 is a block diagram of a structured data generating device in an embodiment of the present invention;

[0019] Figure 3 is a schematic diagram of a correlation matrix formed by correlation values ​​between each pair of features in an embodiment of the present invention;

[0020] Figure 4 It is a flowchart of a method for generating structured data based on language model and feature correlation in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and the accompanying drawings specifically illustrate the method and device for generating structured data based on language model and feature correlation of the present invention.

[0022] This embodiment provides a structured data generation device based on a language model and feature correlation, hereinafter referred to as a structured data generation device, which is used to generate corresponding synthetic structured data according to a structured data set.

[0023] Figure 1 is a schematic diagram of a structured data set in an embodiment of the present invention.

[0024] like Figure 1 As shown, the structured data set contains n=3 samples, each sample contains m=5 features, and the feature names corresponding to the 5 features are "Age", "FrequentFlyer", "IncomeClass", "ServicesOpted" and "Target".

[0025] In this embodiment, the synthetic structured data has a feature distribution similar to or identical to that of the structured data set, and conforms to the relationship between data features. The synthetic structured data can be used as training data for downstream tasks, thereby solving problems such as too few original data samples in the structured data set, or the original data containing sensitive information and being unavailable.

[0026] Figure 2 It is a block diagram of a structured data generating device in an embodiment of the present invention.

[0027] like Figure 2 As shown, the structured data generating device 100 includes a preprocessing module 11, a correlation calculation module 12, a correlation screening module 13, a related text generating module 14, an initial text generating module 15, a training text generating module 16, a model training module 17, a synthetic data generating module 18 and a general control module 19 for controlling the above modules.

[0028] The preprocessing module 11 is used to preprocess the values ​​corresponding to each feature of each sample to obtain corresponding preprocessing values.

[0029] Among them, the preprocessing module 11 performs one-hot encoding on the text type value to obtain the corresponding one-hot encoding as the corresponding preprocessing value, and for the value of the numerical type, uses the value as the corresponding preprocessing value, that is, for the value of the numerical type, retains the original representation as the preprocessing value.

[0030] The correlation calculation module 12 is used to obtain the correlation values ​​between all the features in pairs through a correlation calculation method according to all the pre-processed values ​​corresponding to the features in the structured data set.

[0031] The correlation calculation method is one of the Pearson correlation coefficient calculation method, the Spearman correlation coefficient calculation method, and the Kendall correlation coefficient calculation method.

[0032] In this embodiment, the correlation calculation method is the Pearson correlation coefficient calculation method, so the correlation value between the kth feature and the lth feature is The calculation expression is:

[0033]

[0034] Where E[] is the function for calculating mathematical expectation, Y k is the kth feature, Y k The mean value of Y l is the lth feature, Y l The mean of Y k The standard deviation of Y l The standard deviation of .

[0035] According to the definition of Pearson correlation coefficient, we know that The closer the value of is to 1, the k With Y l The stronger the positive correlation; The closer the value is to -1, the k With Y l The stronger the negative correlation; The closer the value of is to 0, the k With Y l The linear relationship is weaker.

[0036] Figure 3 It is a schematic diagram of a correlation matrix formed by correlation values ​​between each pair of features in an embodiment of the present invention.

[0037] like Figure 3 As shown, in the correlation matrix, the correlation values ​​corresponding to the same two features are set to 1, and the Pearson correlation coefficient is used to calculate the correlation values ​​between different features, and it is found that "FrequentFlyer" and "IncomeClass", and "Target" and "IncomeClass" have a strong linear relationship.

[0038] The correlation screening module 13 is used to select two features as correlation feature pairs from all two features according to a preset threshold and a correlation value.

[0039] The correlation screening module 13 takes the features whose absolute values ​​of correlation values ​​are greater than a preset threshold as correlation feature pairs, and the expression thereof is:

[0040]

[0041] In the formula is the correlation value The absolute value of , threshold is the preset threshold.

[0042] The related text generation module 14 is used to generate, for each related feature pair, a text describing the relationship between the corresponding two features as a relationship description text.

[0043] The relationship description text generated by the related text generation module 14 is a text that connects the feature names of two features through a preset relationship conjunction. The expression for the relationship description text corresponding to the kth feature and the lth feature generated by the related text generation module 14 is:

[0044] R k,l =[F k ,C kl ,F l ],

[0045] Where F k is the feature name of the kth feature, F l is the feature name of the lth feature, C kl is the relational conjunction between the kth feature and the lth feature, R k,l It is the relationship description text corresponding to the k-th feature and the l-th feature.

[0046] In this embodiment, the relational conjunction is "depends on". When the relevant feature pair includes the feature names "IncomeClass" and "FrequentFlyer", the corresponding relational description text is "IncomeClass depends onFrequentFlyer". In other embodiments, different relational conjunctions may be selected according to different correlation matrix calculation methods, or when two features do not have a linear correlation, specific descriptive words may be used as corresponding relational conjunctions to describe the relationship between the two features.

[0047] The initial text generation module 15 is used to convert the feature name and value of each feature corresponding to each sample into the initial text corresponding to the feature of the sample through preset connecting words.

[0048] Initial text generation module 15 Initial text generation of the i-th sample D i The expression of the j-th feature of is:

[0049] T j (D i )=[F j ,V j ,D i,j ],

[0050] Where F j is the feature name of the jth feature, D i,j is the value of the jth feature, V j It is the conjunction corresponding to the feature name and value of the j-th feature.

[0051] In this embodiment, the initial text generation module 15 uses the conjunction "is". The initial text generation module 15 performs initial text conversion for the feature name "Age" and the value "34", and the obtained initial text is "Age is 34". In other embodiments, different conjunctions can be preset for different features according to actual needs.

[0052] The training text generation module 16 is used to combine all the corresponding initial texts and all the relationship description texts for each sample to obtain the corresponding training sample.

[0053] Among them, in the training samples, the relationship description text corresponding to the relevant feature pair is set between the initial texts corresponding to the two features corresponding to the relevant feature pair.

[0054] In this embodiment, when "FrequentFlyer" and "IncomeClass", "Target" and "IncomeClass" have corresponding relationship description texts, the training sample corresponding to the first sample of the structured data set is "Age is 34, FrequentFlyer is No, IncomeClass depends on FrequentFlyer, IncomeClass is MiddleIncome, ServicesOpted is 6, Target depends on IncomeClass, Target is 0".

[0055] In other embodiments, the features before and after the relational conjunction have a mutual dependence relationship, that is, When , according to the dependency relationship, the initial text corresponding to the dependent feature is set before the corresponding relationship description text.

[0056] The model training module 17 is used to construct a language model and train the language model according to all training samples to obtain a trained language model.

[0057] Among them, the generation process of the language model can be expressed as:

[0058] T z =q(ω 1 ,…,ω s-1 ),

[0059] Where ω 1 ,…,ω s-1 is the token sequence of the input text, then the probability distribution of the language model inferring the output token ω at the sth position is:

[0060]

[0061] Where T is the sampling temperature, which is a parameter used to control the creativity level of AI-generated text. The higher the sampling temperature, the more freedom the model has in selecting output tokens, and the more diverse the generated results will be.

[0062] The synthetic data generation module 18 is used to generate synthetic structured data through the trained language model. In this embodiment, the synthetic data generation module 18 decodes the data output by the trained language model to obtain synthetic structured data.

[0063] The master control module 19 stores a control program for controlling the operation of each module.

[0064] The following describes the process of using the structured data generating device 100 to perform a structured data generating method based on a language model and feature correlation in conjunction with the accompanying drawings.

[0065] Figure 4 It is a flowchart of a method for generating structured data based on language model and feature correlation in an embodiment of the present invention.

[0066] like Figure 4 As shown, the structured data generation method based on the language model and feature correlation includes the following steps:

[0067] Step S1 is a preprocessing step, in which the preprocessing module 11 is used to preprocess the values ​​corresponding to each feature of each sample to obtain corresponding preprocessing values.

[0068] Step S2 is a correlation calculation step, in which the correlation calculation module 12 obtains the correlation values ​​between all the features according to all the pre-processed values ​​corresponding to the features in the structured data set through a correlation calculation method.

[0069] Step S3 is a correlation screening step, in which the correlation screening module 13 selects two features as correlation feature pairs from all two features according to a preset threshold and a correlation value.

[0070] Step S4 is a related text generation step, in which the related text generation module 14 generates text describing the relationship between the two corresponding features for each related feature pair as a relationship description text.

[0071] Step S5 is the initial text generation step, in which the initial text generation module 15 converts the feature name and value of each feature corresponding to each sample into the initial text corresponding to the feature of the sample through preset conjunctions.

[0072] Step S6 is a training text generation step, in which the training text generation module 16 combines all corresponding initial texts and all relationship description texts for each sample to obtain a corresponding training sample.

[0073] Step S7 is a model training step, in which a language model is constructed using the model training module 17, and the language model is trained according to all training samples to obtain a trained language model.

[0074] Step S8 is a synthetic data generation step, in which the synthetic data generation module 18 is used to generate synthetic structured data using a trained language model.

[0075] In this embodiment, steps S2-S4 are performed first and then step S5. In other embodiments, step S5 may be performed first and then steps S2-S4, or steps S2-S4 and step S5 may be performed simultaneously as needed.

[0076] In this embodiment, the existing GReaT method is compared with the Top3 method and Avg method based on the structured data generation method based on the language model and feature correlation on the existing "Tour&Travels Customer ChurnPrediction" dataset, and LR linear regression, DT decision tree and RF random forest are used to evaluate the synthetic structured data generated by each method, and compared with the original data Original used to construct the structured dataset.

[0077] Among them, the Top3 method is to select the three pairs of features with the highest correlation values ​​in the structured data generation method based on the language model and feature correlation to generate relationship description text, and the Avg method is to generate relationship description text corresponding to all features with higher correlation values ​​than the average in the structured data generation method based on the language model and feature correlation. The distilgpt2 model is selected as the language model in each method.

[0078] Then, each method generates 5,000 synthetic data as the accuracy of synthetic structured data in different machine learning model classifiers, as shown in the following table:

[0079] Original GReaT method Top 3 Methods Avg method LR Linear Regression 0.822 0.753 0.759 0.732 DT Decision Tree 0.901 0.842 0.905 0.916 RF Random Forest 0.905 0.832 0.901 0.916

[0080] In the above table, the first column is the classifiers of different machine learning models, and the second to fifth columns are the original data Original, as well as the accuracy of synthetic structured data generated by the GReaT method, the Top3 method, and the Avg method on different machine learning model classifiers. It can be seen that using specific relationship description text can generate more accurate and effective synthetic structured data.

[0081] Functions and Effects of the Embodiments

[0082] According to the structured data generation method and device based on language model and feature correlation involved in this embodiment, the correlation value between each feature is calculated by the correlation calculation method, and then the relevant feature pairs are screened according to the correlation value and the corresponding relationship description text is generated, and finally the language model is trained using the relationship description text. In short, this method can generate more accurate and effective synthetic structured data.

[0083] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A structured data generation method based on language model and feature correlation, used to generate corresponding synthetic structured data according to a structured data set, wherein the structured data set includes n samples, each of which includes m features, and is characterized in that: The following steps are involved: A preprocessing step of preprocessing the values ​​corresponding to each of the features of each of the samples to obtain corresponding preprocessed values; A correlation calculation step, according to all the pre-processed values ​​corresponding to each of the features in the structured data set, obtaining correlation values ​​between all the features by a correlation calculation method; A correlation screening step, selecting two features from all two features as relevant feature pairs according to a preset threshold and the correlation value; A related text generating step, for each related feature pair, generating a text describing the relationship between the two corresponding features as a relationship description text; The initial text generation step is to convert the feature name and value of each feature of each sample into the initial text corresponding to the feature of the sample through the preset connecting words; A training text generation step, for each of the samples, combining all the corresponding initial texts and all the relationship description texts to obtain a corresponding training sample; A model training step, constructing a language model, and training the language model according to all the training samples to obtain a trained language model; The synthetic data generating step generates the synthetic structured data through the trained language model.

2. The method for generating structured data based on language model and feature correlation according to claim 1, characterized in that: in, In the correlation calculation step, the correlation calculation method is one of a Pearson correlation coefficient calculation method, a Spearman correlation coefficient calculation method, and a Kendall correlation coefficient calculation method.

3. The method for generating structured data based on language model and feature correlation according to claim 2, characterized in that: in, When the correlation calculation method is the Pearson correlation coefficient calculation method, the correlation value between the kth feature and the lth feature is The calculation expression is: Where E[] is the function for calculating mathematical expectation, Y k is the kth feature, Y k The mean value of Y l is the lth feature, Y l The mean of Y k The standard deviation of Y l The standard deviation of .

4. The method for generating structured data based on language model and feature correlation according to claim 3, characterized in that: in, In the correlation screening step, the pairwise features whose absolute values ​​of the correlation values ​​are greater than the preset threshold are taken as the correlation feature pairs.

5. The method for generating structured data based on language model and feature correlation according to claim 1, characterized in that: in, In the related text generation step, the relationship description text is a text that connects the feature names of two features through a preset relationship conjunction.

6. The method for generating structured data based on language model and feature correlation according to claim 1, characterized in that: in, In the training sample, the relationship description text corresponding to the relevant feature pair is set between the initial texts corresponding to the two features corresponding to the relevant feature pair.

7. The method for generating structured data based on language model and feature correlation according to claim 1, characterized in that: in, In the preprocessing step, the value of the text type is one-hot encoded to obtain the corresponding one-hot encoding as the preprocessing value, For the value of the numerical type, the value is used as the corresponding preprocessing value.

8. A structured data generation device based on language model and feature correlation, used to generate corresponding synthetic structured data according to a structured data set, wherein the structured data set includes n samples, each of which includes m features, and is characterized in that: include: A preprocessing module, used for preprocessing the values ​​corresponding to each of the features of each of the samples to obtain corresponding preprocessed values; A correlation calculation module, used for obtaining correlation values ​​between all the features in the structured data set by a correlation calculation method according to all the pre-processed values ​​corresponding to each of the features; A correlation screening module, used for selecting two features as relevant feature pairs from all two features according to a preset threshold and the correlation value; A related text generation module, used for generating, for each related feature pair, a text describing the relationship between the two corresponding features as a relationship description text; An initial text generation module, for converting the feature name and value of each feature of each sample into the initial text corresponding to the feature of the sample through a preset conjunction; A training text generation module, for combining all the corresponding initial texts and all the relationship description texts for each of the samples to obtain a corresponding training sample; A model training module is used to construct a language model and train the language model according to all the training samples to obtain a trained language model; The synthetic data generation module is used to generate the synthetic structured data through the trained language model.