Joint multi-task table semantic parsing method based on pre-training model
Through the joint multi-task table semantic analysis method based on pre-trained models, the problem that users with non-technical backgrounds find it difficult to extract information from table data is solved, and high accuracy and complex query generation capabilities are achieved, and complex table structures and query requirements are adapted to complex table structures and query requirements.
Patent Information
- Application Number
- CN202510085101.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to achieve the efficient extraction of valuable information from table data by users without technology background, mainly due to the complexity of SQL statements and the limitations of traditional text-to-SQL parsing methods.
Using a joint multitask table semantic analysis method based on pre-trained models, SQL statements and tables are converted into natural language text through a large language model, and MLNaT model with 12-layer relationship-aware Transformer architecture is constructed to perform joint learning of mask language model, column prediction and SQL generation.
It improves the accuracy of table semantic analysis and the generation ability of complex queries, and can handle multi-table joining and aggregation operations more effectively, adapting to complex scenarios in practical applications.
Smart Images

Figure CN120011390A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing and database technology, and more specifically, to a joint multi-task table semantic parsing method based on a pre-training model. Background Art
[0002] In today's era of accelerating digitalization, tabular data has become one of the key carriers of information storage and transmission. For business operations, various data tables in financial statements accurately record detailed information on core indicators such as revenue, cost, and profit, and are an important basis for management to make strategic decisions; in the field of scientific research, experimental data tables orderly present observation results under different variable conditions, providing indispensable materials for researchers to explore laws and verify hypotheses. However, for the vast number of users with non-technical backgrounds, there are significant obstacles to extracting valuable information from these tables with the help of SQL statements.
[0003] SQL is a highly professional database query language with complex syntax and rigorous logic, requiring users to have solid programming knowledge and rich practical experience. For example, when constructing multi-table association queries, it is necessary not only to accurately specify the connection conditions, but also to reasonably use aggregate functions and filter clauses to obtain the desired results. This technical threshold often discourages many non-technical personnel from dealing with massive amounts of table data, making it difficult to fully explore the potential value behind the data, greatly limiting the effective use and sharing of data.
[0004] Traditional text-to-SQL parsing methods are unable to cope with this dilemma. Although the rule-based matching method can play a certain role in specific simple scenarios, it is difficult to achieve accurate and universal parsing in the face of complex and changeable natural language expressions and diverse table structures due to the limitations of its rule base. The template generation method also has many disadvantages. It is highly dependent on manually designed templates, which not only consumes a lot of manpower and time costs, but also lacks flexibility when dealing with new fields or special needs. It cannot meet the growing complexity and diversity requirements in practical applications. There is an urgent need for more efficient and intelligent solutions to break this deadlock. Summary of the invention
[0005] In view of the shortcomings of the prior art, the object of the present invention is to provide a joint multi-task table semantic parsing method based on a pre-trained model.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A joint multi-task table semantic parsing method based on a pre-trained model includes the following steps:
[0008] After crawling SQL statements from the specified website, the large language model is used to convert SQL statements and tables into natural language text, and columns and tables are extracted as positive and negative samples. At the same time, the acquired experimental data format is converted to be consistent with the Spider dataset. By creating a prompt word template and using a few-sample framework, the task of generating natural language questions and database schemas from SQL statements and tables is completed;
[0009] Build an MLNaT model with a 12-layer relation-aware Transformer architecture, concatenate the sentence and the column name in the pattern in a specific format and input them, set up three learning tasks: masked language model, column prediction, and SQL generation, and perform pre-training;
[0010] The model is experimentally evaluated on the Spider dataset, with the exact set matching rate as the evaluation indicator, and the RAT-SQL model as the baseline model.
[0011] Preferably, the use of a large language model to convert SQL statements and tables into natural language texts is specifically to use the large language model ERNIE4.0 or ERNIE3.5. For SQL-to-statement conversion, the original SQL statement and a preset prompt word template are input; for table-to-statement conversion, the column name, column value and corresponding relationship of the table are input.
[0012] Preferably, the extracting of columns and tables as positive samples and negative samples is to extract columns and tables from each SQL as positive samples, and extract irrelevant columns and tables from other SQLs as negative samples to form triplets.
[0013] Preferably, the acquired experimental data format is converted to be consistent with the Spider data set by calling the Wenshengwen interface of the large language model, creating a prompt word template with specific content, using a few-sample framework, and setting inference hyperparameters, wherein the inference hyperparameters include temperature=0.8 and top_p=0.1.
[0014] Preferably, the inputting of splicing the statement with the column name in the pattern in a specific format is to splice the statement L with the column name column in the pattern S in a format of X={<t>L<co1>c1<co1>c2<co1>< / t>}.
[0015] Preferably, the masked language model task is to use 15% of the tokens in the text <mask>Replacement, let the model predict the original token of the replaced position.
[0016] Preferably, the column prediction task is to perform an average pooling operation on all tokens corresponding to the column name to obtain a representation vector, input the representation vector into a two-layer MLP for binary classification to determine whether the column name is used, and use a binary cross entropy function as a loss function.
[0017] Preferably, in the SQL generation task, the decoder generates a target SQL token with the help of a data dictionary storing column names and SQL statement keywords, wherein SQL keyword embedding is randomly initialized and trained in a pre-training phase, and the column representation is obtained by averaging the sub-token representations of the column. In each decoding step, the decoder generates a hidden vector and performs a dot product operation to generate a probability distribution of the target vocabulary on the vocabulary data set.
[0018] Compared with the prior art, the present invention has the following beneficial effects:
[0019] 1. In the present invention, a joint multi-task learning framework is adopted. Through the synergy of the three learning tasks of masked language model, column prediction and final SQL generation, the model can more fully learn the semantic association between statements and table structures, improve the accuracy of table semantic parsing, and more effectively generate high-quality SQL query statements compared with existing single-task learning methods.
[0020] 2. In the present invention, the method of generating training data based on the large language model ERNIE4.0 / 3.5 is proposed to solve the problem of poor quality of some existing pre-training data, enrich the diversity of training data, lay the foundation for training a model with strong generalization ability, and help the model to accurately perform table semantic analysis in different database scenarios and diversified natural language queries.
[0021] 3. In the present invention, by selecting the BART initialization model and designing an encoder semantic parser based on a 12-layer relation-aware Transformer, combined with reasonable pre-training and fine-tuning strategies, the overall performance of the model is further improved, and its ability to handle complex table structures and complex query requirements is enhanced, showing good adaptability and robustness when facing complex situations such as multi-table connections and aggregation operations in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flowchart of a joint multi-task table semantic parsing method based on a pre-training model is proposed for the present invention;
[0023] Figure 2 This is a schematic diagram of the MLNaT model framework proposed in the present invention;
[0024] Figure 3 A schematic diagram of a single example of the Spider dataset proposed in the present invention. DETAILED DESCRIPTION
[0025] Reference Figures 1 to 3 .
[0026] The embodiment further illustrates a joint multi-task table semantic parsing method based on a pre-training model proposed by the present invention.
[0027] A joint multi-task table semantic parsing method based on a pre-trained model includes the following steps:
[0028] After crawling SQL statements from the specified website, the large language model is used to convert SQL statements and tables into natural language text, and columns and tables are extracted as positive and negative samples. At the same time, the acquired experimental data format is converted to be consistent with the Spider dataset. By creating a prompt word template and using a few-sample framework, the task of generating natural language questions and database schemas from SQL statements and tables is completed;
[0029] Build an MLNaT model with a 12-layer relation-aware Transformer architecture, concatenate the sentence and the column name in the pattern in a specific format and input them, set up three learning tasks: masked language model, column prediction, and SQL generation, and perform pre-training;
[0030] The model is experimentally evaluated on the Spider dataset, with the exact set matching rate as the evaluation indicator, and the RAT-SQL model as the baseline model.
[0031] The method of converting SQL statements and tables into natural language text using a large language model specifically uses the large language model ERNIE4.0 or ERNIE3.5. For SQL-to-statement conversion, the original SQL statement and a preset prompt word template are input; for table-to-statement conversion, the column name, column value and corresponding relationship of the table are input.
[0032] The Spider dataset is a multi-database, multi-table, single-round query text-to-SQL parsing dataset, including 10,181 natural language (NL) questions and 5,693 unique SQL query statements generated in 200 databases, 1,020 tables involving 138 different fields; its evaluation indicator is the exact set matching rate; the benchmark parser uses the RAT-SQL model as the baseline model; the RAT-SQL model is one of the most advanced parsers on the Spider dataset, which uses Relation-aware Self-Attention; at the same time, both explicit relations (Schema) and implicit relations (Linking between Question and Schema) are taken into account in the encoding, which improves the representation ability of the model.
[0033] The extracting of columns and tables as positive samples and negative samples is to extract columns and tables from each SQL as positive samples, and extract irrelevant columns and tables from other SQLs as negative samples to form triplets.
[0034] The obtained experimental data format is converted to be consistent with the Spider data set by calling the Wenshengwen interface of the large language model, creating a prompt word template with specific content, using a few-sample framework, and setting inference hyperparameters, wherein the inference hyperparameters include temperature=0.8 and top_p=0.1.
[0035] The pre-training stage is specifically:
[0036] In the processing flow of the MLNaT model, the operation will first be carried out for the given sentence L and pattern S; specifically, the model will concatenate the sentence L with the column name column in the pattern S, and the specific format of the concatenation is set to X={<t>L<co1>c1<co1>c2<co1>< / t>}; after the concatenation is completed, the data will be input into the architecture built by the 12-layer relation-aware Transformer; each token in the input can be encoded as a contextual representation, and then these representations are used by different decoders for different learning tasks.
[0037] The masked language model is a very important pre-training task technology in the field of natural language processing, and is often used in the pre-training stage of language models such as BERT. For text data, 15% of the tokens in the text are randomly replaced with a special symbol. <mask>to replace; and the task of the masked language model is to let the model predict these replaced <mask>What token should be in the position? The model needs to learn the contextual semantics of the entire sentence and output a reasonable result. <mask>The labeled text allows the model to perform predictive training, and the model can learn the deep semantic relationship, grammatical structure and other knowledge between words in the text.
[0038] The column predictions include:
[0039] First, for all tokens corresponding to a column name, we will obtain a representation vector through the average pooling operation; the calculation formula is: Among them, the sequence x of length n has a pooling output y for position i i ;
[0040] After obtaining the representation vector of the column name, it will be input into a two-layer multi-layer perceptron (MLP); MLP is a common neural network structure composed of multiple neurons, which transforms the input data layer by layer, extracts features, and other operations; the two-layer MLP is used here to make the representation vector obtained from the column name pass through the two layers of neurons for calculation, further mining and refining features, so as to adapt to the subsequent specific classification tasks; and this specific task is to perform a binary classification of whether it is used; that is, based on the vector processed by the two layers of MLP, it is judged whether the content corresponding to the column name belongs to the "used" category or the "unused" category in the specific application scenario or model requirement. The final output result is one of the two cases of the binary classification, so as to realize the classification judgment of the column name related availability. The loss function used in this binary classification task is the binary cross entropy function, and its calculation formula is: Among them, y i is a binary label 0 or 1, p(y i ) The output is the label y i probability.
[0041] The SQL statement output includes:
[0042] One way for the MLNaT model decoder to generate the target SQL token is to use a data dictionary that stores column names and SQL statement keywords. The data dictionary is a knowledge base from which the decoder selects appropriate elements to form the SQL token to be output. At the same time, the embeddings of SQL keywords are randomly initialized, and these initialized embeddings are trained in the pre-training stage, so that the model gradually learns the appropriate vector representation of each keyword, so that it can be accurately placed in the appropriate position when generating SQL statements.
[0043] When obtaining the column representation, the method used is to average the sub-token representations of the column; this is the same as the way to obtain the column representation in the column prediction learning objective; by averaging the sub-token representation vectors corresponding to the column, a vector that can represent the overall semantic features of the column is finally obtained, preparing for the accurate use of column-related information in subsequent operations such as generating SQL statements;
[0044] In each decoding step, the decoder first generates a hidden vector, which can be seen as a mathematical representation of the current decoding stage state. It contains the information that has been processed and potential guidance for the subsequent generation of content. The decoder uses this hidden vector to perform a dot product operation to generate the probability distribution of the target vocabulary in the vocabulary data set, that is, to determine the probability of each keyword, column name, etc. in the SQL statement at the current decoding position. Finally, the actual output is determined based on this probability distribution, and the generation of the entire SQL statement and other content is gradually completed.
[0045] The model is experimentally evaluated on the Spider dataset, specifically:
[0046] In the pre-training phase, we use the underlying transformer initialized with BART to train our MLNaT model; BART (Bidirectional and Auto-Regressive Transformers) is a powerful pre-trained language model with an encoder-decoder architecture that can perform well in a variety of natural language processing tasks; by using BART as initialization, we can take advantage of its powerful bidirectional context modeling capabilities and autoregressive decoding capabilities, thus providing a good starting point for the training of the MLNaT model; this pre-training phase is mainly trained through large-scale text data using self-supervised learning methods.
[0047] In the fine-tuning stage, we further utilize the pre-trained model structure, especially the encoder semantic parser in the MLNaT model; this encoder is designed based on the 12-layer relation-aware Transformer architecture; the Transformer model is one of the most advanced sequence modeling methods currently, and its self-attention mechanism can effectively capture the dependencies between elements in the sequence, while the relation-aware Transformer further enhances the model's ability to handle complex semantic relationships; in addition, the 12-layer relation-aware Transformer structure can maintain high accuracy and robustness when facing complex language tasks through multi-level processing and fine-grained semantic modeling.
[0048] For downstream tasks, experiments were conducted on a single dataset to verify the effectiveness of the framework. The comparative experiments are as follows:
[0049] In the Spider dataset experiment, we reproduced the RAT-SQL+BERT model and achieved an exact set match rate of 0.647 on the development set, which is close to the result of the RAT-SQL v2+BERT model, but not as good as its v3 version. After that, after replacing the BERT encoder with the BART encoder, the accuracy on the development set and test set increased to 0.659 and 0.628 respectively. Although the performance of the model based on the BART encoder on the hidden test set is close to that of the RAT-SQL v3+BERT model, the BART encoder only contains 12 layers of Transformer, which is more streamlined than the 24 layers of the BERT large model.
[0050] By introducing the MLNaT framework, the exact set matching rate of the RAT-SQL model on the development set and hidden test set is further improved to 0.665 and 0.631; this result shows that the MLNaT framework has significant advantages in improving the column prediction and SQL generation performance in the text-to-SQL parsing task, verifying its ability to generate high-quality parsing results; Table 1 shows the comparative experimental results for different learning objectives.
[0051] Table 1 Comparative experiments with different learning objectives
[0052]
[0053] The experiment studies the exact set matching rate of three different learning tasks, namely the Masked Language Model (MLM) task, the column prediction (CPred) task, and the final SQL generation (FinSQL) task. The ablation experiment is as follows:
[0054] Table 2 shows the experimental results of three baseline systems based on the IRNet model: IRNet+BERT, IRNet+TaBERT and IRNet+RoBERTa; The experimental results show that improving the quality of the semantic parser encoder is an effective direction to improve performance; In order to further verify the contribution of each learning objective, an ablation experiment was conducted, and Table 3 shows the results of the ablation experiment; The results show that without the FinSQL learning objective, the accuracy of the two learning tasks MLM and CPred is improved by 0.3% and 0.1% respectively compared with the baseline (IRNet+BART); This shows that these learning objectives have improved the encoding quality of the Transformer encoder; Based on the standard unsupervised learning objective MLM combined with the CPred learning task, the model accuracy is further improved to 0.699, which is 1.9% higher than the baseline;
[0055] Table 2 Ablation experiments based on different learning objectives of IRNet model
[0056]
[0057] Table 3 Ablation experiments of different learning objectives
[0058]
[0059] When the FinSQL learning objective is included, the three learning objectives are compared based on a baseline accuracy of 0.699, showing the value of the FinSQL learning task. Under this condition, we observe that the MLM learning objective improves the accuracy by 0.4%, while the CPred learning objective improves the accuracy by 0.2%. The combination of MLM and CPred learning tasks further optimizes the generation of complex queries, and the model accuracy is improved to 0.710, enabling the model to handle multi-table joins and aggregation operations more accurately, demonstrating its potential for application in practical semantic parsing tasks.
[0060] The table example is as follows:
[0061]
[0062] In summary, this paper explores two problems in the text-to-SQL semantic parsing task and proposes a multi-task joint learning framework for natural language sentence and table structure pre-training, which contains three different learning objectives. Experimental results show that the framework is effective on the Spider dataset.
[0063] In the future, we will focus on the following directions: First, expand the datasets of high-quality SQL and table structures to enrich the training data source of the model and improve the generalization ability of the MLNaT model in diverse tasks; second, study efficient processing methods for large-scale table structures to cope with more complex database scenarios; finally, in view of the data privacy issues involved in the fine-tuning and reasoning process of current large language models, we will also explore data privacy protection solutions in the future to ensure data security when using large language model interfaces.
[0064] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.< / mask> < / mask> < / mask> < / mask>
Claims
1. A joint multi-task table semantic parsing method based on a pre-trained model, characterized in that: The following steps are involved: After crawling SQL statements from the specified website, the large language model is used to convert SQL statements and tables into natural language text, and columns and tables are extracted as positive and negative samples. At the same time, the acquired experimental data format is converted to be consistent with the Spider dataset. By creating a prompt word template and using a few-sample framework, the task of generating natural language questions and database schemas from SQL statements and tables is completed; Build an MLNaT model with a 12-layer relation-aware Transformer architecture, concatenate the sentence and the column name in the pattern in a specific format and input them, set up three learning tasks: masked language model, column prediction, and SQL generation, and perform pre-training; The model is experimentally evaluated on the Spider dataset, with the exact set matching rate as the evaluation indicator, and the RAT-SQL model as the baseline model.
2. According to the pre-trained model-based joint multi-task table semantic parsing method of claim 1, characterized in that: The method of converting SQL statements and tables into natural language text using a large language model specifically uses the large language model ERNIE4.0 or ERNIE3.
5. For SQL-to-statement conversion, the original SQL statement and a preset prompt word template are input; for table-to-statement conversion, the column name, column value and corresponding relationship of the table are input.
3. The method for joint multi-task table semantic parsing based on a pre-training model according to claim 2, characterized in that: The extracting of columns and tables as positive samples and negative samples is to extract columns and tables from each SQL as positive samples, and extract irrelevant columns and tables from other SQLs as negative samples to form triplets.
4. The method for joint multi-task table semantic parsing based on a pre-trained model according to claim 3, characterized in that: The obtained experimental data format is converted to be consistent with the Spider data set by calling the Wenshengwen interface of the large language model, creating a prompt word template with specific content, using a few-sample framework, and setting inference hyperparameters, wherein the inference hyperparameters include temperature=0.8 and top_p=0.
1.
5. The method for joint multi-task table semantic parsing based on a pre-training model according to claim 4, characterized in that: The inputting of splicing the statement with the column name in the pattern in a specific format is to splice the statement L with the column name column in the pattern S in the format of X={<t>L<co1>c1<co1>c2<co1>< / t>}.
6. The method for joint multi-task table semantic parsing based on a pre-trained model according to claim 5, characterized in that: The masked language model task is to use 15% of the tokens in the text <mask> Replacement, let the model predict the original token of the replaced position.< / mask> 7. The method for joint multi-task table semantic parsing based on a pre-trained model according to claim 6, characterized in that: The column prediction task is to perform average pooling operation on all tokens corresponding to the column name to obtain a representation vector, input the representation vector into a two-layer MLP for binary classification to determine whether the column name is used, and use a binary cross entropy function as the loss function.
8. The method for joint multi-task table semantic parsing based on a pre-trained model according to claim 7, characterized in that: In the SQL generation task, the decoder generates the target SQL token with the help of a data dictionary storing column names and SQL statement keywords, wherein SQL keyword embedding is randomly initialized and trained in the pre-training stage, and the column representation is obtained by averaging the sub-token representations of the column. In each decoding step, the decoder generates a hidden vector and performs a dot product operation to generate the probability distribution of the target vocabulary on the vocabulary data set.
Citation Information
Cited By
Table-to-text generation method based on multi-agent collaboration
CN120724976A
A table-to-text generation method based on multi-agent collaboration
CN120724976B
Large language model interpretation generation method and system based on fine tuning and joint tasks
CN121257553A
Fine-tuning and joint task-based large language model explanation generation method and system
CN121257553B