Text2sql method based on large model ensemble learning

By obtaining the optimal hyperparameter group and fine-tuning the Schema-linking and SQL generation models using the large-scale model Lora+Moe algorithm, the problem of insufficient SQL generation accuracy in the prior art is solved, and higher accuracy is achieved.

CN120494090APending Publication Date: 2025-08-15HAIKOU PORT COMM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510539217.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing large-model-based text2sql tasks are difficult to find a balance between ensuring high recall and high accuracy, resulting in insufficient SQL generation accuracy.

Method used

By obtaining the optimal hyperparameter group, fine-tuning the Schema-linking model and training the SQL generation model, the big model Lora+Moe algorithm is used for model training and supplementation, and the generation of table names and column names is optimized by combining Beam-Search and n-gram algorithms.

Benefits of technology

The accuracy of SQL generation is improved, and the accuracy of SQL generation models is improved by optimizing hyperparameters and fine-tuning of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494090A_ABST
    Figure CN120494090A_ABST
Patent Text Reader

Abstract

According to the text2sql method based on large model ensemble learning, after hyper-parameters K, RT, TN, RC, CN and MoeN are obtained, the hyper-parameters are divided into two hyper-parameter groups, the hyper-parameter groups are used for training, the K can be used for fine adjustment of a Schema-linking model, an initial PreTC set containing an initial table name and a column name is obtained by processing a user question, and the initial PreTC set is used for training. The initial PreTC set can be supplemented according to RT, TN, RC and CN to obtain an initial FinTC set, MoeN is used for training an SQL generation model, the trained SQL generation model can process a user problem and the initial FinTC set, an optimal hyper-parameter set with a good effect is obtained based on comparison of processing results, and the optimal hyper-parameter set can be used for processing the user problem and the initial FinTC set. K and MoeN in the optimal hyper-parameter set are used for generating more candidate tables and columns by the Schema-linking model and training the SQL generation model, RT, RC or TN and CN can supplement the set, and the accuracy of SQL generation can be improved through careful design of the candidate tables and columns screened by the Schema-linking model and adjustment training of the SQL generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a text2sql method based on large model ensemble learning. Background Art

[0002] Currently, text2sql tasks based on large model technology can be divided into two categories. One category involves prompt engineering technology based on large models. Through few-shot learning and thought chaining, similar questions and SQL structures in prompt word examples are used to guide large models to generate accurate SQL statements. However, this method requires large models with over 100 billion parameters to achieve an accuracy rate of over 70%, which requires high computing power resources. Furthermore, it requires the construction of a high-quality historical text2sql question-and-answer database to meet the requirements of few-shot learning.

[0003] Another type of method is an ensemble learning approach based on efficient supervised fine-tuning of large models. This approach first fine-tunes the schema-linking large model, then uses it to select the table and field names relevant to the question. Based on the table and column names selected by schema-linking, the large SQL generation model is then called to generate the SQL statement corresponding to the question. This type of approach already existed before the introduction of large model technology. For example, the RESDSQL algorithm first fine-tunes the discriminant NLP model Roberta to score the table and column names relevant to the question. It then selects the names of the four highest-scoring tables and the five highest-scoring columns within these tables. These are then input into the SQL generation model along with the question to generate the final SQL statement. The idea behind this ensemble approach is to improve the accuracy of the SQL generation model by pre-selecting the table and column names relevant to the question.

[0004] RESDSQL uses a discriminative model to score problem-related table and column names, while ensemble learning methods based on efficient supervised fine-tuning of large models use generative models to directly generate table and column names corresponding to the problem. For example, DTS-SQL first generates problem-related table names by fine-tuning the large model, and then uses the problem-related table and column names selected by the large model as input to the large SQL generation model. For the second type of ensemble learning method, the accuracy of SQL model generation also depends on the accuracy of schema-linking predictions. Both RESDSQL and DTS-SQL ensure recall by introducing redundant table and column names. For example, RESDSQL directly specifies the four highest-scoring tables and the five highest-scoring fields in each table as input to the SQL generation model. Although this method achieves a recall of 0.997, its precision is only 15%. Finding a balance between high recall and high precision is key to improving SQL generation accuracy. Summary of the Invention

[0005] In view of this, the present invention proposes a text2sql method based on large-model ensemble learning. By obtaining the optimal hyperparameter group to fine-tune the Schema-linking model and train the SQL generation model, the accuracy of SQL language generation can be improved.

[0006] The technical solution of the present invention is achieved as follows:

[0007] A text2sql method based on large model ensemble learning includes the following steps:

[0008] Step S1: Initialize hyperparameters K, RT, T_N, RC, C_N, and Moe_N, and divide the hyperparameters into a first hyperparameter group {K, RT, RC, Moe_N} and a second hyperparameter group {K, T_N, C_N, Moe_N}, where K is the number of beam searches, RT is the dynamic value of the number of tables in the input SQL generation model, T_N is the specified value of the number of tables in the input SQL generation model, RC is the dynamic value of the number of columns in the input SQL generation model, C_N is the specified value of the number of columns in the input SQL generation model, and Moe_N is the number of expert models.

[0009] Step S2: Use the large-model Lora algorithm to train the Schema-linking model, input the user question into the Schema-linking model, and use the Beam-Search method to output the corresponding initial table names and column names based on the hyperparameter K in the first hyperparameter group and the second hyperparameter group, and output the initial table names and column names as the corresponding initial PreTC sets;

[0010] Step S3: Supplement the initial PreTC set based on RT and RC in the first hyperparameter group and T_N and C_N in the second hyperparameter group, and obtain an initial FinTC set. Use the large-model Lora+Moe algorithm to train the SQL generation model, where the number of experts in the Moe algorithm is determined by the hyperparameter Moe_N in the first and second hyperparameter groups. Input the user question and the initial FinTC set into the trained SQL generation model, and compare the processing results of the SQL generation model to determine the optimal hyperparameter set.

[0011] Step S4: Input the user question into the Schema-linking model. The Schema-linking model uses the Beam-Search method to output relevant table names and column names based on the hyperparameter K in the optimal hyperparameter group, and form the optimal PreTC set.

[0012] Step S5: Supplement the optimal PreTC set based on RT, RC or T_N, C_N in the optimal hyperparameter group to obtain the optimal FinTC set. Use the large model Lora+Moe algorithm to train the SQL generation model based on the hyperparameter Moe_N in the optimal hyperparameter group. Input the user question and the optimal FinTC set into the trained SQL generation model to generate SQL query statements.

[0013] Preferably, after initializing the hyperparameters K, RT, T_N, RC, C_N, and Moe_N in step S1, a value range or a discrete value set is defined for each hyperparameter, where:

[0014] The value set of the hyperparameter K is: K∈{2,3,4,5,6,7};

[0015] The value set of the hyperparameter RT is: RT∈{0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1.0};

[0016] The value set of the hyperparameter T_N is: T_N∈{2,3,4,5};

[0017] The value set of the hyperparameter RC is: RC∈{0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1.0};

[0018] The value set of the hyperparameter C_N is: C_N∈{2,3,4,5,6};

[0019] The value set of the hyperparameter Moe_N is: Moe_N∈{2,3,4,5}.

[0020] Preferably, the specific steps of step S2 are:

[0021] Step S21: Use the large model Lora algorithm to train the Schema-linking model;

[0022] Step S21: The Schema-linking model uses the Beam-Search method to output corresponding initial table names and column names based on the hyperparameter K in the first hyperparameter group and the second hyperparameter group respectively.

[0023] Step S22: Use the n-gram algorithm to find the proposed table names and column names based on the user question, and combine them with the initial table names and column names corresponding to the first hyperparameter group and the second hyperparameter group to form an initial PreTC set.

[0024] Preferably, the specific steps of the Schema-linking model using the Beam-Search method to obtain table names and column names in step S2 and step S4 are as follows: when Beam-Search generates a partial sequence in each step, K optimal candidates are retained, and these candidate sequences are gradually expanded until a complete output sequence is generated. The K candidate sequences finally output are de-duplicated and merged as the output of the Schema-linking model.

[0025] Preferably, the specific steps of supplementing the initial PreTC set based on RT and RC in the first hyperparameter group and obtaining the initial FinTC set in step S3 include:

[0026] Step S31: For the first hyperparameter group, calculate the number of tables and columns required to be input into the SQL generation model for each sample based on the hyperparameters RT and RC;

[0027] Step S32: Use the RESDSQL scoring model to find the table and column with the highest score that does not appear in the initial PreTC set, and add them to the initial PreTC set to form the initial FinTC set.

[0028] Preferably, the specific steps of supplementing the initial PreTC set based on T_N and C_N in the second hyperparameter group and obtaining the initial FinTC set in step S3 include:

[0029] Step S33: For the second hyperparameter set, determine the number of tables and columns that are ultimately input into the SQL generation model based on T_N and C_N;

[0030] Step S34: determine whether the number of tables or columns for a certain sample prediction in the initial PreTC set is less than the hyperparameters T_N and C_N;

[0031] Step S35: If the result is less than the hyperparameters T_N and C_N, the RESDSQL scoring model is used to find the table and column with the highest score that does not appear in the initial PreTC set, and the table and column are added to the initial PreTC set to form the initial FinTC set.

[0032] Step S36: When the judgment result is greater than the hyperparameters T_N and C_N, the initial PreTC set is maintained unchanged and output as the initial FinTC set.

[0033] Preferably, in step S3, the user question and the initial FinTC set are input into the trained SQL generation model, and the processing results of the SQL generation model are compared to determine the optimal hyperparameter set. Specific steps also include:

[0034] Step S37: Input the initial FinTC set and the user question into the SQL generation model to predict the SQL statement;

[0035] Step S38: Compare the predicted SQL statement query results in the database with the results of the SQL statement query in Ground Truth to see if they are consistent, and determine the optimal hyperparameter set based on the comparison results.

[0036] Preferably, the specific steps of using the large model Lora+Moe algorithm to train the SQL generation model in steps S3 and S5 are:

[0037] Get Lora's expression:

[0038]

[0039] Where W+ΔW represents the sum of the original weight matrix W and the weight increment ΔW, which is the updated weight matrix. ΔW=BA means that the weight increment ΔW is decomposed into the product of two low-rank matrices B and A. is a low-rank matrix, where d in is the input dimension, r is the dimension of the low-rank factor, is another low-rank matrix, where d out is the output dimension;

[0040] The Moe architecture is introduced based on the Lora architecture to replace the linear layer. The architecture of the Lora+Moe algorithm is:

[0041]

[0042] Where W0 is the original weight matrix, x is the input vector, W0x represents the linear transformation of the input using the original weight matrix, o = W0x + ΔWx is the W + ΔW part of the Lora expression, which expresses the sum of the weight increments after Lora fine-tuning based on the weights of the original large model, and ΔWx is the weighted summation of the product of the original single low-rank matrix of Lora into the product of multiple low-rank matrices processed by multiple experts, that is, the part replaced by Moe:

[0043]

[0044] Where N is the number of experts, and its value is equal to the hyperparameter Moe_N, G(x) i is the gating function that determines which expert i, E, to assign the input x to. i (x) is the processing result of the i-th expert on input x.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] The present invention discloses a text2sql method based on large model ensemble learning. After initializing the hyperparameters K, RT, T_N, RC, C_N and Moe_N, the hyperparameters are divided into a first hyperparameter group and a second hyperparameter group. The Schema-linking model is used to first obtain the initial table name and column name based on the user question and output them as an initial PreTC set. Then, the initial PreTC set can be supplemented based on the first hyperparameter group and the second hyperparameter group to obtain an initial FinTC set. The initial FinTC set and the corresponding hyperparameter group can be input into the SQL generation model. The optimal hyperparameter group is determined by comparing the results of the SQL generation model. After obtaining the optimal hyperparameter group, Schema generation can be performed based on the optimal hyperparameter group. Fine-tune the Schema-linking model and train the SQL generation model. The trained Schema-linking model can first obtain relevant table and column names based on user questions and generate the optimal PreTC set. Then, based on the hyperparameter set, the optimal PreTC set is supplemented to obtain the optimal FinTC set. After the SQL generation model is trained, the user question and the optimal FinTC set can be input into the SQL generation model to generate SQL query statements. After limiting the range and value of the hyperparameters, a traversal method is used to determine the optimal hyperparameter set to achieve fine-tuning of the Schema-linking model and training of the SQL generation model, which can improve the accuracy of SQL language generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only preferred embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0048] Figure 1 This is a flowchart of a text2sql method based on large model ensemble learning of the present invention;

[0049] Figure 2 Flowchart of step S2 of the text2sql method based on large model ensemble learning of the present invention;

[0050] Figure 3 This is a flowchart of step S3 of the text2sql method based on large model ensemble learning of the present invention. DETAILED DESCRIPTION

[0051] In order to better understand the technical content of the present invention, a specific embodiment is provided below, and the present invention is further described in conjunction with the accompanying drawings.

[0052] See also Figures 1 to 3 The present invention provides a text2sql method based on large model ensemble learning, comprising the following steps:

[0053] Step S1: Initialize hyperparameters K, RT, T_N, RC, C_N, and Moe_N, and divide the hyperparameters into a first hyperparameter group {K, RT, RC, Moe_N} and a second hyperparameter group {K, T_N, C_N, Moe_N}, where K is the number of beam searches, RT is the dynamic value of the number of tables in the input SQL generation model, T_N is the specified value of the number of tables in the input SQL generation model, RC is the dynamic value of the number of columns in the input SQL generation model, C_N is the specified value of the number of columns in the input SQL generation model, and Moe_N is the number of expert models.

[0054] Step S2: Use the large-model Lora algorithm to train the Schema-linking model, input the user question into the Schema-linking model, and use the Beam-Search method to output the corresponding initial table names and column names based on the hyperparameter K in the first hyperparameter group and the second hyperparameter group, and output the initial table names and column names as the corresponding initial PreTC sets;

[0055] Step S3: Supplement the initial PreTC set based on RT and RC in the first hyperparameter group and T_N and C_N in the second hyperparameter group, and obtain an initial FinTC set. Use the large-model Lora+Moe algorithm to train the SQL generation model, where the number of experts in the Moe algorithm is determined by the hyperparameter Moe_N in the first and second hyperparameter groups. Input the user question and the initial FinTC set into the trained SQL generation model, and compare the processing results of the SQL generation model to determine the optimal hyperparameter set.

[0056] Step S4: Input the user question into the Schema-linking model. The Schema-linking model uses the Beam-Search method to output relevant table names and column names based on the hyperparameter K in the optimal hyperparameter group, and form the optimal PreTC set.

[0057] Step S5: Supplement the optimal PreTC set based on RT, RC or T_N, C_N in the optimal hyperparameter group to obtain the optimal FinTC set. Use the large model Lora+Moe algorithm to train the SQL generation model based on the hyperparameter Moe_N in the optimal hyperparameter group. Input the user question and the optimal FinTC set into the trained SQL generation model to generate SQL query statements.

[0058] The present invention provides a text2sql method based on large model ensemble learning, which is improved in the traditional ensemble learning method based on large model efficient supervised fine-tuning. The overall idea is to fine-tune the Schema-linking model, obtain the relevant table names and column names, and then use the set of relevant table names and column names and user questions as the input of the SQL generation model. The SQL generation model generates SQL query statements. However, there is a certain gap between the SQL generation model and the traditional SQL generation model. The present invention adopts the large model Lora+Moe algorithm to train the SQL generation model, in which the fine-tuning of the Schema-linking model and the training of the SQL generation model are Hyperparameters K and Moe_N are required during the training process. In addition to K and Moe_N, other hyperparameters include RT, T_N, RC, and C_N. The present invention optimizes the six hyperparameters to obtain the optimal hyperparameter group, and then fine-tunes, trains, and supplements the model based on the optimal hyperparameter group. That is, the method of the present invention can be divided into two processes: training and prediction. The methods used in the two processes are basically the same. The difference is that there are multiple hyperparameters in the training process, and traversal calculations are required. Accurate hyperparameters can be obtained based on the effect of SQL generation model, while the prediction process only performs fine-tuning training of the model and supplements the set under an accurate set of hyperparameters, thereby improving the accuracy of the prediction.

[0059] After initializing the above 6 hyperparameters, they are divided into the first hyperparameter group and the second hyperparameter group. Since the hyperparameters have different value ranges, the first hyperparameter group and the second hyperparameter group have different combinations, that is, there are several groups of hyperparameter combinations. The training process needs to train each hyperparameter group. The overall training process is: first, the Schema-linking model processes the user question, and based on the hyperparameter K in each first hyperparameter group and each second hyperparameter group, the corresponding initial table name and column name are obtained, and the corresponding initial table name and column name are output as the initial PreTC set. The number of the initial PreTC set is the same as the total number of the first hyperparameter group and the second hyperparameter group. In order to improve the processing accuracy of the SQL generation model, after obtaining the initial PreTC set Afterwards, the corresponding initial PreTC set can be supplemented based on RT and RC in the first hyperparameter group and T_N and C_N in the second hyperparameter group, and an initial FinTC set can be obtained. Then, the SQL generation model can be trained based on the hyperparameter Moe_N in the first hyperparameter group and the second hyperparameter group, that is, multiple trained SQL generation models can be obtained. After the training is completed, the corresponding initial FinTC set and user questions can be input into the SQL generation model, and the SQL generation model obtains the initial prediction results. The initial prediction results of the SQL generation model are compared to obtain a set of hyperparameter groups with better effects, which is the optimal hyperparameter group. At this point, the training process is completed. The six hyperparameters can be optimized through grid search to obtain the optimal hyperparameter group.

[0060] The optimal hyperparameter group includes K, Moe_N and RT, RC or T_N, C_N. At this time, the prediction process can be carried out. The prediction process is the same as the training process, except that the required hyperparameters have been determined. K and Moe_N can be used to fine-tune the Schema-linking model and train the SQL generation model, while RT, RC or T_N, C_N can be used to supplement the set. The Schema-linking model can process user questions and obtain relevant table names and column names, and then output the optimal PreTC set. The fine-tuned Schema-linking model can filter out some Some table and column names that are irrelevant to the question can be added to improve the accuracy of the SQL generation model. The optimal PreTC set can then be supplemented based on RT, RC or T_N, C_N to obtain the optimal FinTC set. The SQL generation model is trained based on the optimal hyperparameter Moe_N combined with the Lora+Moe algorithm. This can improve the problem that the Lora fine-tuned SQL generation model has poor recognition effect on some keywords, further improving the accuracy of the SQL generation model. The optimal FinTC set and user questions can be used as inputs to the SQL generation model, which generates accurate SQL query statements.

[0061] Preferably, after initializing the hyperparameters K, RT, T_N, RC, C_N, and Moe_N in step S1, a value range or a discrete value set is defined for each hyperparameter, where:

[0062] The value set of the hyperparameter K is: K∈{2,3,4,5,6,7};

[0063] The value set of the hyperparameter RT is: RT∈{0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1.0};

[0064] The value set of the hyperparameter T_N is: T_N∈{2,3,4,5};

[0065] The value set of the hyperparameter RC is: RC∈{0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1.0};

[0066] The value set of the hyperparameter C_N is: C_N∈{2,3,4,5,6};

[0067] The value set of the hyperparameter Moe_N is: Moe_N∈{2,3,4,5}.

[0068] Define a reasonable value range or discrete value set for each hyperparameter to avoid wasting computing resources due to an overly large search space. The value range is determined based on experimental data.

[0069] Preferably, the specific steps of step S2 are:

[0070] Step S21: Use the large model Lora algorithm to train the Schema-linking model;

[0071] Step S21: The Schema-linking model uses the Beam-Search method to output corresponding initial table names and column names based on the hyperparameter K in the first hyperparameter group and the second hyperparameter group respectively.

[0072] Step S22: Use the n-gram algorithm to find the proposed table names and column names based on the user question, and combine them with the initial table names and column names corresponding to the first hyperparameter group and the second hyperparameter group to form an initial PreTC set.

[0073] In order to increase the amount of data, the present invention introduces an n-gram algorithm. The n-gram algorithm can directly find the proposed table names and column names based on the user question. The found table names and column names are combined with the initial table names and column names to form the initial PreTC set.

[0074] Preferably, the specific steps of the Schema-linking model using the Beam-Search method to obtain table names and column names in step S2 and step S4 are as follows: when Beam-Search generates a partial sequence in each step, K optimal candidates are retained, and these candidate sequences are gradually expanded until a complete output sequence is generated. The K candidate sequences finally output are de-duplicated and merged as the output of the Schema-linking model.

[0075] The general Schema-linking model expression is: It means finding an optimal model through fine-tuning technology The probability under this model Maximum, where T and C represent q (question) and (Database schema information) The generated set of related table names and related column names.

[0076] The Beam-search method is used to allow Schema-linking to output different candidate table and column names. When generating partial sequences at each step, BeamSearch retains the K best candidates. It gradually expands these candidate sequences until a complete output sequence is generated. The final K candidate sequences are de-duplicated and merged as the staged output of the Schema-linking model.

[0077]

[0078] The above formula is the general expression of the beam-search algorithm, where represents the beam (i.e., the set of candidate sequences) at time step t, and Top-k represents the operator used to select the most likely K sequences from all possible candidate sequences. Each candidate sequence consists of the sequence itself and the probability of generating the sequence. The Top-k operation only retains those candidate sequences with the highest generation probability, thereby continuing to expand these most promising sequences in subsequent time steps. represents the output of the i-th sequence from time step 1 to t, Indicates the candidate sequence given the schema-linking model input X. Here X is the problem in the text2sql task and the database schema.

[0079] Preferably, the specific steps of supplementing the initial PreTC set based on RT and RC in the first hyperparameter group and obtaining the initial FinTC set in step S3 include:

[0080] Step S31: For the first hyperparameter group, calculate the number of tables and columns required to be input into the SQL generation model for each sample based on the hyperparameters RT and RC;

[0081] Step S32: Use the RESDSQL scoring model to find the table and column with the highest score that does not appear in the initial PreTC set, and add them to the initial PreTC set to form the initial FinTC set.

[0082] The specific steps of supplementing the initial PreTC set based on T_N and C_N in the second hyperparameter group and obtaining the initial FinTC set include:

[0083] Step S33: For the second hyperparameter set, determine the number of tables and columns that are ultimately input into the SQL generation model based on T_N and C_N;

[0084] Step S34: determine whether the number of tables or columns for a certain sample prediction in the initial PreTC set is less than the hyperparameters T_N and C_N;

[0085] Step S35: If the result is less than the hyperparameters T_N and C_N, the RESDSQL scoring model is used to find the table and column with the highest score that does not appear in the initial PreTC set, and the table and column are added to the initial PreTC set to form the initial FinTC set.

[0086] Step S36: When the judgment result is greater than the hyperparameters T_N and C_N, the initial PreTC set is maintained unchanged and output as the initial FinTC set.

[0087] The first hyperparameter group {K, RT, RC, Moe_N} and the second hyperparameter group {K, T_N, C_N, Moe_N} both contain the hyperparameter K for Schema-linking model fine-tuning and the hyperparameter Moe_N for SQL generation model training. Hyperparameters have several value ranges, and it is necessary to determine the optimal values of K and Moe_N. To further improve the accuracy of the SQL generation model, the initial PreTC set needs to be supplemented. Different hyperparameter groups have different supplementation methods. For the first hyperparameter group, after calculating the number of tables and columns for each sample based on the hyperparameters RT and RC, the RESDSQL scoring model is used to find the sample with the highest score that is not in the initial PreTC set. The tables and columns that appeared in the combination are added to the initial PreTC set to form the initial FinTC set. For the second hyperparameter group, the number of tables and columns finally input into the SQL generation model is determined based on T_N and C_N, and it is judged whether the number of tables or columns predicted by a sample in the initial PreTC set is less than the hyperparameters T_N and C_N. When the judgment result is less than, the same method as the first hyperparameter group is used to supplement and form the initial FinTC set. When the judgment result is greater than, the initial PreTC set is maintained unchanged and directly output as the initial FinTC set. For the second hyperparameter group, all samples in the generated initial FinTC set have the same number of tables and columns.

[0088] Preferably, in step S3, the user question and the initial FinTC set are input into the trained SQL generation model, and the processing results of the SQL generation model are compared to determine the optimal hyperparameter set. Specific steps also include:

[0089] Step S37: Input the initial FinTC set and the user question into the SQL generation model to predict the SQL statement;

[0090] Step S38: Compare the predicted SQL statement query results in the database with the results of the SQL statement query in Ground Truth to see if they are consistent, and determine the optimal hyperparameter set based on the comparison results.

[0091] The initial FinTC set and hyperparameter Moe_N are used by the SQL generation model for prediction. Moe_N can be used to train the SQL generation model. After training, the SQL generation model predicts SQL statements based on the initial FinTC set and user question processing, and compares the predicted results with the results of the SQL statement query in Ground Truth. The prediction effect is judged based on whether they are consistent, thereby obtaining a hyperparameter set with better results and outputting it as the optimal hyperparameter set. K and Moe_N in the optimal hyperparameter set can be used for fine-tuning and training the Schema-linking model, while RT, RC or T_N, C_N can be used to supplement the set.

[0092] Preferably, the specific steps of using the large model Lora+Moe algorithm to train the SQL generation model in steps S3 and S5 are:

[0093] Get Lora's expression:

[0094]

[0095] Where W+ΔW represents the sum of the original weight matrix W and the weight increment ΔW, which is the updated weight matrix. ΔW=BA means that the weight increment ΔW is decomposed into the product of two low-rank matrices B and A. is a low-rank matrix, where d in is the input dimension, r is the dimension of the low-rank factor, is another low-rank matrix, where d out is the output dimension;

[0096] The Moe architecture is introduced based on the Lora architecture to replace the linear layer. The architecture of the Lora+Moe algorithm is:

[0097]

[0098] Where W0 is the original weight matrix, x is the input vector, W0x represents the linear transformation of the input using the original weight matrix, o = W0x + ΔWx is the W + ΔW part of the Lora expression, which expresses the sum of the weight increments after Lora fine-tuning based on the weights of the original large model, and ΔWx is the weighted summation of the product of the original single low-rank matrix of Lora into the product of multiple low-rank matrices processed by multiple experts, that is, the part replaced by Moe:

[0099]

[0100] Where N is the number of experts, and its value is equal to the hyperparameter Moe_N, G(x) iis the gating function that determines which expert i, E, the input x is assigned to. i (x) is the processing result of the i-th expert on input x.

[0101] During the experiment, it was found that SQL statements with keywords such as "union", "except", and "intersect" had poor results under the traditional large-model Lora fine-tuning method. Therefore, the present invention combines the Moe-type architecture to improve the accuracy of SQL statements. The Moe architecture is introduced on the basis of the original Lora to replace the linear layer, thereby allowing the expert models within Moe to collaboratively learn and update the matrix. The Moe layer contains N expert models. For a given input x, the forward propagation process of the Moe layer can be mathematically represented as a weighted combination based on these expert models. There are two steps involved: (1) The participation program of each expert model is determined through the gating network of the input feature. The gating network is trained based on the sample data of text2sql; (2) The input x is sent to all experts respectively, and the output of each expert is weighted and summed according to the weight determined in the first step to obtain the final output.

[0102] After the SQL generation model is trained, the relevant table names and column names obtained by the Schema-linking model and user questions can be input, and the Schema-linking model can process them to generate SQL query statements.

[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A text2sql method based on large model ensemble learning, characterized in that: The following steps are involved: Step S1: Initialize hyperparameters K, RT, T_N, RC, C_N, and Moe_N, and divide the hyperparameters into a first hyperparameter group {K, RT, RC, Moe_N} and a second hyperparameter group {K, T_N, C_N, Moe_N}, where K is the number of beam searches, RT is the dynamic value of the number of tables in the input SQL generation model, T_N is the specified value of the number of tables in the input SQL generation model, RC is the dynamic value of the number of columns in the input SQL generation model, C_N is the specified value of the number of columns in the input SQL generation model, and Moe_N is the number of expert models. Step S2: Use the large-model Lora algorithm to train the Schema-linking model, input the user question into the Schema-linking model, and use the Beam-Search method to output the corresponding initial table names and column names based on the hyperparameter K in the first hyperparameter group and the second hyperparameter group, and output the initial table names and column names as the corresponding initial PreTC sets; Step S3: Supplement the initial PreTC set based on RT and RC in the first hyperparameter group and T_N and C_N in the second hyperparameter group, and obtain an initial FinTC set. Use the large-model Lora+Moe algorithm to train the SQL generation model, where the number of experts in the Moe algorithm is determined by the hyperparameter Moe_N in the first and second hyperparameter groups. Input the user question and the initial FinTC set into the trained SQL generation model, and compare the processing results of the SQL generation model to determine the optimal hyperparameter set. Step S4: Input the user question into the Schema-linking model. The Schema-linking model uses the Beam-Search method to output relevant table names and column names based on the hyperparameter K in the optimal hyperparameter group, and form the optimal PreTC set. Step S5: Supplement the optimal PreTC set based on RT, RC or T_N, C_N in the optimal hyperparameter group to obtain the optimal FinTC set. Use the large model Lora+Moe algorithm to train the SQL generation model based on the hyperparameter Moe_N in the optimal hyperparameter group. Input the user question and the optimal FinTC set into the trained SQL generation model to generate SQL query statements.

2. A text2sql method based on large model ensemble learning according to claim 1, characterized in that: After initializing the hyperparameters K, RT, T_N, RC, C_N, and Moe_N in step S1, a value range or discrete value set is defined for each hyperparameter, where: The value set of the hyperparameter K is: K∈{2,3,4,5,6,7}; The value set of the hyperparameter RT is: RT∈{0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1.0}; The value set of the hyperparameter T_N is: T_N∈{2,3,4,5}; The value set of the hyperparameter RC is: RC∈{0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1.0}; The value set of the hyperparameter C_N is: C_N∈{2,3,4,5,6}; The value set of the hyperparameter Moe_N is: Moe_N∈{2,3,4,5}.

3. The text2sql method based on large model ensemble learning according to claim 1 is characterized in that: The specific steps of step S2 are: Step S21: Use the large model Lora algorithm to train the Schema-linking model; Step S21: The Schema-linking model uses the Beam-Search method to output corresponding initial table names and column names based on the hyperparameter K in the first hyperparameter group and the second hyperparameter group respectively. Step S22: Use the n-gram algorithm to find the proposed table names and column names based on the user question, and combine them with the initial table names and column names corresponding to the first hyperparameter group and the second hyperparameter group to form an initial PreTC set.

4. The text2sql method based on large model ensemble learning according to claim 1 is characterized in that: The specific steps for the Schema-linking model in steps S2 and S4 to obtain table names and column names using the Beam-Search method are as follows: when Beam-Search generates a partial sequence in each step, it retains K optimal candidates and gradually expands these candidate sequences until a complete output sequence is generated. The final K candidate sequences are de-duplicated and merged as the output of the Schema-linking model.

5. The text2sql method based on large model ensemble learning according to claim 1 is characterized in that: The specific steps of supplementing the initial PreTC set based on RT and RC in the first hyperparameter group and obtaining the initial FinTC set in step S3 include: Step S31: For the first hyperparameter group, calculate the number of tables and columns required to be input into the SQL generation model for each sample based on the hyperparameters RT and RC; Step S32: Use the RESDSQL scoring model to find the table and column with the highest score that does not appear in the initial PreTC set, and add them to the initial PreTC set to form the initial FinTC set.

6. A text2sql method based on large model ensemble learning according to claim 5, characterized in that: The specific steps of supplementing the initial PreTC set based on T_N and C_N in the second hyperparameter group and obtaining the initial FinTC set in step S3 include: Step S33: For the second hyperparameter set, determine the number of tables and columns that are ultimately input into the SQL generation model based on T_N and C_N; Step S34: determine whether the number of tables or columns for a certain sample prediction in the initial PreTC set is less than the hyperparameters T_N and C_N; Step S35: If the result is less than the hyperparameters T_N and C_N, the RESDSQL scoring model is used to find the table and column with the highest score that does not appear in the initial PreTC set, and the table and column are added to the initial PreTC set to form the initial FinTC set. Step S36: When the judgment result is greater than the hyperparameters T_N and C_N, the initial PreTC set is maintained unchanged and output as the initial FinTC set.

7. The text2sql method based on large model ensemble learning according to claim 6 is characterized in that: In step S3, the user question and the initial FinTC set are input into the trained SQL generation model, and the processing results of the SQL generation model are compared to determine the optimal hyperparameter set. Specific steps also include: Step S37: Input the initial FinTC set and the user question into the SQL generation model to predict the SQL statement; Step S38: Compare the predicted SQL statement query results in the database with the results of the SQL statement query in Ground Truth to see if they are consistent, and determine the optimal hyperparameter set based on the comparison results.

8. The text2sql method based on large model ensemble learning according to claim 1, characterized in that: The specific steps of using the large model Lora+Moe algorithm to train the SQL generation model in steps S3 and S5 are as follows: Get Lora's expression: Where W+ΔW represents the sum of the original weight matrix W and the weight increment ΔW, which is the updated weight matrix. ΔW=BA means that the weight increment ΔW is decomposed into the product of two low-rank matrices B and A. is a low-rank matrix, where d in is the input dimension, r is the dimension of the low-rank factor, is another low-rank matrix, where d out is the output dimension; The Moe architecture is introduced based on the Lora architecture to replace the linear layer. The architecture of the Lora+Moe algorithm is: Where W0 is the original weight matrix, x is the input vector, W0x represents the linear transformation of the input using the original weight matrix, o = W0x + ΔWx is the W + ΔW part of the Lora expression, which expresses the sum of the weight increments after Lora fine-tuning based on the weights of the original large model, and ΔWx is the weighted summation of the product of the original single low-rank matrix of Lora into the product of multiple low-rank matrices processed by multiple experts, that is, the part replaced by Moe: Where N is the number of experts, and its value is equal to the hyperparameter Moe_N, G(x) i is the gating function that determines which expert i, E, to assign the input x to. i (x) is the processing result of the i-th expert on input x.