Text-to-SQL (Structured Query Language) generation method and system based on large language model fine tuning

By generating inference paths and performing direct preference optimization training, the existing Text-to-SQL methods fail to make full use of language reasoning capabilities and preference learning algorithms, which significantly improves the model's performance and generalization capabilities on Text-to-SQL tasks.

CN120144614APending Publication Date: 2025-06-13RENMIN UNIVERSITY OF CHINA
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510183465.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing fine-tuning-based Text-to-SQL approach fails to fully utilize language reasoning capabilities and preference learning algorithms, resulting in insufficient performance and generalization capabilities of models in complex inference tasks and more challenging scenarios.

Method used

By collecting triple data sets and generating inference paths, creating an optimized data set, supervised fine-tuning training for the basic model, obtaining the reference model, and then performing feedback on the database corresponding to the output results of the reference model for direct preference optimization training, obtaining the generative model.

Benefits of technology

It significantly improves the performance and generalization capabilities of the model on Text-to-SQL tasks, and the generated SQL query statements are more accurate, meet user needs, and significantly improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144614A_ABST
    Figure CN120144614A_ABST
Patent Text Reader

Abstract

The invention relates to a text-to-SQL (Structured Query Language) generation method and system based on large language model fine tuning. The method comprises the following steps: collecting a triple data set; generating a reasoning path for the triple data set, and obtaining an optimized data set comprising the triple data set and the corresponding reasoning path; selecting a basic model, and performing supervised fine tuning training on the basic model through the optimized data set to obtain a reference model; executing feedback through a database corresponding to an output result of the reference model, and performing direct preference optimization training on the reference model to obtain a generative model; and performing SQL generation on a to-be-generated text through the generation model to obtain a generation result. The method utilizes the advantages of the model language reasoning ability and the preference learning algorithm to improve the accuracy of the model in the aspect of text-to-SQL (Structured Query Language) generation tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method and system for text-to-SQL generation based on fine-tuning of a large language model. Background Art

[0002] The text-to-SQL generation (Text-to-SQL) task aims to translate natural language questions into executable SQL queries and has attracted extensive attention from both the natural language processing and database fields in recent years. With the rise of large language models (LLMs), the research paradigm of Text-to-SQL has gradually shifted to using a multi-agent framework to prompt closed-source large language models. However, in practical applications, using closed-source models will cause a series of problems such as high usage costs, data privacy risks, and slow generation speeds.

[0003] However, existing fine-tuning-based Text-to-SQL methods have two major limitations that restrict the performance of the model:

[0004] Insufficient utilization of language reasoning ability: Compared with directly generating answers, encouraging large language models to generate step-by-step chain-of-thought (CoT) can significantly improve the performance of complex reasoning tasks. CoT helps unlock the language reasoning ability obtained by the model during the pre-training stage by decomposing complex tasks into simpler logical steps. However, current Text-to-SQL evaluation sets usually do not provide the CoT path from the question to the target SQL query. Therefore, existing fine-tuning methods only train the model to directly generate SQL queries, skipping the intermediate reasoning steps. This omission causes the model to fail to fully utilize its pre-training knowledge, restricting its performance on existing evaluation sets and its generalization ability in more challenging scenarios.

[0005] Insufficient exploration of preference learning algorithms: After initial supervised fine-tuning (SFT), direct preference optimization (DPO) is one of the most widely adopted post-training strategies and has shown potential in complex tasks such as math word problems and code generation, which can further improve the performance of the model. However, although the database can provide accurate feedback to help construct paired training data for DPO, the exploration of this algorithm by existing technologies is very limited, and its role in improving the performance of Text-to-SQL models is not yet clear. Summary of the Invention

[0006] The present invention provides a method and system for text-to-SQL generation based on fine-tuning of a large language model to solve the defects of the prior art.

[0007] The present invention provides a method for generating text-to-SQL based on fine-tuning of a large language model, including:

[0008] S1: Collect a triple dataset;

[0009] S2: Generate an inference path for the triple dataset to obtain an optimized dataset including the triple dataset and the corresponding inference path;

[0010] S3: Select a base model and perform supervised fine-tuning training on the base model through the optimized dataset to obtain a reference model;

[0011] S4: Perform direct preference optimization training on the reference model through the feedback of the database corresponding to the output result of the reference model to obtain a generation model;

[0012] S5: Generate SQL for the text to be generated through the generation model to obtain a generation result.

[0013] In the method for generating text-to-SQL based on fine-tuning of a large language model provided by the present invention, the triple dataset in step S1 includes: input questions, database meta-information, and SQL.

[0014] In the method for generating text-to-SQL based on fine-tuning of a large language model provided by the present invention, step S3 further includes:

[0015] S31: Select a base model;

[0016] S32: Through a database schema retriever, retrieve database prompt words according to the optimized dataset;

[0017] S33: Train the base model through the optimized dataset and the generated database prompt words to obtain a reference model.

[0018] In the method for generating text-to-SQL based on fine-tuning of a large language model provided by the present invention, the base model is the XLM-RoBERTa-XL model.

[0019] In the method for generating text-to-SQL based on fine-tuning of a large language model provided by the present invention, step S32 further includes:

[0020] S321: Through a database schema retriever, identify relevant tables and relevant columns according to the input questions in the optimized dataset;

[0021] S322: Through a database value matching method, retrieve relevant values from the corresponding positions of the database according to the relevant tables and the relevant columns;

[0022] S323: Output the related table, the related column, the related value, and the corresponding primary key and foreign key as database prompts.

[0023] According to a method for generating text-to-SQL based on fine-tuning of a large language model provided by the present invention, the expression of the loss function of the reference model in step S33 is:

[0024]

[0025] where L SFT is the supervised fine-tuning loss function of the reference model, q is the input question in the optimization dataset, d is the database prompt, c is the inference path, π base is the base model, D S is the optimization dataset, and E represents the next-step prediction under the condition.

[0026] According to a method for generating text-to-SQL based on fine-tuning of a large language model provided by the present invention, step S4 further includes:

[0027] S41: Answer the input question through the reference model to generate an output result;

[0028] S42: Based on the database execution feedback corresponding to the output result, construct a preference data pair including the correct execution result and the incorrect execution result;

[0029] S43: Directly optimize and train the reference model through the preference data pair to obtain a generation model.

[0030] According to a method for generating text-to-SQL based on fine-tuning of a large language model provided by the present invention, step S42 further includes:

[0031] S421: Sample the question responses included in the output result to obtain response samples;

[0032] S422: Extract the predicted SQL from the response samples and execute the predicted SQL to obtain a predicted execution result;

[0033] S423: Judge the predicted execution result. When the predicted execution result is the same as the standard execution result of the standard SQL, the corresponding predicted execution result is the correct execution result. When the predicted execution result is different from the standard execution result of the standard SQL, the corresponding predicted execution result is the incorrect execution result;

[0034] S424: Construct a preference data pair including the correct execution result and the incorrect execution result.

[0035] A text-to-SQL generation method based on fine-tuning of a large language model according to the present invention, the expression of the loss function of the generation model in step S43 is:

[0036]

[0037] where L DPO is the loss function of the generation model, E represents the next-step prediction under the condition, q is the input question in the optimization dataset, d is the database prompt, c + is the question response corresponding to the correct execution result, c - is the question response corresponding to the wrong execution result, σ(·) is the Sigmoid function, β is the hyperparameter controlling the KL divergence penalty intensity, π DPO is the policy model initialized by the reference model for DPO training, π SFT is the reference model.

[0038] The present invention also provides a text-to-SQL generation system based on fine-tuning of a large language model, which is used to execute a text-to-SQL generation method based on fine-tuning of a large language model as described in any one of the above, including:

[0039] A collection module, which is used to collect a triple dataset;

[0040] An inference module, which is configured as a large language model and is used to generate an inference path for the triple dataset to obtain an optimization dataset including the triple dataset and the corresponding inference path;

[0041] A first training module, which is used to perform supervised fine-tuning training on the selected base model through the optimization dataset to obtain a reference model;

[0042] A second training module, which is used to perform direct preference optimization training on the reference model through the database execution feedback corresponding to the output result of the reference model to obtain a generation model;

[0043] A generation module, which is configured as the generation model trained by the second training module and is used to perform SQL generation on the text to be generated to obtain a generation result.

[0044] A method and system for text-to-SQL generation based on fine-tuning of a large language model provided by the present invention collect a data set and generate an inference path for it, creating an optimized data set, providing richer and more accurate information for subsequent model training, helping the model learn how to accurately extract key information from natural language questions and generate corresponding SQL query statements; secondly, through supervised fine-tuning training with the optimized data set, the model can perform customized learning for the generation task (text-to-SQL), thereby improving its performance in related tasks. At the same time, by introducing database prompt words, the model can better understand the structure and context information of the database, further improving the accuracy of SQL generation; in addition, by referring to the database execution feedback corresponding to the output result of the reference model, the present invention constructs preference data pairs and performs direct preference optimization training on the reference model, enabling the model to adjust its SQL generation strategy according to the correctness of the execution result, thus being more in line with the user's needs in practical applications; during the training process, the present invention constructs high-quality preference data pairs, providing strong support for the direct preference optimization of the model. Since the present invention can generate more accurate SQL query statements that meet the user's needs, it can significantly improve the user experience in practical applications and help users obtain the required data information more quickly and accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 It is a schematic flowchart of a method for text-to-SQL generation based on fine-tuning of a large language model provided by an embodiment of the present invention;

[0047] Figure 2 It is a schematic structural diagram of a system for text-to-SQL generation based on fine-tuning of a large language model provided by an embodiment of the present invention.

[0048] Reference numerals: 100, collection module; 200, inference module; 300, first training module; 400, second training module; 500, generation module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them, and they should not be construed as limitations on the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.

[0050] The embodiments of the present invention will be described below in conjunction with the drawings.

[0051] As Figure 1 shown, the present invention provides a text-to-SQL generation method based on fine-tuning of a large language model, including:

[0052] S1: Collect a triple dataset.

[0053] Among them, the triple dataset in step S1 includes: input questions, database meta-information, and SQL.

[0054] Further, the triple dataset in step S1 is collected based on existing publicly available Text-to-SQL datasets. The above input questions are questions or query requests put forward by users in natural language. The above database meta-information is information about the database structure, which describes the structure and relationships of tables, columns, data types, etc. in the database. For the Text-to-SQL task, database meta-information is necessary because it is required to help the model understand how to extract the required information from the database. The above SQL refers to the database query statement corresponding to the input question.

[0055] S2: Generate an inference path for the triple dataset to obtain an optimized dataset including the triple dataset and the corresponding inference path.

[0056] Further, different from most math word problem evaluation sets (such as MATH and GSM8K) that provide step-by-step solutions connecting the question to the final answer, the Text-to-SQL evaluation set lacks an inference path. To bridge this gap and reduce the cost of manual annotation, the present invention uses a large language model in step S2 to generate CoT inference paths for existing Text-to-SQL training data. Specifically, for each training sample, the present invention uses GPT-4o-mini to generate K = 16 diverse CoT inference paths, manifesting the process of gradually converting natural language questions into SQL queries.

[0057] S3: Select a base model and perform supervised fine-tuning training on the base model using the optimized dataset to obtain a reference model.

[0058] The purpose of step S3 is to train a smaller base model with the CoT-enhanced training set obtained in step S2, enabling the model to generate diverse answers with reasoning processes in the Text-to-SQL task.

[0059] Among them, step S3 further includes:

[0060] S31: Select a base model.

[0061] In step S31, the present invention selects the XLM-RoBERTa-XL model as the base model. Subsequently, the optimized dataset obtained in step S2 is input into the base model. The input dataset includes input questions, database meta-information, SQL query statements, and corresponding reasoning paths.

[0062] Among them, the base model is the XLM-RoBERTa-XL model.

[0063] XLM-RoBERTa-XL has shown strong performance in multiple natural language processing tasks. To improve the accuracy of recall, the present invention replaces the base model of the classifier from RoBERTa-Large (355M parameters) with XLM-RoBERTa-XL (3.5B parameters).

[0064] S32: Through a database schema retriever, retrieve database prompt words according to the optimized dataset.

[0065] Among them, step S32 further includes:

[0066] S321: Through a database schema retriever, identify relevant tables and relevant columns according to the input questions in the optimized dataset.

[0067] In step S321, first use the database schema retriever to identify the database tables and columns related to the question by matching the keywords in the input question with the table and column names in the database schema according to the input questions in the optimized dataset.

[0068] S322: Through a database value matching method, retrieve relevant values from the corresponding positions of the database according to the relevant tables and the relevant columns.

[0069] In step S322, after determining the relevant tables and relevant columns in step S321, use the database value matching method to retrieve the values related to these tables and columns from the corresponding positions of the database. The obtained relevant values are used to provide the model to construct more accurate SQL query statements.

[0070] S323: Output the related table, the related column, the related value, and the corresponding primary key and foreign key as database prompts.

[0071] In step S323, the identified related tables, related columns, related values, and the corresponding primary key and foreign key information are integrated and output as database prompts, which will be provided to the base model as additional input information to help it generate more accurate SQL query statements.

[0072] S33: Train the base model with the optimized dataset and the generated database prompts to obtain a reference model.

[0073] In step S3, the language model is fine-tuned using the CoT-enhanced dataset as a whole, enabling it to answer the Text-to-SQL task in a step-by-step manner. The input sequence includes not only the question itself but also the meta-information of the database, such as table and column names, primary key and foreign key relationships, and potentially useful database values.

[0074] The present invention first uses a database schema linker to identify the most relevant tables and columns according to the question; subsequently, the present invention adopts a "coarse-to-fine" database value matching method to retrieve the values related to the question from the database. These retrieved tables, columns, values, and the remaining primary key and foreign key together constitute the database prompts.

[0075] Finally, formally, the present invention denotes the question as q, the database prompt as d, and the output CoT inference path as c, and then uses the Conditional Next-token Prediction loss function for supervised fine-tuning to obtain the reference model.

[0076] Among them, the expression of the loss function of the reference model in step S33 is:

[0077]

[0078] Among them, L SFT is the supervised fine-tuning loss function of the reference model, q is the input question in the optimized dataset, d is the database prompt, c is the inference path, π base is the base model, D S is the optimized dataset, and E represents Conditional Next-token Prediction.

[0079] S4: Perform direct preference optimization training on the reference model through the database execution feedback corresponding to the output result of the reference model to obtain a generation model.

[0080] The main purpose of step S4 is to further optimize and train the reference model based on the execution feedback of the output result of the reference model on the database. The training method is direct preference optimization, which uses the actual execution effect of the model output to guide the training of the model, enabling the model to generate more accurate and effective SQL query statements.

[0081] Among them, step S4 further includes:

[0082] S41: Answer the input question through the reference model to generate an output result.

[0083] The present invention uses a reference model to answer a given input question. The input question is a query request put forward by the user in natural language. The reference model will generate one or more possible SQL query statements as the output result based on the knowledge it has learned. Specifically, for each input question, the DPO algorithm requires a pair of answers: a correct answer and a wrong answer. The reference model π obtained from step S3 in the present invention SFT samples N = 16 responses for each question.

[0084] S42: Based on the database execution feedback corresponding to the output result, construct a preference data pair including the correct execution result and the wrong execution result.

[0085] Among them, step S42 further includes:

[0086] S421: Sample the question responses included in the output result to obtain response samples.

[0087] S422: Extract the predicted SQL from the response samples and execute the predicted SQL to obtain the predicted execution result.

[0088] S423: Judge the predicted execution result. When the predicted execution result is the same as the standard execution result of the standard SQL, the corresponding predicted execution result is the correct execution result. When the predicted execution result is different from the standard execution result of the standard SQL, the corresponding predicted execution result is the wrong execution result.

[0089] S424: Construct a preference data pair including the correct execution result and the wrong execution result.

[0090] In step S42 and the corresponding steps S421 to S424, the predicted SQL is extracted from the CoT-style response obtained from step S1 and executed on the original database. Only when the execution result of the predicted SQL is consistent with the execution result of the standard SQL, the answer is regarded as correct; otherwise, it is regarded as wrong. For questions containing both correct and wrong answers, a correct response and a wrong response are randomly selected to form the preference data pair required for DPO training.

[0091] S43: Directly optimize and train the reference model using the preference data to obtain a generation model.

[0092] Formally, for a question q and its corresponding database prompt d, the present invention records the correct (or selected) response as c + , and the incorrect (or rejected) response as c - . The goal of DPO is to maximize the gap between the log-likelihoods of the selected and rejected answers, while ensuring that the deviation of the model from the reference model is not too large. The specific loss function is as follows.

[0093] Among them, the expression of the loss function of the generation model in step S43 is:

[0094]

[0095] Among them, L DPO is the loss function of the generation model, E represents the next-step prediction under the condition, q is the input question in the optimization dataset, d is the database prompt word, c + is the question response corresponding to the correct execution result, c - is the question response corresponding to the incorrect execution result, σ(·) is the Sigmoid function, β is the hyperparameter controlling the KL divergence penalty intensity, π DPO is the policy model initialized by the reference model for DPO training, and π SFT is the reference model.

[0096] S5: Use the generation model to perform SQL generation on the text to be generated to obtain a generation result.

[0097] To evaluate the robustness of the method proposed by the present invention, the present invention has conducted tests on a total of ten models with different model families (covering Deepseek, Qwen, Llama, and CodeS), different professional skills (general models, models dedicated to code or SQL generation), and different sizes (from 6.7 billion to 15 billion parameters).

[0098] The baseline algorithm compared by the present invention is to perform supervised fine-tuning on the training set of the Bird dataset according to the CodeS method, and the model directly returns SQL when receiving questions.

[0099] The present invention selects the Bird dataset as the evaluation set of the present invention. Bird is one of the most difficult and most realistic datasets widely used in the text-to-SQL generation task, and can fully reflect the capabilities of the model.

[0100] The present invention synthesizes the chain of thought, performs supervised fine-tuning, collects preference data, and conducts preference optimization learning on the training set of the Bird dataset to obtain the final model. In the test phase, all models are inferred using greedy decoding (temperature parameter is 0). The present invention is compared with the baseline method in terms of performance on the development set of the Bird dataset. The results are shown in Table 1. As can be seen from Table 1, the performance of the present invention on multiple base models is significantly better than the baseline algorithm.

[0101] Table 1 Comparison of the performance results of the method of the present invention with other base models

[0102] Base model Baseline algorithm The present invention Deepseek-llm-7b-chat 51.8 55.9 Llama-3.1-instruct-7b 59.0 61.2 Qwen2.5-7B-Instruct 58.8 61.9 Qwen2.5-14B-Instruct 64.3 65.3 Deepseek-coder-6.7b-instruct 60.6 63.8 CodeLlama-7b-Instruct-hf 57.0 61.3 CodeLlama-13b-Instruct-hf 60.0 63.9 Qwen2.5-Coder-7B-Instruct 61.6 63.4 CodeS-7b 56.8 57.5 CodeS-15b 58.3 61.1

[0103] As Figure 2 shown, the present invention also provides a text-to-SQL generation system based on large language model fine-tuning, including:

[0104] A collection module 100, which is used to collect triple datasets;

[0105] An inference module 200, which is configured as a large language model and is used to generate an inference path for the triple dataset to obtain an optimized dataset including the triple dataset and the corresponding inference path;

[0106] A first training module 300, which is used to perform supervised fine-tuning training on the selected base model through the optimized dataset to obtain a reference model;

[0107] A second training module 400, which is used to perform direct preference optimization training on the reference model through the database execution feedback corresponding to the output result of the reference model to obtain a generation model;

[0108] A generation module 500, which is configured as the generation model trained by the second training module and is used to perform SQL generation on the text to be generated to obtain a generation result.

[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0110] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0111] A method and system for text-to-SQL generation based on fine-tuning of a large language model provided by the present invention proposes a synthesis of a chain of thought (CoT) and a strategy for direct preference optimization (DPO), which solves the problem that the existing text-to-SQL generation methods based on fine-tuning of a large language model cannot fully utilize the model's language reasoning ability and the advantages of preference learning algorithms. The method proposed by the present invention can further improve the performance of the model in the text-to-SQL generation task.

[0112] A method and system for text-to-SQL generation based on fine-tuning of a large language model provided by the present invention provides a new solution idea for the Text-to-SQL task by introducing a CoT (Conditioned Next-token Prediction) inference path. The CoT inference path can explicitly convert the natural language problem into an SQL query step by step, which not only enhances the interpretability of the model but also helps the model learn more accurate and effective conversion rules. Therefore, compared with the traditional Text-to-SQL method, the present invention has significantly improved the accuracy and efficiency of generating SQL query statements.

[0113] When constructing the triple dataset, the present invention makes full use of the existing publicly available Text-to-SQL datasets and automatically generates CoT inference paths through a large language model, thus avoiding a large amount of manual annotation work. This not only saves time and labor costs but also improves the diversity and richness of the dataset, providing more comprehensive data support for the training of the model.

[0114] The XLM-RoBERTa-XL model selected in the present invention is used as the base model, and supervised fine-tuning training and direct preference optimization training are carried out on this basis. The XLM-RoBERTa-XL model itself has shown strong performance in multiple natural language processing tasks. After being trained by the present invention, the model can better adapt to different database structures and query requirements, thus showing stronger generalization ability in the Text-to-SQL task.

[0115] In the direct preference optimization training stage, the present invention further optimizes the reference model by constructing preference data pairs containing correct execution results and incorrect execution results, enabling the model to more accurately evaluate the execution effects of different SQL query statements and tend to generate SQL query statements that can be correctly executed and return the expected results. At the same time, due to the introduction of the CoT reasoning path, the model also shows higher diversity when generating SQL query statements.

[0116] Finally, the present invention generates SQL for the text to be generated through the generation model, obtains the generation result, simplifies the usage process of the Text-to-SQL task, enables users to obtain the corresponding SQL query statement only by inputting a natural language question, not only improves the convenience of using the task, but also provides a more user-friendly query interface.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A text-to-SQL generation method based on fine-tuning of a large language model, characterized in that: include: S1: Collect triplet dataset; S2: Generate a reasoning path for the triple data set to obtain an optimized data set including the triple data set and the corresponding reasoning path; S3: Select a basic model, and perform supervised fine-tuning training on the basic model using the optimized data set to obtain a reference model; S4: performing feedback through a database corresponding to the output result of the reference model, performing direct preference optimization training on the reference model, and obtaining a generation model; S5: Generate SQL for the text to be generated by using the generation model to obtain a generation result.

2. A text-to-SQL generation method based on large language model fine-tuning according to claim 1, characterized in that: The triplet data set in step S1 includes: input question, database metadata, and SQL.

3. The method for generating text to SQL based on fine-tuning of a large language model according to claim 1, characterized in that: Step S3 further comprises: S31: Select a basic model; S32: using a database pattern recaller, searching and obtaining a database prompt word according to the optimized data set; S33: The basic model is trained using the optimized data set and the generated database prompt words to obtain a reference model.

4. The method for generating text to SQL based on fine-tuning of a large language model according to claim 3, characterized in that: The basic model is the XLM-RoBERTa-XL model.

5. The method for generating text to SQL based on fine-tuning of a large language model according to claim 3, characterized in that: Step S32 further includes: S321: using a database pattern recaller, identifying and obtaining relevant tables and relevant columns according to input questions in the optimized data set; S322: using a database value matching method, according to the relevant table and the relevant column, searching for relevant values ​​from corresponding positions in the database; S323: Output the related table, the related column, the related value and the corresponding primary key and foreign key as a database prompt word.

6. The method for generating text to SQL based on fine-tuning of a large language model according to claim 3, characterized in that: The expression of the loss function of the reference model in step S33 is: Among them, L SFT is the supervised fine-tuning loss function of the reference model, q is the input question in the optimization dataset, d is the database prompt word, c is the reasoning path, and π base As the basic model, D S To optimize the dataset, E represents the conditional next step prediction.

7. The method for generating text to SQL based on fine-tuning of a large language model according to claim 1, characterized in that: Step S4 further comprises: S41: answering the input question through the reference model to generate an output result; S42: constructing a preference data pair including a correct execution result and an incorrect execution result based on the database execution feedback corresponding to the output result; S43: Performing direct preference optimization training on the reference model using the preference data to obtain a generated model.

8. The method for generating text to SQL based on fine-tuning of a large language model according to claim 7, characterized in that: Step S42 further includes: S421: Sampling the responses to the questions included in the output result to obtain response samples; S422: extracting a predicted SQL from the response sample and executing the predicted SQL to obtain a predicted execution result; S423: judging the predicted execution result; when the predicted execution result is the same as the standard execution result of the standard SQL, the corresponding predicted execution result is a correct execution result; when the predicted execution result is different from the standard execution result of the standard SQL, the corresponding predicted execution result is an incorrect execution result; S424: Construct a preference data pair including a correct execution result and an incorrect execution result.

9. The method for generating text to SQL based on fine-tuning of a large language model according to claim 7, characterized in that: The expression of the loss function of the generation model in step S43 is: Among them, L DPO is the loss function of the generated model, E represents the conditional next step prediction, q is the input question in the optimization dataset, d is the database prompt word, and c + The correct execution result corresponds to the question response, c - is the problem response corresponding to the error execution result, σ(·) is the Sigmoid function, β is the hyperparameter controlling the KL divergence penalty strength, and π DPO The policy model for DPO training initialized by the reference model, π SFT For reference model.

10. A text-to-SQL generation system based on large language model fine-tuning, used to execute a text-to-SQL generation method based on large language model fine-tuning as claimed in any one of claims 1 to 9, characterized in that: include: A collection module, wherein the collection module is used to collect a triplet data set; A reasoning module, wherein the reasoning module is configured as a large language model, and is used to generate a reasoning path for the triple data set, and obtain an optimized data set including the triple data set and the corresponding reasoning path; A first training module, the first training module is used to perform supervised fine-tuning training on the selected basic model through the optimized data set to obtain a reference model; A second training module, the second training module is used to perform feedback on the database corresponding to the output result of the reference model, perform direct preference optimization training on the reference model, and obtain a generated model; A generation module is configured as a generation model trained by the second training module, and performs SQL generation on the text to be generated to obtain a generation result.

Citation Information

Cited By

  • Mathematical test question generation and solution collaborative enhancement method and system based on large model

    CN120822602A

  • Table question and answer method based on large model and retrieval enhancement

    CN120929485A

  • A table-based question answering method based on large models and retrieval enhancement

    CN120929485B

  • Content security reasoning auditing method based on knowledge enhancement

    CN121189459A