Multi-LLM model integrated reasoning method for Text2SQL task

Through the multi-LLM model integration inference method, selecting the best model and conducting context learning, the existing Text-to-SQL method has solved the shortcomings in accuracy, speed and resource consumption, and achieved more efficient data query.

CN120218227APending Publication Date: 2025-06-27FUDAN UNIVERSITY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510160645.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing Text-to-SQL methods have insufficient accuracy and generation speed, and are costly to consume a lot of resources.

Method used

Multi-LLM model integrated inference method is adopted, and by establishing a simulation database, matching similar problems, selecting the best model, and using context learning to generate SQL statements.

Benefits of technology

Improve the accuracy and generation speed of Text-to-SQL tasks, reduce resource consumption, and achieve more efficient data query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218227A_ABST
    Figure CN120218227A_ABST
Patent Text Reader

Abstract

The invention provides a multi-LLM model integrated reasoning method for a Text2SQL task, in the method, a plurality of different LLM models are integrated, the most suitable LLM model is selected to answer a to-be-processed question according to the performance of each LLM model on similar questions, and a small number of sample prompts in a prompt word project are adopted to further improve the reasoning accuracy, so that the reasoning efficiency is improved. Therefore, each LLM can answer the most skilled problem, the reasoning accuracy and reasoning speed of the Text-to-SQL task are considered, fine tuning or re-training of the LLM model is not needed, and the resource consumption is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field, and particularly to a multi-LLM model integrated reasoning method for the Text2SQL task. Background Art

[0002] In the digital revolution era, data has become an important production factor driving human activities. Data is everywhere, from business operations to scientific research. However, with the popularization of electronic devices and the explosive growth of data volume, querying and exploring this data has become increasingly complex and difficult, even for experts. Existing data query methods are either form-based, which are easy to use but have limited query capabilities, or low-level tools that allow users to synthesize queries using the underlying database query language (such as SQL), but these tools are only applicable to a few people (such as SQL experts). This situation makes it difficult for ordinary people to access and utilize data, and even for practitioners, it is very troublesome to write a large number of query statements in different database and application scenarios and ensure their correctness.

[0003] To lower the threshold of data query, the Text-to-SQL task emerged as the times require. Text-to-SQL converts natural language queries into structured query language (SQL) commands, enabling users to interact with databases using natural language. In this way, we can eliminate the technical barriers hindering data access, break the barrier between natural language and structured data, and enable everyone to access, use, understand, and extract value from data. Figure 4 Shows a problem description of Text-to-SQL. Figure 5 Shows the general Text-to-SQL process in the prior art.

[0004] Given input:

[0005] A natural language question: Q = {q1, q2, ……, q n}, where q i represents the i-th question in this question. For example: What is the name of the employee with the highest salary?

[0006] A database schema: S = {T1, T2, ……, T n}, where T i represents the i-th table in the database, and each table T i includes columns C i1 、C i2 , …, C ip (where C ip is the P-th column attribute of T i ). For example: Table: Employees(ID, Name, Salary)

[0007] The goal is to generate an SQL query:

[0008] SQL query: S = {y1, y2, ……, y n}, where y i is the i-th token in the generated query. For example: SELECT Name FROM Employees ORDER BY Salary DESC LIMIT 1.

[0009] For such tasks, traditional text-to-SQL methods are mainly LSTM-based and Transformers-based. The LSTM-based method is one of the earliest deep learning methods applied to the text-to-SQL task. The LSTM model generates an SQL query by learning the context representations of the input natural language question and the database table. It can effectively model the sequential dependencies between the natural language question and the SQL query. For example, Bi-LSTM is used to learn the semantic representations of question-SQL pairs. LSTM and its variants capture the context through sequential processing, but it is relatively difficult to handle long-distance dependencies in complex queries. However, the Transformer-based method, by introducing the self-attention mechanism, can handle long-distance dependencies more effectively, thus significantly improving the ability to generate accurate SQL queries, especially when dealing with complex database schemas. For example, models like GraPPa, GAP, StruG, etc. enrich schema understanding through enhanced pre-training and context representations. It can introduce innovative technologies such as attention mechanisms and pre-training, which can greatly improve performance, especially in schema linking and complex query generation tasks. However, Transformers require a large amount of high-quality training data and have high requirements for data diversity.

[0010] The Text-to-SQL with Prompt method uses prompt engineering to enhance the performance and reliability of language models by carefully designing input prompts to guide the model to generate relevant and meaningful outputs. Different from traditional LSTM- and Transformer-based methods, prompt engineering relies on large-scale pre-trained language models (LLMs) without additional fine-tuning. This method directly creates SQL by identifying specific prompts or sentences and gradually generating the next high-probability output. Its advantages are that there is no additional computational cost and the results can be obtained quickly. This technique is particularly crucial in solving problems such as "hallucinations". The disadvantage is that it is highly dependent on data and the training model. If the ability of the training model itself is not strong enough or the data selection of the prompt samples is not accurate enough, the generated SQL may not be accurate enough. The methods of prompt engineering mainly include the following:

[0011] I. Zero-shot Prompt: Without task-specific training data, the model directly makes preliminary judgments and generates SQL based on a large amount of data. This method is suitable for quickly applying new tasks, but the accuracy may not be high.

[0012] II. Few-shot Prompt: Given a small number of examples, the model generates more accurate SQL, especially suitable for complex tasks. Through multiple prompts, the model can explore more parameter spaces and thus output the most suitable SQL.

[0013] III. Chain of Thought (CoT): Stimulates the model's complex thinking ability and generates accurate SQL through step-by-step reasoning. Combining with example prompts, this method can significantly improve the accuracy of generating complex or ambiguous SQL.

[0014] The Text-to-SQL with Fine-Tuning method relies on a pre-trained model and further fine-tunes it to adapt to specific tasks and domains. Compared with low-cost prompt methods, this method can not only significantly improve the performance of the pre-trained model in a specific domain but also perform well in handling complex queries and cross-domain tasks. The disadvantages are the dependence on pre-trained models, the need for fine-tuning training data, and computing resources. The fine-tuning methods are mainly divided into two categories:

[0015] I. Full-parameter fine-tuning: Fine-tune all parameters of the model, suitable for tasks that require high precision. By further supervised training on a specific dataset, the performance of the model in a specific task is optimized, such as Knowledge-to-SQL, etc.

[0016] II. Parameter-efficient fine-tuning: Only fine-tune some parameters of the model, usually specific layers or modules. This method can effectively reduce training time and computing resource consumption while maintaining high performance, such as DAIL-SQL.

[0017] The Text-to-SQL with Task-Training method trains a specific Text-to-SQL model from scratch, using training strategies similar to large language models (LLMs), such as Transformers and mixture-of-experts models. This method does not need to rely on pre-trained models like the Fine-Tuning method but directly trains from scratch and is suitable for SQL generation for specific tasks. Through specially designed training strategies, this method can better adapt to complex database structures and diverse query requirements. The disadvantage is that it requires a much larger amount of training data and computing resources, and the training process is complex and time-consuming.

[0018] The Text-to-SQL with LLM agent method collaborates multiple agents to dynamically generate and correct SQL queries, improving query accuracy and execution efficiency. This method uses a decomposition agent to break down complex SQL queries into sub-problems and gradually generate the final SQL query; a selection agent is used to filter databases and reduce interference from specific data. This method can repair incorrect SQL queries through external tools to ensure that the generated queries are consistent with expectations. Some systems such as MAGIC automatically generate self-correction guidelines by iteratively correcting incorrect queries, improving the reliability of query generation. The disadvantage is that there are many intermediate links in generating SQL and error correction through iteration, so the time to output SQL will be relatively long.

[0019] The above briefly describes the existing Text-to-SQL methods. Generally speaking, the existing related methods have the following disadvantages:

[0020] Firstly, the accuracy is not high. The inventor found through testing that the accuracy of the model generating SQL is indeed positively correlated with the model size; however, beyond a certain accuracy range, simply increasing the model size is difficult to further improve the accuracy of the model. Improving the accuracy through the Text-to-SQL with Prompt method by optimizing the prompt words rather than the model can only achieve limited results.

[0021] Secondly, the inference speed is slow. Although the accuracy of the model generating SQL is positively correlated with the model size, the model size cannot be blindly increased because a larger model means more inference resources are required during inference and also slower inference speed. Another example is the Text-to-SQL with LLM agent method, which is a method that sacrifices time for accuracy.

[0022] Thirdly, the resource consumption is large. For the Text-to-SQL with Fine-Tuning and Text-to-SQL with Task-Training methods, both require a large amount of accurate and diverse training data, and both consume a large amount of computing resources during the fine-tuning or retraining process. Summary of the Invention

[0023] The present invention is made to solve the above problems, and aims to provide a new method that can simultaneously balance the accuracy and generation speed of Text-to-SQL and has low resource consumption. The present invention adopts the following technical solutions:

[0024] The present invention provides a multi-LLM model integrated inference method for the Text2SQL task, which has the following technical features and includes the following steps: Step S1, establish a simulated database according to the application scenario, where the simulated database contains text descriptions of multiple questions and corresponding correct SQL statements; Step S2, match the text description of the problem to be processed with the text descriptions of the problems in the simulated database, and select the top n problems with the highest similarity to the problem to be processed as similar problems; Step S3, use each of the integrated multiple large language models to generate SQL statements for the n similar problems respectively, and score the generation results of each large language model; Step S4, based on the scores, select one of the integrated multiple large language models as the target large language model for answering the problem to be processed; Step S5, add some of the n similar problems to the prompt template of the target large language model to enable the target large language model to perform context learning; Step S6, use the target large language model that has undergone context learning to answer the problem to be processed, thereby generating the corresponding SQL statement.

[0025] The multi-LLM model integrated inference method for the Text2SQL task provided by the present invention may also have the following technical features, wherein Step S2 includes the following sub-steps: Step S2-1, perform domain keyword masking on the text descriptions of each of the problems in the simulated database and the text description of the problem to be processed; Step S2-2, convert the masked text descriptions of each of the problems and the masked text description of the problem to be processed into corresponding text vectors respectively; Step S2-3, calculate the similarity between the text vector of the problem to be processed and the text vectors of each of the problems according to the similarity calculation algorithm; Step S2-4, according to the calculated similarity and the predetermined number n of similar problem selections, select the n problems with the highest similarity to the problem to be processed as the similar problems.

[0026] The multi-LLM model integrated inference method for the Text2SQL task provided by the present invention may also have the following technical features, wherein in Step S2-1, based on the preset keywords, match and mask the corresponding vocabulary in the text descriptions of the problems and the text description of the problem to be processed, and in Step S2-4, use cosine similarity for the matching of the text vectors:

[0027]

[0028] In the formula, is the text vector of the problem to be processed, is the text vector of the problem in the simulated database, x iis a vector The value of the i-th dimension, y i is a vector The value of the i-th dimension, and n is the dimension size of the vector. Find the top n of the above problems that minimize cos(θ) as the similar problems.

[0029] The multi-LLM model integrated inference method for the Text2SQL task provided by the present invention may also have such a technical feature that in step S2-4, n = 5.

[0030] The multi-LLM model integrated inference method for the Text2SQL task provided by the present invention may also have such a technical feature that in step S3, the multiple large language models integrated include SQLCoder, DeepSeek-22b, Llama3, and Qwen2.

[0031] The multi-LLM model integrated inference method for the Text2SQL task provided by the present invention may also have such a technical feature that in step S3, for each of the integrated large language models, determine whether the SQL statement generation result of the large language model on the n similar problems is correct according to the corresponding correct SQL statement, and count the number of the similar problems answered correctly as the score.

[0032] The multi-LLM model integrated inference method for the Text2SQL task provided by the present invention may also have such a technical feature that in step S3, for each of the integrated large language models, determine whether the SQL statement generation result of the large language model on the n similar problems is correct according to the corresponding correct SQL statement, and calculate the score by weighting different similar problems:

[0033]

[0034] In the formula, r i Indicates whether the output SQL result of the model for the i-th similar problem is correct, corresponding to 0 or 1; t i Is the weight corresponding to the i-th problem, and this weight corresponds to the similarity between the problem and the problem to be processed.

[0035] The multi-LLM model integrated inference method for the Text2SQL task provided by the present invention may also have such a technical feature that in step S4, select the large language model with the highest score as the target large language model. When there are multiple large language models with the highest score, select the large language model with the largest number of parameters among the multiple large language models as the target large language model.

[0036] The multi-LLM model integrated reasoning method for the Text2SQL task provided by the present invention may also have the following technical feature: in step S5, the database table structure involved in the problem to be processed and at most m similar problems with similarity scores higher than a predetermined threshold in the similar problems are added to the prompt template of the target large language model.

[0037] The multi-LLM model integrated reasoning method for the Text2SQL task provided by the present invention may also have the following technical feature: in step S5, m = 5.

[0038] Functions and effects of the invention

[0039] According to the multi-LLM model integrated reasoning method for the Text2SQL task provided by the present invention, compared with the general Text2SQL generation methods in the prior art, it has the following aspects

[0040] Beneficial effects:

[0041] First, the reasoning effect is better. Since the most suitable LLM model is selected to answer the problem to be processed according to the performance of each LLM model on similar problems, that is, each LLM answers the questions it is best at, and few-shot prompting in prompt engineering is used to further improve the reasoning accuracy, the reasoning accuracy is high.

[0042] Second, the reasoning speed is fast. Although the LLM model with a large number of parameters has a high accuracy, due to its large number of parameters, it requires a lot of hardware resources for reasoning, so the speed is very slow, while the LLM model with a small number of parameters is the opposite. The method of the present invention can well retain the respective advantages of the large model and the small model, so as to maintain a fast reasoning speed while ensuring the reasoning accuracy.

[0043] Third, the resource consumption is small. Since the most suitable LLM model is selected according to the performance of each LLM model on similar problems, and a small number of high-quality similar problems are used to provide certain context information for it, without the need to fine-tune or retrain the LLM model, the computational resource consumption in this regard can be significantly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a schematic diagram of the multi-LLM model integrated reasoning architecture for the Text2SQL task in an embodiment of the present invention;

[0045] Figure 2 is a flowchart of the multi-LLM model integrated reasoning method for the Text2SQL task in an embodiment of the present invention;

[0046] Figure 3It is the flowchart of step S2 in the embodiment of the present invention;

[0047] Figure 4 It is a schematic diagram of the problem description of Text-to-SQL in the prior art;

[0048] Figure 5 It is a schematic flowchart of general Text-to-SQL in the prior art. Detailed implementation manners

[0049] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the multi-LLM model integrated inference method for the Text2SQL task of the present invention will be specifically described below in conjunction with embodiments and the accompanying drawings.

[0050] <Embodiment>

[0051] This embodiment provides a multi-LLM model integrated (collaborative) inference method for the Text2SQL task. For the convenience of explaining the method flow, the model architecture involved in this method will be briefly described below.

[0052] Figure 1 It is a schematic diagram of the multi-LLM model collaborative inference architecture for the Text2SQL task in this embodiment.

[0053] As Figure 1 shown, the model architecture involved in this method includes a similar problem selection module 10, a large language model (LLM) selection module 20, and an in-context learning module 30.

[0054] Among them, the similar problem selection module 10 includes a simulation problem library 11 (simulation database) in a specific field. This module is used to select several similar problems with the highest similarity to the problem to be processed from the simulation problem library 11, and provide these similar problems to the large language model selection module 20 and the in-context learning module 30 respectively.

[0055] The large language model selection module 20 statistically analyzes the performance of multiple integrated large language models on the selected similar problems, and selects a large language model with the best performance to answer the problem to be processed.

[0056] The in-context learning module 30 generates corresponding in-context learning prompts based on the selected similar problems to provide some detailed knowledge to the selected large language model, further improving the accuracy of the large language model in answering the problem to be processed.

[0057] The specific algorithms involved in the above modules will be specifically described in the following corresponding steps.

[0058] Figure 2 This is the flowchart of the collaborative inference method of the multi-LLM model for the Text2SQL task in this embodiment.

[0059] As Figure 2 shown, the method of this embodiment includes the following steps:

[0060] Step S1, establish a simulated database according to the application scenario, where the simulated database contains text descriptions of multiple questions and corresponding correct SQL statements.

[0061] Step S2, match the text description of the question to be processed with the text descriptions of the questions in the simulated database, and select the top n questions with the highest similarity to the question to be processed as similar questions.

[0062] Step S3, use each of the integrated multiple large language models (LLMs) to generate SQL statements for the n similar questions respectively, and score the generation results of each large language model.

[0063] Step S4, select one large language model from the integrated multiple large language models based on the score as the target large language model for answering the question to be processed.

[0064] Step S5, add some of the multiple similar questions to the prompt template of the target large language model to enable the target large language model to perform context learning.

[0065] Step S6, use the target large language model after context learning to answer the question to be processed, thereby generating the corresponding SQL statement.

[0066] That is to say, the method of this embodiment first establishes a simulated database according to the application scenario, and then records the specific performance of the integrated multiple LLMs on these questions. When a Text-to-SQL question is input, this method will retrieve the most similar questions in this simulated database, and then count the performance of these LLMs on these similar questions, and select a model with the best performance to answer this question to obtain an accurate SQL statement generation result. The following will elaborate on each of the above steps.

[0067] Step S1, establish a simulated database according to the application scenario, where the simulated database contains text descriptions of multiple questions and corresponding SQL statements.

[0068] Among them, the simulated database is a collection of domain-similar questions, and its main stored content is: the text description of each question and the corresponding SQL statement, and also includes certain noun markers.

[0069] In this embodiment, a simulation database for the financial scenario is built based on Cspider. Here, Cspider is a large-scale Chinese complex and cross-domain semantic interpretation and text-to-SQL dataset, which is translated from Spider by NLP researchers and computer science students.

[0070] Step S2: Match the text description of the problem to be processed with the text descriptions of the problems in the simulation database, and select the top n problems with the highest similarity to the problem to be processed as similar problems.

[0071] Figure 3 It is the flowchart of step S2 in this embodiment.

[0072] As Figure 3 shown, step S2 specifically includes the following sub-steps:

[0073] Step S2-1: Perform domain keyword masking on the text descriptions of each problem in the simulation database and the text description of the problem to be processed respectively.

[0074] Among them, the domain keyword masking is to match and mask the corresponding vocabulary in the text description of the problem and the text description of the problem to be processed based on the preset keywords.

[0075] For example, the text descriptions of two problems are respectively query the number of people in XX school and query the principal of XX school, which involve the same school name. If the text descriptions of these two problems are directly matched, it is very likely that the school name will be included in the calculation of the matching degree, resulting in a high similarity between these two problems. However, in fact, the SQL structures of these two problems are completely different. But in the method of this embodiment, by masking the specific domain keywords and numerals and other words that do not affect the overall semantics, the model can pay more attention to the SQL structure expressed by the problem, so as to obtain a matching result that is more in line with the overall semantics, rather than just a matching result of text similarity.

[0076] Step S2-2: Convert the masked text descriptions of each problem and the masked text description of the problem to be processed into corresponding text vectors respectively.

[0077] Step S2-3: Calculate the similarity between the text vector of the problem to be processed and the text vectors of each problem according to the similarity calculation algorithm.

[0078] Step S2-4: Select the top n problems with the highest similarity to the problem to be processed as similar problems according to the calculated similarity and the predetermined number n of similar problems to be selected.

[0079] Among them, for the similarity calculation algorithm, the classical cosine similarity is used in this embodiment to match text vectors. That is, two text problem features are vectorized, and the cosine value of the included angle between the two feature vectors is calculated. The specific calculation formula is as follows:

[0080]

[0081] In the formula, is the text vector of the problem that needs to be converted into an SQL statement (i.e., the problem to be processed), is the text vector of the problem recorded in the database, x i is the i-th dimension value of the vector and y i is the i-th dimension value of the vector and n is the dimension size of the vector. The top n problems that minimize cos(θ) in the above formula are found in the simulated database as similar problems. In this embodiment, n = 5.

[0082] Step S3: Use each of the integrated multiple large language models (LLMs) to separately generate SQL statements for the n similar problems, and score the generation results of each large language model respectively.

[0083] That is, for each large language model, based on the SQL statement generated by it for each similar problem and the correct SQL statement corresponding to the similar problem in the simulated database, it is judged whether the SQL statement generation result of the large language model for the n similar problems is correct, and a score is given. Generally speaking, it is the number of similar problems answered correctly.

[0084] In an alternative solution, scores are calculated by weighting different similar problems, and the calculation formula is as follows:

[0085]

[0086] In the formula, r i indicates whether the output SQL result of the model for the i-th similar problem is correct, corresponding to 0 or 1; t i is the weight corresponding to the i-th problem, and this weight can be directly given by the similarity between the problem and the problem to be processed.

[0087] In this embodiment, the four integrated large language models are SQLCoder, DeepSeek-22b, Llama3, and Qwen2. Preferably, SQLCoder is SQLCoder-7b, Llama3 is Llama3-8b, and Qwen2 is Qwen2-7b. In an alternative solution, the integrated multiple large language models can also be other large language models in the prior art.

[0088] Step S4: Select one large language model from the integrated multiple large language models based on the scores as the target large language model for answering the model to be processed.

[0089] Among them, select the large language model with the highest score to answer the question to be processed.

[0090] When the scores of two or more large language models are all the highest, select the model with the largest number of parameters among these models to answer the question to be processed. This is because the large language model with a large number of model parameters usually has a better effect in processing the Text2SQL task. Therefore, when the scores are the same, it is default that the model with a larger number of parameters is more likely to output the correct SQL statement under the same circumstances.

[0091] Step S5: Add some of the multiple similar questions to the prompt template of the target large language model to enable the target large language model to perform in-context learning.

[0092] Among them, after the two key links of question matching in step S3 and scoring in step S4, it has been determined which large language model is used to answer the question to be processed. Step S5 is the in-context learning link, and the purpose is to provide some detailed knowledge to the selected large language model to further improve the accuracy of the target large language model in answering specific questions to be processed.

[0093] Specifically, in step S5, add the database table structure involved in the question to be processed and at most m questions with similarity scores higher than a predetermined threshold in its similar questions to the prompt template of the target large language model. The similarity score is also the similarity value between the text descriptions of each similar question and the text description of the question to be processed calculated by the above similarity algorithm. In this embodiment, m = 5.

[0094] Among them, selecting questions with scores higher than a certain score is to control the quality of sample learning; while selecting at most 5 questions to limit the maximum number of sample questions is because the inventor observed that too many samples will "distract" the large language model, which will instead make the output result of the large language model deviate more from the correct result.

[0095] Step S6: Use the target large language model that has undergone in-context learning to answer the question to be processed, thereby generating the corresponding SQL statement.

[0096] The following Table 1 shows the corresponding pseudocode for the above steps:

[0097] Table 1 Pseudocode for the multi-LLM model integration inference method for the Text2SQL task

[0098]

[0099]

[0100] For example, as Figure 1 shown, the problem to be processed is "What are the names and capacities of the stadiums with the highest average attendance rate?".

[0101] Through the above step S2, the 5 similar problems selected from the simulated database according to the text description similarity are "What are the names of the players who got higher than the average score?" (similarity score 0.9), "What are the names of all the players who scored higher than the average score?" (similarity score 0.85), "What are the names of the buildings with the building height arranged in descending order?" (similarity score 0.8), "Find the names of the rooms with a price higher than the average price." (similarity score 0.75), "What's the name of the product with the highest price?" (similarity score 0.7).

[0102] Furthermore, through the above step S3, the integrated four large language models are used to answer these 5 similar problems respectively, generate the corresponding SQL statements, and compare them with the corresponding SQL statements in the simulated database, so as to score the generation results of each large language model. For example, the score of DeepSeek-22b is 4, the score of SQLCoder-7b is 3.3, the score of Llama3-8b is 3.2, and the score of Qwen2-7b is 2.3.

[0103] Then, in step S4, according to the scores of each model, the DeepSeek-22b with the highest score is selected as the target large language model to answer the problem to be processed.

[0104] Then, in step S5, 3 similar problems with similarity scores higher than a predetermined threshold (for example, 0.75) are selected from the 5 similar problems and added to the prompt template of DeepSeek-22b to provide reference examples for the model to perform context learning.

[0105] Finally, in step S6, the DeepSeek-22b after context learning is used to answer the problem to be processed and generate the corresponding SQL statement:

[0106] SELECT s.name,s.capacity FROM stadium ORDERED BY s.average DESC NULLSLAST LIMIT 1;

[0107] This SQL statement is the correct SQL statement corresponding to the problem to be processed.

[0108] <Comparative Example>

[0109] This comparative example provides multiple inference methods for the Text2SQL task based on a single existing model, which are used to compare the effects with the inference method for the Text2SQL task that integrates multiple large language models provided in the embodiment. The single existing models used for comparison are SQLCoder, DeepSeek-22b, Llama3, and Qwen2, which respectively correspond to the four large language models integrated in the method of the embodiment.

[0110] In this comparative example, the Text-to-SQL generation accuracies of these four single existing models and the integrated multi-model of the embodiment were respectively tested on 1034 questions, and the test results are summarized in Table 2 below:

[0111] Table 2 Accuracy data table of single existing models and integrated multi-LLM models

[0112]

[0113] The results of this comparative experiment show that the inference accuracy of the integrated inference scheme of multi-LLM models is much higher than that of single models. Regarding the test results of single models, there is no model with an accuracy higher than 0.700. However, in the integrated inference method of multi-LLM models in the embodiment, each LLM only answers the questions it is good at, and the overall accuracy of generating SQL can reach 0.734, far higher than the performance of each single model.

[0114] In addition, in this comparative example, a comparative experiment on the inference speed of the schemes of the four single models and the integrated multi-LLM model scheme of the embodiment was also carried out on an A100-80G graphics card, and the experimental results are summarized in Table 3 below:

[0115] Table 3 Inference speed data table of single existing models and integrated multi-LLM models

[0116]

[0117] The results of this comparative experiment show that the inference efficiency of the integrated inference of the integrated multi-LLM model in the embodiment is more than 3 times that of a single DeepSeek-22b model, and at the same time, it also ensures that the inference accuracy rate is much higher than that of the DeepSeek-22b model. Although small models such as SQLCoder, Llama3, and Qwen2 have great advantages in inference speed, the inference accuracy rates of these small models are generally low, significantly lower than the scheme of the embodiment. For example, the inference accuracy rate of the Qwen2 model is only 0.403.

[0118] Therefore, the integrated inference scheme of the integrated multi-LLM model in the embodiment not only ensures that the inference accuracy rate exceeds the schemes of all single models, but also ensures high efficiency in inference.

[0119] Functions and Effects of the Embodiment

[0120] Compared with the inference methods for the Text2SQL task in the prior art, the multi-LLM model integrated inference method for the Text2SQL task provided in this embodiment has the following advantages in multiple aspects:

[0121] First, the inference effect is better. Since the most suitable LLM model is selected to answer the question to be processed according to the performance of each LLM model on similar questions, that is, each LLM answers the question it is best at, and the few-shot prompting in prompt engineering is used to further improve the inference accuracy, the inference accuracy is high. In the comparative example, an accuracy comparison experiment was conducted between the integrated multi-LLM model inference scheme of the embodiment and the other four single-model inference schemes. The experimental results show that the scheme of the embodiment is far higher than the inference effects of all single-model schemes in terms of accuracy.

[0122] Second, the inference speed is fast. Although the LLM model with a large number of parameters has a high accuracy, due to its large number of parameters, it requires a lot of hardware resources for inference, so the speed will be very slow. On the contrary, the LLM model with a small number of parameters. The method of the present invention can well retain the respective advantages of the large model and the small model, so as to maintain a relatively fast inference speed while ensuring the inference accuracy. In the comparative example, an inference efficiency comparison experiment was also conducted between the integrated multi-LLM model inference scheme of the embodiment and the other four single-model inference schemes. The experimental results show that compared with the DeepSeek-22b model, the inference efficiency of the scheme of the embodiment is more than 3 times that of it; although the three small models of SQLCoder, Llama3, and Qwen2 have natural advantages in inference speed, their inference effects are very poor. The scheme of the embodiment can achieve both inference effect and inference efficiency.

[0123] Third, the resource consumption is less. Since the most suitable LLM model is selected according to the performance of each LLM model on similar questions, and a small number of high-quality similar questions are used to provide certain context information for it, without the need to fine-tune or retrain the LLM model, the computational resource consumption in this regard can be significantly reduced.

[0124] The above embodiments are only used to illustrate the specific implementation manners of the present invention, and the present invention is not limited to the description scope of the above embodiments. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-LLM model integrated reasoning method for Text2SQL tasks, characterized by: The following steps are involved: Step S1, establishing a simulation database according to an application scenario, wherein the simulation database includes text descriptions of multiple problems and corresponding correct SQL statements; Step S2, matching the text description of the problem to be processed with the text description of the problem in the simulation database, and selecting the first n problems with the highest similarity to the problem to be processed as similar problems; Step S3, using the integrated multiple large language models to generate SQL statements for the n similar questions respectively, and scoring the generated results of each of the large language models respectively; Step S4, selecting one of the integrated plurality of large language models based on the score as a target large language model for answering the model to be processed; Step S5, adding some of the multiple similar questions to the prompt word template of the target large language model, so that the target large language model performs context learning; Step S6, using the target large language model that has undergone context learning to answer the question to be processed, thereby generating a corresponding SQL statement.

2. According to claim 1, the multi-LLM model integrated reasoning method for the Text2SQL task, Features: Wherein, step S2 includes the following sub-steps: Step S2-1, respectively performing domain keyword masking on the text description of each of the problems in the simulation database and the text description of the problem to be processed; Step S2-2, converting the masked text description of each of the questions and the masked text description of the problem to be processed into corresponding text vectors respectively; Step S2-3, calculating the similarity between the text vector of the question to be processed and the text vectors of each of the questions according to a similarity calculation algorithm; Step S2-4, according to the calculated similarity and the predetermined number n of similar questions to be selected, select n questions with the highest similarity to the question to be processed as the similar questions.

3. The multi-LLM model integrated reasoning method for the Text2SQL task according to claim 2 is characterized in that: in, In step S2-1, the text description of the problem and the corresponding words in the text description of the problem to be processed are matched and masked based on preset keywords. In step S2-4, the text vectors are matched using cosine similarity: In the formula, is the question text vector to be processed, is the text vector of the question in the simulation database, x i is a vector The i-th dimension value, y i is a vector The i-th dimension value of , n is the dimension size of the vector, and the first n problems that minimize cos(θ) are found as the similar problems.

4. The multi-LLM model integrated reasoning method for the Text2SQL task according to claim 3 is characterized in that: in, In step S2-4, n=5.

5. The multi-LLM model integrated reasoning method for Text2SQL task according to claim 1, characterized in that: in, In step S3, the integrated multiple large language models include SQLCoder, DeepSeek-22b, Llama3, and Qwen2.

6. The multi-LLM model integrated reasoning method for Text2SQL task according to claim 1, characterized in that: in, In step S3, for each of the integrated large language models, whether the SQL statement generation result of the large language model on the n similar questions is correct is determined according to the corresponding correct SQL statement, and the number of the similar questions answered correctly is counted as the score.

7. The multi-LLM model integrated reasoning method for Text2SQL task according to claim 1, characterized in that: in, In step S3, for each of the integrated large language models, whether the SQL statement generation result of the large language model on the n similar questions is correct is judged according to the corresponding correct SQL statement, and the score is calculated by weighting different similar questions: In the formula, r i Is the correctness of the output SQL result of the model for the ith similar question, corresponding to 0 or 1; t i is the weight corresponding to the i-th question, which corresponds to the similarity between the question and the problem to be processed.

8. The multi-LLM model integrated reasoning method for Text2SQL task according to claim 1, characterized in that: in, In step S4, the large language model with the highest score is selected as the target large language model. When there are multiple large language models with the highest scores, the large language model with the largest number of parameters among the multiple large language models is selected as the target large language model.

9. The multi-LLM model integrated reasoning method for Text2SQL task according to claim 1, characterized in that: in, In step S5, the database table structure involved in the problem to be processed and at most m similar problems with similarity scores higher than a predetermined threshold are added to the prompt word template of the target large language model.

10. The multi-LLM model integrated reasoning method for Text2SQL task according to claim 9, characterized in that: in, In step S5, m=5.

Citation Information

Cited By

  • Vehicle network interaction data query method based on large language model

    CN120994683A