Method and device for implementing text2sql model based on thinking chain, computer equipment and readable storage medium

Through the text2sql model implementation method based on thinking chain, multiple open source large language models and target large models are used for evaluation and fine-tuning training, the traditional text-to-SQL conversion method is solved in terms of accuracy and flexibility, and efficient and accurate SQL statement generation is achieved.

CN119938697AInactive Publication Date: 2025-05-06DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510004468.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional text-to-SQL conversion methods have insufficient accuracy and flexibility, and it is difficult to efficiently handle complex and changeable user instructions, and cannot make full use of large-scale data and advanced language models, resulting in the low quality of generated SQL statements.

Method used

The text2sql model implementation method based on thinking chain is adopted, and RAG search and build prompt words are obtained by obtaining sample SQL query instruction text, multiple open source large language models and target large models are called for evaluation and fine-tuning training, and high-quality SQL statements are generated.

Benefits of technology

It realizes efficient and accurate text-to-SQL conversion, can handle complex and changeable user instructions, make full use of large-scale data and advanced language models, and the quality of generated SQL statements is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938697A_ABST
    Figure CN119938697A_ABST
Patent Text Reader

Abstract

The invention discloses a thinking chain-based text2sql model implementation method and device, computer equipment and a readable storage medium, and the method comprises the steps: firstly obtaining a sample SQL query instruction text, carrying out RAG retrieval, constructing cue words, calling at least two open source large language models for output, carrying out the evaluation of a target large model to obtain a step-by-step and reward training data set, and carrying out the calculation of the step-by-step and reward training data set; and performing fine tuning training on the initial step-by-step and reward model. And after the target instruction input by the user is obtained, reasoning is performed in combination with the trained step-by-step and reward model to obtain the corresponding SQL statement, so that efficient and accurate text-to-SQL conversion is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model applications, and in particular to a text2sql model implementation method, device, computer equipment and readable storage medium based on thought chain. Background Art

[0002] With the rapid growth of data volume and the increasing complexity of user data query requirements, traditional text-to-SQL conversion methods are insufficient in accuracy and flexibility. Existing technologies are difficult to efficiently process complex and changeable user instructions, and cannot fully utilize large-scale data and advanced language models, resulting in low-quality generated SQL statements that cannot meet actual application needs. Summary of the invention

[0003] The object of the present invention is to provide a method, device, computer equipment and readable storage medium for implementing a text2sql model based on thought chain.

[0004] In a first aspect, an embodiment of the present invention provides a method for implementing a text2sql model based on a thought chain, comprising:

[0005] Obtaining a sample SQL query instruction text, performing a RAG search on the sample SQL query instruction text to obtain a RAG search result, and constructing a prompt word based on the sample SQL query instruction text and the RAG search result;

[0006] Calling at least two pre-configured open source large language models to perform an output operation on the prompt word, and obtaining output results corresponding to the at least two open source large language models respectively;

[0007] Calling a pre-configured target large model to evaluate the output results corresponding to the at least two open source large language models respectively, to obtain a step-by-step training data set and a reward training data set;

[0008] Fine-tune the initial step-by-step model based on the step-by-step training data set, and fine-tune the initial reward model based on the reward training data set to obtain a trained step-by-step model and reward model;

[0009] The target SQL query instruction text input by the user is obtained, and the model reasoning based on the thinking chain is performed on the target SQL query instruction text in combination with the step-by-step model and the reward model to obtain the SQL statement corresponding to the target SQL query instruction text.

[0010] In a possible implementation, the acquiring of the sample SQL query instruction text, performing RAG search on the sample SQL query instruction text to obtain a RAG search result, and constructing a prompt word based on the sample SQL query instruction text and the RAG search result includes:

[0011] Get the sample SQL query command text;

[0012] Performing a table structure search on the sample SQL query instruction text, and based on the similarity of the table structure descriptions, recalling a preset number of table structure description information related to the sample SQL query instruction text;

[0013] Performing SQL query retrieval on the sample SQL query instruction text, and recalling a preset number of SQL query statements related to the sample SQL query instruction text based on the similarity of historical SQL query statements;

[0014] The sample SQL query instruction text, the preset number of table structure description information and the preset number of SQL query statements are integrated to construct the prompt word.

[0015] In a possible implementation, calling at least two pre-configured open source large language models to perform an output operation on the prompt word to obtain output results corresponding to the at least two open source large language models respectively includes:

[0016] At least two pre-configured open source large language models are called to perform an output operation on the prompt word, and counter data, step content data, self-reflection data, final answer data, and solution evaluation data corresponding to the at least two open source large language models are obtained as the output results.

[0017] In a possible implementation, the calling of the pre-configured target large model to evaluate the output results corresponding to the at least two open source large language models respectively to obtain a step-by-step training data set and a reward training data set includes:

[0018] Calling a pre-configured target large model to evaluate the step content data, the self-reflection data, the final answer data, and the solution evaluation data respectively corresponding to the at least two open source large language models to obtain a full evaluation result;

[0019] Constructing the step-by-step training data set from the full evaluation results based on the prompt word content and the step-by-step output results with correct evaluation representation;

[0020] The reward training data set is constructed from the full evaluation results based on the prompt word content, the content output at each step, and the evaluation results of the content output at each step.

[0021] In a possible implementation, fine-tuning the initial step-by-step model based on the step-by-step training data set, and fine-tuning the initial reward model based on the reward training data set to obtain a trained step-by-step model and reward model, includes:

[0022] Based on the step-by-step training data set, the initial step-by-step model is fine-tuned using LoRA to obtain a trained step-by-step model;

[0023] The initial reward model is fine-tuned based on the reward training data set to obtain a trained reward model.

[0024] In a possible implementation, the step of obtaining a target SQL query instruction text input by a user, and performing a model reasoning based on a thinking chain on the target SQL query instruction text in combination with the step-by-step model and the reward model to obtain a SQL statement corresponding to the target SQL query instruction text includes:

[0025] Get the target SQL query instruction text entered by the user;

[0026] Performing a table structure search on the target SQL query instruction text, and based on the similarity of the table structure descriptions, recalling a preset number of table structure description information related to the target SQL query instruction text;

[0027] Performing SQL query retrieval on the target SQL query instruction text, and recalling a preset number of SQL query statements related to the target SQL query instruction text based on the similarity of historical SQL query statements;

[0028] Integrate the target SQL query instruction text, the preset number of table structure description information and the preset number of SQL query statements to construct a target prompt word;

[0029] Inputting the target prompt word into the step-by-step model for reasoning to obtain preliminary model reasoning content;

[0030] Inputting the preliminary model reasoning content into the reward model for verification to obtain a verification result;

[0031] In the case where the verification result is characterized as an error, the target prompt word and the error information are fed back to the step-by-step model, and the step of inputting the target prompt word into the step-by-step model for reasoning to obtain the preliminary model reasoning content is re-executed;

[0032] When the verification result is characterized as correct, SQL verification is performed on the preliminary model inference content, and when the verification passes, the preliminary model inference content is output as an SQL statement corresponding to the target SQL query instruction text.

[0033] In a possible implementation, the performing SQL verification on the preliminary model reasoning content includes:

[0034] Determining whether the preliminary model inference content complies with SQL grammar rules;

[0035] If yes, it is determined that the preliminary model reasoning content verification has passed;

[0036] If not, it is determined that the preliminary model reasoning content verification fails;

[0037] In the case where the preliminary model reasoning content verification fails, the target prompt word and error information are fed back to the step-by-step model, and the step of inputting the target prompt word into the step-by-step model for reasoning to obtain the preliminary model reasoning content is re-executed.

[0038] In a second aspect, an embodiment of the present invention provides a text2sql model implementation device based on thought chain, comprising:

[0039] An acquisition module is used to acquire a sample SQL query instruction text, perform RAG search on the sample SQL query instruction text to obtain a RAG search result, and construct a prompt word based on the sample SQL query instruction text and the RAG search result; call at least two pre-configured open source large language models to perform an output operation on the prompt word to obtain output results corresponding to the at least two open source large language models; call a pre-configured target large model to evaluate the output results corresponding to the at least two open source large language models to obtain a step-by-step training data set and a reward training data set; fine-tune the initial step-by-step model based on the step-by-step training data set, and fine-tune the initial reward model based on the reward training data set to obtain a trained step-by-step model and a reward model;

[0040] The implementation module is used to obtain the target SQL query instruction text input by the user, and perform model reasoning based on the thinking chain on the target SQL query instruction text in combination with the step-by-step model and the reward model to obtain the SQL statement corresponding to the target SQL query instruction text.

[0041] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor and a non-volatile memory storing computer instructions, wherein when the computer instructions are executed by the processor, the computer device executes the method described in the first aspect.

[0042] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, wherein the readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method described in the first aspect.

[0043] Compared with the prior art, the beneficial effects provided by the present invention include: adopting a text2sql model implementation method, device, computer equipment and readable storage medium based on thought chain disclosed in the present invention, by obtaining sample SQL query instruction text, searching and constructing prompt words through RAG, calling at least two open source large language models for output, and then using the target large model for evaluation to obtain step-by-step and reward training data sets, and fine-tuning the initial step-by-step and reward models. After obtaining the target instruction input by the user, reasoning is performed in combination with the trained step-by-step and reward models to obtain the corresponding SQL statement, thereby realizing efficient and accurate text-to-SQL conversion. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative work.

[0045] Figure 1 A schematic diagram of the steps of a method for implementing a text2sql model based on a thought chain provided in an embodiment of the present invention;

[0046] Figure 2 A schematic structural diagram of a text2sql model implementation device based on thought chain provided in an embodiment of the present invention;

[0047] Figure 3 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0049] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0050] In order to solve the technical problems in the aforementioned background technology, Figure 1 The present invention provides a flowchart of a method for implementing a text2sql model based on a thought chain according to an embodiment of the present invention. The method for implementing a text2sql model based on a thought chain is introduced in detail below.

[0051] Step S201, obtaining a sample SQL query instruction text, performing a RAG search on the sample SQL query instruction text to obtain a RAG search result, and constructing a prompt word based on the sample SQL query instruction text and the RAG search result;

[0052] Step S202, calling at least two pre-configured open source large language models to perform an output operation on the prompt word, and obtaining output results corresponding to the at least two open source large language models respectively;

[0053] Step S203, calling the pre-configured target large model to evaluate the output results corresponding to the at least two open source large language models respectively, to obtain a step-by-step training data set and a reward training data set;

[0054] Step S204, fine-tuning the initial step-by-step model based on the step-by-step training data set, and fine-tuning the initial reward model based on the reward training data set, to obtain a trained step-by-step model and reward model;

[0055] Step S205, obtaining the target SQL query instruction text input by the user, combining the step-by-step model and the reward model to perform model reasoning based on the thinking chain on the target SQL query instruction text, and obtaining the SQL statement corresponding to the target SQL query instruction text.

[0056] In the embodiment of the present invention, for example, it is assumed that our server is processing a database query task about sales data.

[0057] First, the server obtains a sample SQL query instruction text, for example: "Query the top 10 products in sales in East China last month". After receiving this instruction, the server immediately performs a RAG search on it. In terms of table structure retrieval, the server calculates the similarity algorithm and recalls the five table structure descriptions that may be involved from the numerous table structure descriptions in the database, such as "product sales table", "region table", "timetable", etc. In terms of SQL query retrieval, the server identifies and recalls three SQL query statements similar to the current instruction based on the similarity of historical query statements, such as "Query the top five products in sales in North China last quarter" and "Query the products with the highest sales in the western region this month".

[0058] Based on the sample SQL query command text and RAG search results, the server starts to build prompt words, which include detailed instruction instructions, possible table structure descriptions, similar historical query statements, and strict requirements for output formats.

[0059] Next, the server calls at least two pre-configured open source large language models, such as qwen2-72b-instruct and llama-3.1-70B-instruct, to perform output operations on the constructed prompt words. The output of the qwen2-72b-instruct model may be: <count> [Starting Budget]< / count> <step> First determine the time range of the previous month< / step> <count> [Remaining budget]< / count> <step> Filter records from East China from the region table< / step> ...... <answer> The final top 10 product information is [specific product list]< / answer> <reflection> The whole reasoning process is relatively clear and the steps are reasonable< / reflection> The output of the llama-3.1-70B-instruct model might be: <count> [Starting Budget]< / count> <step> First get the sales data of all products< / step> <count> [Remaining budget]< / count> <step> Filter data from East China by region< / step> ...... <answer> The top 10 products are [specific product list]< / answer> <reflection> Some steps can be further optimized to improve efficiency< / reflection> .

[0060] Then, the server calls the pre-configured target large model, such as the closed-source kimi model, to evaluate the output results of the two open-source large language models. The kimi model will carefully analyze the rationality of each step, the accuracy of the final answer, and the effectiveness of self-reflection. After evaluation, the step-by-step output results and related information that are deemed correct and effective will be organized into a step-by-step training data set, such as detailed prompt words, reasonable step-by-step steps, accurate final answers, etc. At the same time, the prompt words, the content of each step output, and the evaluation results of each step output content are organized into a reward training data set.

[0061] Based on the generated step-by-step training dataset, the server fine-tunes the initial step-by-step model. For example, during the training process, the model will learn how to more accurately determine the time range, more efficiently filter regional data, and other steps. By continuously adjusting the model's parameters, it can better respond to similar query instructions. Similarly, the initial reward model is fine-tuned based on the reward training dataset, allowing the reward model to more accurately evaluate the quality of each step of the step-by-step model output.

[0062] When a user enters a target SQL query command text, such as "Query the five product categories with the fastest sales growth in South China in the first half of this year", the server first receives this command. Then, like the previous sample command, it performs a RAG search, recalls the relevant table structure and historical query statements, and constructs a prompt word. Then, the prompt word is input into the step-by-step model, and the step-by-step model outputs the reasoning results step by step, such as:<count> [Starting Budget]< / count> <step> Determine the time period for the first half of this year< / step> <count> [Remaining budget]< / count> <step> Get relevant data from the product category table< / step> ....... Afterwards, the reward model verifies these inference results. If it is found that one of the steps is inaccurate, such as the determination of the time interval is incorrect, the server will feedback the error information to the step-by-step model and ask it to re-infer. If the verification passes, the server will submit the generated SQL statement to the database for verification. If the verification finds a syntax error, the server will ask the step-by-step model to re-infer again until the generated SQL statement passes the verification, and finally obtains accurate and reliable query results and returns them to the user.

[0063] Through this entire process, the server can efficiently and accurately process various complex SQL query instructions and provide users with satisfactory services.

[0064] In an embodiment of the present invention, the obtaining of sample SQL query instruction text, performing RAG search on the sample SQL query instruction text to obtain RAG search results, and constructing prompt words based on the sample SQL query instruction text and the RAG search results can be implemented through the following examples.

[0065] Get the sample SQL query command text;

[0066] Performing a table structure search on the sample SQL query instruction text, and based on the similarity of the table structure descriptions, recalling a preset number of table structure description information related to the sample SQL query instruction text;

[0067] Performing SQL query retrieval on the sample SQL query instruction text, and recalling a preset number of SQL query statements related to the sample SQL query instruction text based on the similarity of historical SQL query statements;

[0068] The sample SQL query instruction text, the preset number of table structure description information and the preset number of SQL query statements are integrated to construct the prompt word.

[0069] In the embodiment of the present invention, the server is running an advanced database query system. At a certain moment, the server obtains a sample SQL query instruction text: "Find the list of sales personnel with the best sales performance in first-tier cities in the past year."

[0070] First, the server starts to retrieve the table structure of this sample SQL query instruction text. It uses a complex similarity algorithm to deeply analyze the numerous table structure descriptions stored in the database. In this process, the server will match and compare the key elements in the sample instruction, such as "sales performance", "sales staff", "first-tier cities", "past year", etc., with each table structure description. After some careful calculations, the server successfully recalled 8 table structure description information that are considered to be highly relevant to the instruction. These 8 tables are "sales performance table", "sales staff information table", "city classification table", "time range table", etc. These table structure descriptions provide detailed information about data storage and organization, laying the foundation for subsequent query operations.

[0071] Next, the server continues to perform SQL query retrieval on the sample SQL query instruction text. It screens from a huge library of historical SQL query statements, and based on the similarities between historical statements and current sample instructions in terms of semantics, data requirements, and logical structure, the server eventually recalls 5 SQL query statements related to the current instruction. For example, "Find a list of salespeople with outstanding sales performance in second-tier cities in the first half of the year," "Query the top three salespeople in new first-tier cities in the past six months," and so on.

[0072] After completing the above two key retrieval steps, the server enters the stage of constructing prompt words. The server takes the obtained sample SQL query instruction text "Find the list of salespeople with the best sales performance in first-tier cities in the past year" as the core part, and then cleverly integrates the recalled 8 table structure description information and 5 related SQL query statements. The constructed prompt words may be as follows:

[0073] You are a database expert and need to generate accurate query results based on the following instructions.

[0074] Please follow the instructions below carefully:

[0075] (1) Read the given question carefully and <count> and< / count> The counter resets between

[0076] (2) Generate a detailed, logical, step-by-step solution.

[0077] (3) Use each step of your solution <step> and< / step> Surrounded by labels.

[0078] (4) You can use a maximum of {budget} steps (starting budget) by <count> and< / count> The label counts down to keep track of it, and when it reaches 0 it stops generating more steps, you don't have to use them all.

[0079] (5) When you are unsure how to proceed, reflect on yourself and, based on your self-reflection and rewards, decide whether you need to go back to the previous steps.

[0080] (6) After completing the solution steps, reorganize and synthesize the steps into the final answer. <answer> and< / answer> Given in the label.

[0081] (7) <reflection> and< / reflection> Within the label, conduct a critical, honest, and subjective self-assessment of your reasoning process.

[0082] The table structure used:

[0083] “Sales Performance Table”: Contains fields such as salesperson ID, sales amount, sales time, etc.

[0084] "Sales Personnel Information Table": Contains fields such as sales personnel ID, name, and city. ......

[0086] "Time Range Table": Contains fields such as time interval identifier, start time, and end time.

[0087] Query record history that may be related to user instructions:

[0088] “Find a list of salespeople with outstanding sales performance in second-tier cities in the first half of the year”

[0089] “Query the top three salespeople in the new first-tier cities in terms of sales performance in the past six months” ......

[0091] User query command:

[0092] “Find the list of salespeople with the best sales performance in first-tier cities in the past year”

[0093] The prompt words constructed in this way provide a comprehensive and accurate information basis for the subsequent call of the open source large language model for output operations, ensuring that the model can fully understand the task requirements and generate high-quality output results. After the server prepares this prompt word, it can pass it to the pre-configured open source large language model to start the next processing flow.

[0094] In the embodiment of the present invention, the calling of at least two pre-configured open source large language models to perform an output operation on the prompt word to obtain output results corresponding to the at least two open source large language models respectively can be implemented through the following examples.

[0095] At least two pre-configured open source large language models are called to perform an output operation on the prompt word, and counter data, step content data, self-reflection data, final answer data, and solution evaluation data corresponding to the at least two open source large language models are obtained as the output results.

[0096] In the embodiment of the present invention, illustratively, the server receives the constructed prompt word and prepares to call at least two pre-configured open source large language models to perform an output operation.

[0097] Assume that the pre-configured open source large language models are qwen2-72b-instruct and llama-3.1-70B-instruct. The server first sends a prompt word to qwen2-72b-instruct to start the output operation.

[0098] After receiving the prompt word, qwen2-72b-instruct begins in-depth analysis and reasoning. It first generates counter data, such as <count>

[10] < / count> , indicating that the initial step budget is 10 steps.

[0099] Next, it generates step content data step by step. The first step may be <step> Filter the time period in the past year from the time range table< / step> ,Then <count> [9]< / count> , indicating that there are 9 budget steps left. The second step may be <step> Extract sales personnel information in first-tier cities from the sales personnel information table< / step> , <count> [8]< / count> .

[0100] During the generation process, qwen2-72b-instruct will also reflect on itself. For example <reflection> The current steps are well planned and can effectively obtain the required data< / reflection> .

[0101] After a series of reasoning and calculations, the final answer data is finally generated <answer> The following is a list of salespeople with the best sales performance in first-tier cities over the past year: [Specific list of people]< / answer> .

[0102] At the same time, the entire solution will be evaluated to generate solution evaluation data <reflection> The entire solution is efficient and accurate, successfully meeting the user's query needs< / reflection> .

[0103] After the server obtains the output result of qwen2-72b-instruct, it immediately sends the same prompt word to llama-3.1-70B-instruct.

[0104] llama-3.1-70B-instruct also follows the same logic to generate corresponding data. Its counter data may be <count> [8]< / count> , step content data such as <step> First associate the sales performance table and salesperson information table< / step> , <count> [7]< / count> , self-reflection data such as <reflection> This step will help to obtain data more accurately in the future.< / reflection> , the final answer data <answer> The list of salespeople with the best sales performance in first-tier cities in the past year is: [another list of specific personnel]< / answer>, Solution Evaluation Data <reflection> The solution was executed smoothly, but the algorithm can be further optimized to improve efficiency when associating data.< / reflection> .

[0105] The server successfully obtained the counter data, step content data, self-reflection data, final answer data, and solution evaluation data corresponding to the two open source large language models, qwen2-72b-instruct and llama-3.1-70B-instruct, as output results. These rich and detailed data will provide important basis and materials for subsequent evaluation and model training.

[0106] Through such a complex and precise processing process, the server fully utilizes the advantages of different open source large language models, laying a solid foundation for generating high-quality database query results. In subsequent processing, the server will further analyze and optimize these output results to continuously improve the system's performance and service quality.

[0107] In an embodiment of the present invention, the calling of the pre-configured target large model to evaluate the output results corresponding to the at least two open source large language models respectively to obtain a step-by-step training data set and a reward training data set can be implemented through the following example.

[0108] Calling a pre-configured target large model to evaluate the step content data, the self-reflection data, the final answer data, and the solution evaluation data respectively corresponding to the at least two open source large language models to obtain a full evaluation result;

[0109] Constructing the step-by-step training data set from the full evaluation results based on the prompt word content and the step-by-step output results with correct evaluation representation;

[0110] The reward training data set is constructed from the full evaluation results based on the prompt word content, the content output at each step, and the evaluation results of the content output at each step.

[0111] In the embodiment of the present invention, the server is working intensively and orderly. At this time, the server has obtained the output results corresponding to at least two open source large language models (such as qwen2-72b-instruct and llama-3.1-70B-instruct), including counter data, step content data, self-reflection data, final answer data, and solution evaluation data.

[0112] The server calls the pre-configured target large model (such as the closed-source kimi model) to evaluate these output results. First, the kimi model begins to evaluate the step content data output by qwen2-72b-instruct. It carefully analyzes the logic and accuracy of each step, such as checking whether the step of "filtering the time period of the past year from the time range table" correctly extracts the required time range, and whether the connection between subsequent steps is reasonable. At the same time, the self-reflection data is also evaluated to determine whether the model's reflection on its own steps is accurate and in-depth. For the final answer data, the kimi model will compare it with the real data in the database to confirm the correctness of the answer. Finally, the solution evaluation data is comprehensively considered to evaluate the integrity and effectiveness of the entire solution.

[0113] After completing the evaluation of qwen2-72b-instruct, the kimi model then performed the same rigorous and detailed evaluation of the output results of llama-3.1-70B-instruct.

[0114] After comprehensive evaluation, the Kimi model obtained the full evaluation results. The server constructed the step-by-step training dataset and the reward training dataset from the full evaluation results.

[0115] When constructing a step-by-step training dataset, the server will filter out the step-by-step output results that are evaluated as correct. For example, the series of steps output by qwen2-72b-instruct, "first extract qualified data from the sales performance table, and then associate it with the salesperson information table", was evaluated by the kimi model as correct and efficient. The server will organize the prompt word content and the corresponding correct step-by-step output results together to form part of the step-by-step training dataset. Similarly, the step-by-step outputs evaluated as correct in llama-3.1-70B-instruct will also be included in the step-by-step training dataset in the same way.

[0116] When constructing the reward training data set, the server will take more factors into consideration. It will integrate the prompt word content, the specific content of each step output and the corresponding evaluation results. For example, although the content output by qwen2-72b-instruct in a certain step is correct, it is not efficient, which is mentioned in the evaluation results. The server will put this prompt word, the output content of this step and the evaluation results into the reward training data set. In this way, the reward training data set can contain richer information, helping the model to better learn how to optimize the output of each step in subsequent training.

[0117] Through such a rigorous and meticulous processing process, the server successfully constructed a step-by-step training data set and a reward training data set, providing high-quality data support for subsequent model fine-tuning training, thereby continuously improving the performance and accuracy of the model to better meet the complex needs of users.

[0118] In an embodiment of the present invention, the initial step-by-step model is fine-tuned based on the step-by-step training data set, and the initial reward model is fine-tuned based on the reward training data set to obtain a trained step-by-step model and a reward model, which can be implemented through the following examples.

[0119] Based on the step-by-step training data set, the initial step-by-step model is fine-tuned using LoRA to obtain a trained step-by-step model;

[0120] The initial reward model is fine-tuned based on the reward training data set to obtain a trained reward model.

[0121] In the embodiment of the present invention, exemplarily, the server is preparing to perform fine-tuning training on the initial step-by-step model and the initial reward model based on the previously constructed step-by-step training data set and the reward training data set.

[0122] First, the server starts fine-tuning the initial step-by-step model using LoRA (Low-Rank Adaptation) based on the step-by-step training dataset. The server inputs a large amount of sample data from the step-by-step training dataset into the initial step-by-step model. These data contain detailed prompt words and step-by-step output results that are evaluated as correct and effective.

[0123] During the training process, the model parameters are adjusted according to the input data. For example, for a step-by-step task involving sales data analysis, the model may learn how to more accurately determine the conditions for data screening, how to more efficiently perform data aggregation and grouping operations, etc. LoRA technology enables the model to adjust only a small number of parameters while maintaining the main structure of the original large model, thereby greatly reducing the computational cost and time of training.

[0124] The server continuously monitors the progress of training and judges the effect of training based on indicators such as the training loss function value and accuracy. After multiple iterations and optimizations, the initial step-by-step model gradually mastered how to generate accurate and reasonable step-by-step solutions based on different query instructions, and finally obtained a trained step-by-step model.

[0125] Next, the server fine-tunes the initial reward model based on the reward training dataset, which contains the prompt words, the output content of each step, and the evaluation results of the output content of each step.

[0126] The server inputs this data into the initial reward model, and the model learns from this data to understand what kind of step-by-step output is excellent and what kind is insufficient. For example, for a specific query task, the model will learn that if the output of a certain step can accurately extract key information and conform to the overall logical framework, then this step should receive a higher reward evaluation; conversely, if the output of a certain step is wrong or unreasonable, it should be given a lower evaluation.

[0127] During the supervised fine-tuning process, the server also pays close attention to the training indicators and results. By continuously adjusting the model parameters, the initial reward model can gradually accurately evaluate each step generated by the step-by-step model, and finally obtain the trained reward model.

[0128] Through such a careful fine-tuning training process, the server successfully obtained a performance-optimized step-by-step model and reward model. These two models can work better together to provide more accurate and efficient services for processing complex database query tasks, improving user experience and satisfaction.

[0129] In an embodiment of the present invention, the target SQL query instruction text input by the user is obtained, and the target SQL query instruction text is subjected to model reasoning based on a thinking chain in combination with the step-by-step model and the reward model to obtain a SQL statement corresponding to the target SQL query instruction text, which can be implemented through the following example.

[0130] Get the target SQL query instruction text entered by the user;

[0131] Performing a table structure search on the target SQL query instruction text, and based on the similarity of the table structure descriptions, recalling a preset number of table structure description information related to the target SQL query instruction text;

[0132] Performing SQL query retrieval on the target SQL query instruction text, and recalling a preset number of SQL query statements related to the target SQL query instruction text based on the similarity of historical SQL query statements;

[0133] Integrate the target SQL query instruction text, the preset number of table structure description information and the preset number of SQL query statements to construct a target prompt word;

[0134] Inputting the target prompt word into the step-by-step model for reasoning to obtain preliminary model reasoning content;

[0135] Inputting the preliminary model reasoning content into the reward model for verification to obtain a verification result;

[0136] In the case where the verification result is characterized as an error, the target prompt word and the error information are fed back to the step-by-step model, and the step of inputting the target prompt word into the step-by-step model for reasoning to obtain the preliminary model reasoning content is re-executed;

[0137] When the verification result is characterized as correct, SQL verification is performed on the preliminary model inference content, and when the verification passes, the preliminary model inference content is output as an SQL statement corresponding to the target SQL query instruction text.

[0138] In the embodiment of the present invention, the server is illustratively ready to provide efficient and accurate services to users at any time. At this time, the server obtains the target SQL query instruction text entered by the user: "Count the total sales of different product categories in each quarter of this year."

[0139] The server then retrieved the table structure of the target SQL query instruction text. It used a sophisticated similarity algorithm to carefully screen the numerous table structure descriptions in the database. After some calculations and comparisons, based on the similarity of the table structure descriptions, it successfully recalled 8 table structure descriptions that were considered closely related to the instruction, such as "sales data table", "product category table", "quarterly timetable", etc. These table structure descriptions provide key architectural information for subsequent query operations.

[0140] Next, the server performs SQL query retrieval on the target SQL query instruction text. It shuttles through the huge historical SQL query statement library and successfully recalls 10 related SQL query statements based on the similarity between the historical statements and the current target instruction. For example, "Count the total sales of different product categories in each month of last year" and "Calculate the total sales of a specific product category in the first half of the year".

[0141] Then, the server integrates the target SQL query instruction text, the recalled 8 table structure description information and 10 SQL query statements, and carefully constructs the target prompt word. This target prompt word contains rich and detailed information, providing comprehensive guidance for subsequent model reasoning.

[0142] After the target prompt is constructed, the server inputs it into the step-by-step model for reasoning. The step-by-step model begins to work and generates preliminary model reasoning content, which may be: <count>

[10] < / count> <step> First, determine the four quarters of this year from the quarterly schedule< / step> <count> [9]< / count> <step> Filter the sales records of the corresponding product categories in the sales data table< / step> ...... <answer> The following are the total sales statistics of different product categories in each quarter of this year [Specific data]< / answer> .

[0143] The server then inputs the preliminary model reasoning content into the reward model for verification. The reward model will carefully review the rationality and accuracy of each step, as well as the correctness of the final answer. If the verification result is characterized as an error, such as finding that the range of a quarter is incorrectly determined, the server will feed back the target prompt word and this error message to the step-by-step model. After receiving the feedback, the step-by-step model will re-infer and generate new preliminary model reasoning content.

[0144] If the validation result of the reward model is correct, the server will perform SQL verification on the preliminary model inference content. During the verification process, the server will strictly check whether the generated SQL statement conforms to the grammatical rules and whether it can be correctly executed in the database. If the verification passes, the server will output this preliminary model inference content as the SQL statement corresponding to the target SQL query instruction text to the user to meet the user's query needs.

[0145] Through such a rigorous and orderly processing flow, the server can provide users with high-quality, accurate SQL query results, improving user experience and trust in the system.

[0146] In the embodiment of the present invention, the SQL verification of the preliminary model reasoning content can be implemented through the following example.

[0147] Determining whether the preliminary model inference content complies with SQL grammar rules;

[0148] If yes, it is determined that the preliminary model reasoning content verification has passed;

[0149] If not, it is determined that the preliminary model reasoning content verification fails;

[0150] In the case where the preliminary model reasoning content verification fails, the target prompt word and error information are fed back to the step-by-step model, and the step of inputting the target prompt word into the step-by-step model for reasoning to obtain the preliminary model reasoning content is re-executed.

[0151] In the embodiment of the present invention, for example, the server is processing a complex query task proposed by a user. Assume that the target SQL query instruction text entered by the user is: "Query the sales quantity and average price of each commodity last month."

[0152] The server performs SQL verification on the preliminary model inference content obtained through step-by-step model and reward model inference. First, the server analyzes the structure and syntax of this preliminary model inference content in detail. It checks whether the keywords in the statement are used correctly, whether the table names and field names are spelled accurately, whether the use of operators and functions complies with the specifications, and whether the overall logic of the statement is clear.

[0153] Assume that the initial model inference content is: "SELECT SUM(sales_quantity),AVG(price)FROMproducts WHERE month=LAST_MONTH();" The server will judge each element one by one. It will confirm that the keywords such as "SELECT", "FROM", and "WHERE" are used appropriately. Then, it will check whether the "products" table name exists in the database and whether the "sales_quantity" and "price" field names are correct. For the "LAST_MONTH()" function, the server will confirm that it is supported and used correctly in the current database environment.

[0154] If after a comprehensive check, the server finds that all syntax elements conform to the SQL syntax rules, then it will determine that the preliminary model inference content has passed the verification. Next, you can execute query operations in the database according to this inference content to obtain the data required by the user.

[0155] However, if during the check, the server finds that there is a situation that does not comply with SQL syntax rules. For example, if it finds that the "LAST_MONTH()" function is not supported in the current database, or the table name "products" is misspelled as "product", then the server will determine that the preliminary model inference content verification has failed.

[0156] Once the verification fails, the server will feed back the target SQL query instruction text entered by the user, the relevant table structure description information, and the target prompt word of the historical query statement, together with the error information, to the step-by-step model. After receiving these feedbacks, the step-by-step model will re-infer.

[0157] For example, the step-by-step model may readjust the reasoning strategy and generate new preliminary model reasoning content, such as: "SELECT SUM(sales_quantity),AVG(price)FROM product_sales WHERE month=DATE_SUB(CURRENT_DATE,INTERVAL 1MONTH);" Then, the server performs SQL verification on the newly generated preliminary model reasoning content again until the verification passes.

[0158] Through such a rigorous and detailed SQL verification process, the server can ensure the grammatical correctness of the query statement finally executed, thereby improving the accuracy and reliability of the query and providing users with better services.

[0159] In order to more clearly describe the solution provided by the embodiment of the present invention, a relatively complete implementation method is provided below.

[0160] In this technical solution, we propose three key links to achieve efficient model training and inference process. These three links include: training data set preparation, model training and model inference.

[0161] First, in the training dataset preparation stage, we used a variety of advanced models as the basis, including qwen2-72b-instruct, LLama-3.1-70B-instruct, and the closed-source kimi model. Through these models, we built an automated way to prepare datasets. This process not only improved the quality and diversity of the dataset, but also ensured that the training data had a wide coverage and could fully meet the needs of subsequent model training. The automated design effectively reduced manual intervention, improved the efficiency of data processing, and ensured that we could quickly iterate and generate a large amount of high-quality training data.

[0162] Next is the model training phase. At this stage, we made full use of the pre-made training datasets based on the open source large models qwen2-72b-instruct and llama-3-8B-sqlcoder. Combined with the large model fine-tuning technology, we successfully trained two models: a step-by-step model and a reward model. The design of the step-by-step model allows the model to decompose complex problems into a series of smaller and more manageable steps, which helps improve the accuracy and efficiency of problem solving. The reward model is used to evaluate the execution results of each step and provide feedback to guide the step-by-step model towards the optimal solution.

[0163] Finally, in the model reasoning stage, we combined the step-by-step model with the reward model and introduced the langchain technology to implement a multi-step circular thinking chain reasoning process. In the model reasoning stage, we used the langchain technology to build a multi-step circular thinking chain reasoning process. By combining the step-by-step model and the reward model, the system can perform complex reasoning processes, simulate the way humans think, and gradually build the optimal solution to the problem. This innovative reasoning mechanism not only enhances the intelligence of the model, but also improves its applicability in practical applications.

[0164] Specific steps:

[0165] Training data set creation:

[0166] In the training dataset preparation phase of this technical solution, we adopted a highly automated approach to generate training data for the step-by-step model and the reward model. The implementation steps of this phase are detailed below, covering the entire process from receiving user query instructions to generating the final dataset (see Figure 1 ).

[0167] 1. Receiving user query instructions

[0168] First, the system receives a query instruction from the user. For example, the user may enter: "Check how many major accidents have occurred in the Northeast this year?"

[0169] 2.RAG (Retrieval-Augmented Generation) retrieval

[0170] Next, we use the RAG method to recall content related to the user query from the vector database. This process consists of two key parts:

[0171] Table structure retrieval: Using the similarity algorithm of table structure description, the 8 table structure description information (DDL content) that may be involved are recalled from the database. This process calculates the similarity between the user query instruction and the table structure description to ensure that the recalled table can effectively support the subsequent SQL query construction.

[0172] SQL query retrieval: By analyzing the similarity of historical SQL query statements, the system can identify and recall 4 SQL query statements similar to the user's query. This step not only provides the context of the user's query, but also provides a reference for the model to generate appropriate SQL statements.

[0173] 3. Construction of prompt words

[0174] After obtaining relevant information about the user's query, the system integrates the user's query instructions, 8 table structure descriptions that may be used, and 4 SQL query statements similar to the user's instructions to construct detailed, step-by-step large language model prompts. These prompts will be passed to two large language models (qwen2-72b-instruct and llama-3.1-70B-instruct models) to perform corresponding queries and reasoning. The prompt template is as follows:

[0175] #You are a database expert who can provide a detailed, step-by-step solution based on the user's instructions.

[0176] ##Please follow the instructions below carefully:

[0177] (1) Read the given question carefully and <count> and< / count> The counter resets between

[0178] (2) Generate a detailed, logical, step-by-step solution.

[0179] (3) Use each step of your solution <step> and< / step> Surrounded by labels.

[0180] (4) You can use a maximum of {budget} steps (starting budget) by <count> and< / count> The label counts down to keep track of it, and when it reaches 0 it stops generating more steps, you don't have to use them all.

[0181] (5) When you are unsure how to proceed, reflect on yourself and, based on your self-reflection and rewards, decide whether you need to go back to the previous steps.

[0182] (6) After completing the solution steps, reorganize and synthesize the steps into the final answer. <answer> and< / answer> Given in the label.

[0183] (7) <reflection> and< / reflection> Within the label, conduct a critical, honest, and subjective self-assessment of your reasoning process.

[0184] ##Table structure used:

[0185] {{table_ddl_array}}

[0186] ##Query record history that may be related to user instructions:

[0187] {{history_query_example_array}}

[0188] ##User query command:

[0189] {{user_input}}

[0190] ##The results of the step-by-step output are:

[0191] 4. Execution and result return of large models

[0192] The system uses two different large models: qwen2-72b-instruct and llama-3.1-70B-instruct. The output is executed according to the prompt words generated in step 3. The results of each step include the following important parts:

[0193] Counter (count): Displays the current remaining step budget so that subsequent processing can be completed within the specified steps.

[0194] Step content (step): Describe each specific step of the solution in detail to ensure the transparency and traceability of the reasoning process.

[0195] Self-reflection: Evaluate the steps that have been executed and provide feedback on the effectiveness and rationality of the steps to facilitate subsequent improvements.

[0196] Final answer: Provide the final answer to the question in a preset format to ensure the standardization of the results.

[0197] Solution evaluation (reflection): Conduct a final self-evaluation of the entire solution, taking into account all aspects of the execution process and providing a basis for quality control of the training data.

[0198] 5. Result Evaluation and Training Data Generation

[0199] Finally, the closed-source large model Kimi was used to conduct a comprehensive evaluation of the "step content", "self-reflection", "final answer" and "solution evaluation" generated by the two models. This evaluation phase not only ensured the accuracy and rationality of the output results, but also generated training data sets for the step-by-step model and the reward model.

[0200] Step-by-step model training data: includes prompt word content and step-by-step output results that are evaluated as correct. Fine-tuning training on this part of the data set can improve the accuracy of the model's understanding of how to decompose complex tasks and gradually generate content.

[0201] Reward model training data: covers the prompt word content, the content of each step output, and the evaluation results of each step output content. Fine-tuning training with these data sets will greatly improve the ability of the reward model in evaluating the content generated by each step of the step-by-step model.

[0202] The training data format of the step-by-step model can be found in Table 1.

[0203] Table 1

[0204]

[0205]

[0206] The training data format of the reward model can be found in Table 2.

[0207] Table 2

[0208]

[0209]

[0210]

[0211]

[0212] Through the above automated training dataset preparation steps, we not only improved the efficiency and quality of data generation, but also laid a solid foundation for the next step of model training.

[0213] Model training:

[0214] During the model training phase, fine-tuning the step-by-step model and reward model is a key step to improve system performance. The step-by-step model is based on qwen2-72b-instruct and is fine-tuned using LoRA (Low-Rank Adaptation) technology through a carefully designed step-by-step training dataset. This method effectively optimizes the parameters of the model, enabling it to better understand and handle the step-by-step parsing of complex tasks.

[0215] At the same time, the reward model is based on llama3-8b-sqlcoder and uses reward training data for supervised fine-tuning (SFT). Through this process, the model can not only learn accurate outputs, but also adjust according to feedback in a dynamic environment.

[0216] During the fine-tuning phase, we paid special attention to the adjustment of key hyperparameters such as learning rate, batch size (batch_size), number of training rounds (epoch), and LoRA low-rank parameter (lora_rank). The optimization of these parameters not only ensures the stability and efficiency of the training process, but also improves the responsiveness and decision-making accuracy of the model in practical applications. Through meticulous tuning, we finally achieved a significant improvement in model accuracy.

[0217] Model Reasoning:

[0218] In this technical solution, the model-based thinking chain reasoning step aims to achieve efficient parsing and response to complex user queries through a series of systematic processing flows. The reasoning process mainly includes the following steps:

[0219] 1. Receiving user query instructions

[0220] First, the system receives the user's query instruction. For example, the user may enter: "Check how many major accidents have occurred in the Northeast this year?"

[0221] 2. Content Recall and Retrieval

[0222] Next, the system uses the RAG (Retrieval-Augmented Generation) method to recall content related to the user's query from the vector database. This process includes two important aspects: Table structure retrieval: The system uses the similarity algorithm of the table structure description to identify 8 tables that may be used. This step ensures that subsequent SQL queries can rely on accurate data models, enhancing the effectiveness and accuracy of queries. Historical SQL query retrieval: By analyzing the similarity of historical SQL query statements, the system can recall 10 SQL statements similar to user queries. This strategy enables the system to draw on existing query experience and further improve the accuracy and efficiency of generating SQL statements.

[0223] 3. Constructing prompt words for large language models

[0224] Using the user query instruction, 8 table structure descriptions that may be used, and 10 SQL query statements similar to the user instruction, the system constructs a detailed, step-by-step large language model prompt word. These prompt words will be sent to the step-by-step model for inference prediction, generating step-by-step execution thinking results. This step is the key link in converting natural language into structured queries.

[0225] 4. Reward model to verify reasoning results

[0226] The system verifies the reasoning results of the step-by-step model step by step through the reward model. This process ensures the accuracy and reliability of the reasoning results. If the verification finds an error, the system will add error information and reassign the task to the step-by-step model for reasoning to correct the error. If the verification is correct, the reasoning result will be passed to the next step for execution.

[0227] 5.SQL verification

[0228] After completing the reasoning verification, the system submits the final SQL statement generated by the large model to the database for execution verification. The main purpose of this link is to ensure that the generated SQL statement meets the grammatical requirements to avoid potential execution errors. If a grammatical error is found in the SQL statement during the verification process, the system will also record the error information and return to the step-by-step model for re-reasoning; if the verification is successful, it will continue to the final execution link. The strict verification at this stage further improves the reliability and stability of the system in actual applications.

[0229] Through the above steps, the system realizes step-by-step thinking chain loop reasoning, which not only improves the accuracy of large models in solving complex problems, but also enhances the robustness of the system and user satisfaction. This iterative and verification process ensures the high quality and reliability of the final result.

[0230] This technical solution combines the step-by-step model with the reward model and introduces the Langchain technology to realize a multi-step circular thinking chain reasoning process. This innovative method significantly improves the intelligence level of the system, enabling it to simulate the human thinking process when dealing with complex problems and gradually derive the optimal solution.

[0231] First, the combination of a step-by-step model and a reward model allows the system to perform self-evaluation and feedback during the reasoning process. The step-by-step model is responsible for breaking down the problem into smaller, processable parts, while the reward model ensures the quality of the output at each step. Through such collaboration, the system can not only generate answers efficiently, but also correct errors in real time and optimize the decision-making process. This dynamic feedback mechanism makes the system more flexible in practical applications and able to adapt to changing user needs.

[0232] Secondly, the introduction of Langchain technology enhances the coherence and logic of the reasoning process, allowing the output of each step to be seamlessly connected to form a complete chain of thought. This structured reasoning method not only improves the interpretability of the model, but also makes it easier for users to understand the decision-making process of the system, thereby increasing trust in the system.

[0233] In summary, the technical effect brought by this technical solution is significant. It not only improves the accuracy and efficiency of the model in complex reasoning scenarios, but also enhances the intelligence and applicability of the system, enabling it to provide reliable support in various practical applications.

[0234] Please refer to Figure 2 , Figure 2 A text2sql model implementation device 110 based on thought chain provided in an embodiment of the present invention includes:

[0235] The acquisition module 1101 is used to acquire a sample SQL query instruction text, perform RAG search on the sample SQL query instruction text to obtain a RAG search result, and construct a prompt word based on the sample SQL query instruction text and the RAG search result; call at least two pre-configured open source large language models to perform an output operation on the prompt word to obtain output results corresponding to the at least two open source large language models; call a pre-configured target large model to evaluate the output results corresponding to the at least two open source large language models to obtain a step-by-step training data set and a reward training data set; fine-tune the initial step-by-step model based on the step-by-step training data set, and fine-tune the initial reward model based on the reward training data set to obtain a trained step-by-step model and a reward model;

[0236] Implementation module 1102 is used to obtain the target SQL query instruction text input by the user, and perform model reasoning based on the thinking chain on the target SQL query instruction text in combination with the step-by-step model and the reward model to obtain the SQL statement corresponding to the target SQL query instruction text.

[0237] It should be noted that the implementation principle of the aforementioned text2sql model implementation device 110 based on the thinking chain can refer to the implementation principle of the aforementioned text2sql model implementation method based on the thinking chain, and will not be repeated here. It should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the text2sql model implementation device 110 based on the thinking chain can be a separately established processing element, or it can be integrated in a chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a processing element of the above device. The function of the text2sql model implementation device 110 based on the thinking chain. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in a processor element or an instruction in software form.

[0238] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more microprocessors (digital signal processors, DSP), or one or more field programmable gate arrays (FPGA), etc. For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0239] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned text2sql model implementation device 110 based on thought chain. Figure 3 As shown, Figure 3 The structure block diagram of the computer device 100 provided in the embodiment of the present invention. The computer device 100 includes a text2sql model implementation device 110 based on thought chain, a memory 111, a processor 112 and a communication unit 113.

[0240] In order to realize data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are electrically connected to each other directly or indirectly. For example, the electrical connection between these elements can be realized through one or more communication buses or signal lines. The text2sql model implementation device 110 based on the thinking chain includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the text2sql model implementation device 110 based on the thinking chain stored in the memory 111, such as the software function modules and computer programs included in the text2sql model implementation device 110 based on the thinking chain.

[0241] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is running, it controls the computer device where the readable storage medium is located to execute the aforementioned text2sql model implementation device 110 based on the thinking chain.

[0242] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.

Claims

1. A text2sql model implementation method based on thought chain, characterized in that: include: Obtaining a sample SQL query instruction text, performing a RAG search on the sample SQL query instruction text to obtain a RAG search result, and constructing a prompt word based on the sample SQL query instruction text and the RAG search result; Calling at least two pre-configured open source large language models to perform an output operation on the prompt word, and obtaining output results corresponding to the at least two open source large language models respectively; Calling a pre-configured target large model to evaluate the output results corresponding to the at least two open source large language models respectively, to obtain a step-by-step training data set and a reward training data set; Fine-tune the initial step-by-step model based on the step-by-step training data set, and fine-tune the initial reward model based on the reward training data set to obtain a trained step-by-step model and reward model; The target SQL query instruction text input by the user is obtained, and the model reasoning based on the thinking chain is performed on the target SQL query instruction text in combination with the step-by-step model and the reward model to obtain the SQL statement corresponding to the target SQL query instruction text.

2. The method according to claim 1, characterized in that The obtaining of the sample SQL query instruction text, performing RAG search on the sample SQL query instruction text to obtain a RAG search result, and constructing a prompt word based on the sample SQL query instruction text and the RAG search result, includes: Get the sample SQL query command text; Performing a table structure search on the sample SQL query instruction text, and based on the similarity of the table structure descriptions, recalling a preset number of table structure description information related to the sample SQL query instruction text; Performing SQL query retrieval on the sample SQL query instruction text, and recalling a preset number of SQL query statements related to the sample SQL query instruction text based on the similarity of historical SQL query statements; The sample SQL query instruction text, the preset number of table structure description information and the preset number of SQL query statements are integrated to construct the prompt word.

3. The method according to claim 1, characterized in that The calling of at least two pre-configured open source large language models to perform an output operation on the prompt word to obtain output results corresponding to the at least two open source large language models respectively includes: At least two pre-configured open source large language models are called to perform an output operation on the prompt word, and counter data, step content data, self-reflection data, final answer data, and solution evaluation data corresponding to the at least two open source large language models are obtained as the output results.

4. The method according to claim 3, characterized in that The calling of the pre-configured target large model to evaluate the output results corresponding to the at least two open source large language models respectively, to obtain a step-by-step training data set and a reward training data set, includes: Calling a pre-configured target large model to evaluate the step content data, the self-reflection data, the final answer data, and the solution evaluation data respectively corresponding to the at least two open source large language models to obtain a full evaluation result; Constructing the step-by-step training data set from the full evaluation results based on the prompt word content and the step-by-step output results with correct evaluation representation; The reward training data set is constructed from the full evaluation results based on the prompt word content, the content output at each step, and the evaluation results of the content output at each step.

5. The method according to claim 1, characterized in that The step of fine-tuning the initial step-by-step model based on the step-by-step training data set and fine-tuning the initial reward model based on the reward training data set to obtain a trained step-by-step model and reward model includes: Based on the step-by-step training data set, the initial step-by-step model is fine-tuned using LoRA to obtain a trained step-by-step model; The initial reward model is fine-tuned based on the reward training data set to obtain a trained reward model.

6. The method according to claim 1, characterized in that The step of obtaining the target SQL query instruction text input by the user, and performing model reasoning based on the thinking chain on the target SQL query instruction text in combination with the step-by-step model and the reward model to obtain the SQL statement corresponding to the target SQL query instruction text includes: Get the target SQL query instruction text entered by the user; Performing a table structure search on the target SQL query instruction text, and based on the similarity of the table structure descriptions, recalling a preset number of table structure description information related to the target SQL query instruction text; Performing SQL query retrieval on the target SQL query instruction text, and recalling a preset number of SQL query statements related to the target SQL query instruction text based on the similarity of historical SQL query statements; Integrate the target SQL query instruction text, the preset number of table structure description information and the preset number of SQL query statements to construct a target prompt word; Inputting the target prompt word into the step-by-step model for reasoning to obtain preliminary model reasoning content; Inputting the preliminary model reasoning content into the reward model for verification to obtain a verification result; In the case where the verification result is characterized as an error, the target prompt word and the error information are fed back to the step-by-step model, and the step of inputting the target prompt word into the step-by-step model for reasoning to obtain the preliminary model reasoning content is re-executed; When the verification result is characterized as correct, SQL verification is performed on the preliminary model inference content, and when the verification passes, the preliminary model inference content is output as an SQL statement corresponding to the target SQL query instruction text.

7. The method according to claim 6, characterized in that The SQL verification of the preliminary model reasoning content includes: Determining whether the preliminary model inference content complies with SQL grammar rules; If yes, it is determined that the preliminary model reasoning content verification has passed; If not, it is determined that the preliminary model reasoning content verification fails; In the case where the preliminary model reasoning content verification fails, the target prompt word and error information are fed back to the step-by-step model, and the step of inputting the target prompt word into the step-by-step model for reasoning to obtain the preliminary model reasoning content is re-executed.

8. A text2sql model implementation device based on thought chain, characterized in that: include: An acquisition module is used to acquire a sample SQL query instruction text, perform RAG search on the sample SQL query instruction text to obtain a RAG search result, and construct a prompt word based on the sample SQL query instruction text and the RAG search result; call at least two pre-configured open source large language models to perform an output operation on the prompt word to obtain output results corresponding to the at least two open source large language models; call a pre-configured target large model to evaluate the output results corresponding to the at least two open source large language models to obtain a step-by-step training data set and a reward training data set; Fine-tune the initial step-by-step model based on the step-by-step training data set, and fine-tune the initial reward model based on the reward training data set to obtain a trained step-by-step model and reward model; The implementation module is used to obtain the target SQL query instruction text input by the user, and perform model reasoning based on the thinking chain on the target SQL query instruction text in combination with the step-by-step model and the reward model to obtain the SQL statement corresponding to the target SQL query instruction text.

9. A computer device, characterized in that: The computer device comprises a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Communication large model construction method and device, equipment and storage medium

    CN117527608A

  • Content evaluation method, device and equipment for large model scene and storage medium

    CN117744664A

  • Method and device for reinforcement learning of large language model

    CN117808120A

  • Structured query language generation method and device, electronic equipment and storage medium

    CN118093621A

  • Problem instruction optimization processing method and device, equipment and storage medium

    CN118626877A

Cited By

  • Judgment document abstract generation method based on three-section type GRPO reinforcement learning

    CN120278126A

  • Text-to-SQL (Structured Query Language) full-link acquisition method

    CN120705250A

  • SQL statement generation model training method and device

    CN120744484A

  • A method and device for training an SQL statement generation model

    CN120744484B

  • Medical field dialogue method and device based on weak supervision and thinking chain, and readable medium

    CN120996180A