Test case self-correction Text-to-SQL method and system and medium
Generate diverse SQL test cases through a large language model and iteratively correct them, which solves the problem of difficult semantic errors in the text-to-SQL method, and improves the accuracy and execution success rate of SQL query statements.
Patent Information
- Application Number
- CN202510741829.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing text-to-SQL method is difficult to detect and correct when facing semantic errors, resulting in semantic errors in the generated SQL query statements, affecting the execution success rate and accuracy rate.
Design a test case self-correcting Text-to-SQL method, use large language models to generate diverse SQL test cases, and generate SQL query statements through iterative corrections, including context retrieval, relational pattern selection, test data and code segment generation, and combine predefined test cases for verification and optimization.
Improve the accuracy of generating SQL query statements, effectively detect and correct semantic errors, ensure that the generated SQL query statements meet the needs of the original problem, and avoid the impact on subsequent data analysis.
Smart Images

Figure CN120256320A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language technology, and in particular to a test case self-correction Text-to-SQL method, system and medium based on large model multi-agents. Background Art
[0002] The research field of text-to-SQL focuses on converting natural language questions into corresponding SQL query statements, which can help business personnel lacking programming skills directly obtain data related to data analysis problems from the database. Currently, common text-to-SQL methods mainly use technologies such as pre-trained language models and large language models, focusing on designing specific processes to improve the accuracy of SQL generation, while ignoring the design of solutions to detect errors existing in the generated SQL. When facing SQL query statements with semantic errors, since the model itself is difficult to detect incorrect semantic understanding deviations from the original data analysis problems, semantic errors always exist in the generated SQL query statements and cannot be corrected, negatively affecting the execution success rate and query accuracy of the SQL query statements. Summary of the Invention
[0003] The main purpose of the present invention is to provide a test case self-correction Text-to-SQL method, system and medium based on large model multi-agents, aiming to solve the problem that semantic errors difficult to correct exist in SQL query statements due to semantic understanding deviations of the original problems, and to improve the accuracy of the generated SQL query statements by providing as accurate test cases as possible to specifically test and check for possible errors in the SQL query statements.
[0004] To achieve the above object, the present invention proposes a test case self-correction Text-to-SQL method, and the method includes the following steps: Step S10, using a large language model to generate diverse SQL test cases; Step S20, analyzing the feedback information of the SQL test cases and generating an SQL query statement through iterative correction.
[0005] A further technical solution of the present invention is that the step S10 includes: Step S101, context retrieval: using a large language model to extract context information related to the SQL query statement to be generated from the database; Step S102, relationship mode selection: using a large language model to select the tables and fields most relevant to the current query from the database schema; Step S103, Test Data and Code Segment Generation: Based on the context information and the relationship pattern selection, generate test data and test code to construct SQL test cases; Among them, the step S101 includes: Using a large language model to extract key query words from the input natural language question, retrieving data content values similar to the key query words in the database content, and retrieving relevant information in the database description information in the vector database; The step S102 includes: Using the natural language question and knowledge through a large language model to identify the tables that the current query may involve, sorting the relevance of these tables, and finally selecting the tables related to the question to exclude irrelevant information; After determining the tables, further identify the fields related to the query and select them; The step S103 includes: Generate test data for verifying the execution result of the SQL statement according to the original data analysis problem and the database schema; Generate test code for obtaining the correct result according to the test data.
[0006] A further technical solution of the present invention is that the step of generating test code for obtaining the correct result according to the test data includes: Execute the generated code through a Python interpreter to verify its correctness. If the execution fails or the query result is empty, the system will adjust according to the feedback information and regenerate the code until the generated code can be executed correctly; If the execution is correct and the query result is obtained, then the query result will be used as the expected output data of the test case.
[0007] A further technical solution of the present invention is that the step S20 includes: Step S201, Generate a preliminary SQL query statement based on the SQL test case; Step S202, Combine the predefined test cases to verify the preliminary generated SQL query statement; Step S203, Optimize and correct the SQL query statement that has completed the execution verification according to the feedback information of the test case.
[0008] A further technical solution of the present invention is that the step S201 includes: Converting the natural language question 𝑄 into an SQL query 𝑌, which can retrieve relevant data from the database. The database can be represented as 𝐷 = ⟨𝐶, 𝑇⟩, where 𝐶 and 𝑇 respectively refer to the information of columns and tables. When processing complex database values, external knowledge 𝐾 must be combined to improve the model's understanding of the database values. The SQL generation stage can be expressed by the following formula, where the function Can represent a large language model with parameters 𝜃: 。
[0009] A further technical solution of the present invention is that the step S202 includes: comparing the preliminarily generated SQL query statement with the pre-defined test cases in the database management system. If the results are consistent, it indicates that the generated SQL query statement has passed the verification; if not, the feedback information is re-input into the model for further correction. Among them, the feedback information for SQL execution verification is divided into two categories: one is the situation where the SQLite execution fails due to SQL syntax errors. In this case, the error message of SQLite is fed back to the SQL correction stage, and the LLM makes targeted modifications. The other is that the SQL is successfully executed but the query result is inconsistent with the expected output of the test case. In this case, the incorrect query result and the expected correct result are fed back to the SQL correction stage together, and the LLM makes corresponding adjustments and corrections.
[0010] A further technical solution of the present invention is that the step S203 includes: correcting the possible errors in the SQL query statement by using the feedback information of the test cases to generate a SQL query statement that meets the expectations. Among them, the following function is used for SQL query statement correction: ; Among them, the function can represent a large language model with parameter 𝜃. The feedback information includes error messages, expected correct results, and existing SQL execution results, and can be expressed as 。
[0011] To achieve the above object, the present invention also proposes a test case self-correction Text-to-SQL system. The system includes a memory, a processor, and a test case self-correction Text-to-SQL program stored on the processor. When the test case self-correction Text-to-SQL program is run by the processor, it executes the steps of the method described above.
[0012] To achieve the above object, the present invention also proposes a computer-readable storage medium. The computer-readable storage medium stores a test case self-correction Text-to-SQL program. When the test case self-correction Text-to-SQL program is run by the processor, it executes the steps of the method described above.
[0013] The beneficial effects of the test case self-correction Text-to-SQL method, system, and medium of the present invention are: 1. The present invention can provide as accurate test cases as possible for targeted testing and checking of possible errors in SQL query statements, thereby improving the accuracy of generating SQL query statements; 2. The present invention can correct the imperceptible semantic errors in the SQL query statement through the design framework of "SQL generation - test case feedback - SQL correction", thereby avoiding the semantic understanding deviation of the original problem as much as possible and ensuring that the query result of the generated SQL query statement will not affect the subsequent data analysis link due to the inconsistency with the requirements of the original problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a schematic flowchart of a preferred embodiment of the test case self-correction Text-to-SQL method of the present invention; Figure 2 is an overall framework diagram of the test case self-correction Text-to-SQL method based on large model multi-agent; Figure 3 is a schematic diagram of the overall architecture for generating SQL test cases based on large language models; Figure 4 is a schematic diagram of the overall architecture for SQL generation of test cases; Figure 5 is a hardware architecture diagram of the test case self-correction Text-to-SQL system of the present invention.
[0015] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0017] The present invention proposes a test case self-correction Text-to-SQL method based on large model multi-agent. As Figure 1 shown, a preferred embodiment of the test case self-correction Text-to-SQL method of the present invention includes the following steps: Step S10, using a large language model to generate diverse SQL test cases.
[0018] The large language model (LLM) is abbreviated as the large model. It is a language model composed of artificial neural networks with many parameters (usually billions of weights or more). By using self-supervised learning or semi-supervised learning to train a large amount of unlabeled text, the large model can "remember" a large amount of facts during training and can capture most of the syntax and semantics of human language, and perform well in a wide range of tasks.
[0019] Step S20, analyzing the feedback information of the SQL test cases and generating SQL query statements through iterative correction.
[0020] Structured Query Language, abbreviated as SQL, is a programming language for a specific purpose, used to manage relational database management systems or perform stream processing in relational stream data management systems.
[0021] Text-to-SQL, also known as NL2SQL, refers to the conversion of natural language questions into corresponding SQL query statements.
[0022] Specifically, in this embodiment, the step S10 includes: Step S101, context retrieval: Using a large language model to extract context information related to the SQL query statement to be generated from the database.
[0023] A database is a computer software system that stores and manages data according to a data structure.
[0024] Step S102, relationship schema selection: Using a large language model to select the tables and fields most relevant to the current query from the database schema.
[0025] A database schema is a structure described in a formal language in a database system. It is a collection of objects and contains different types of objects in different types of databases. In the present invention, the database schema includes four parts: table name fields, column name fields, data content values, and database descriptions. Among them, the table name fields are the names of the data tables stored in the database; the column name fields record the column names included in each data table stored in the database; the data content values refer to the uniquely determined content values actually stored in the database; and the database description is a general description of the information recorded by the specific column name fields in the database.
[0026] Step S103, test data and code segment generation: Based on the context information and the relationship schema selection, generate test data and test code to construct SQL test cases.
[0027] A test case is a set of test inputs, execution conditions, and expected results prepared for a specific goal, used to verify whether a specific software requirement is met. The test cases designed for the Text-to-SQL task in the present invention mainly include two parts: test data and test code. Among them, the test data requires the LLM to generate code with a specific format according to the natural language question and the database schema, which is used to verify the execution result of the SQL statement; the test code requires the LLM to generate a Python code segment for obtaining the correct result according to the generated test data.
[0028] Among them, the step S101 includes: using a large language model to extract key query words from the input natural language question, retrieving data content values similar to the key query words in the database content, and retrieving relevant information in the database description information of the vector database; The step S102 includes: using the natural language question and knowledge through a large language model to identify the tables that the current query may involve, sorting the relevance of these tables, finally selecting the tables related to the question, and excluding irrelevant information; after determining the tables, the fields related to the query will be further identified and selected; The step S103 includes: Generating test data for verifying the execution result of the SQL statement according to the original data analysis problem and the database schema; Generating test code for obtaining the correct result according to the test data.
[0029] In this embodiment, the step of generating test code for obtaining the correct result according to the test data includes: Executing the generated code through a Python interpreter to verify its correctness. If the execution goes wrong or the query result is empty, the system will adjust according to the feedback information and regenerate the code until the generated code can be executed correctly; if the execution is correct and the query result is obtained, then the query result will be used as the expected output data of the test case.
[0030] Furthermore, in this embodiment, the step S20 includes: Step S201, generating a preliminary SQL query statement based on the SQL test case; Step S202, verifying the preliminary generated SQL query statement by combining predefined test cases; Step S203, optimizing and correcting the SQL query statement that has completed the execution verification according to the feedback information of the test case.
[0031] Among them, the step S201 specifically includes: converting the natural language question 𝑄 into an SQL query 𝑌, which can retrieve relevant data from the database. The database can be represented as 𝐷 = ⟨𝐶, 𝑇⟩, where 𝐶 and 𝑇 refer to column and table information respectively. When processing complex database values, external knowledge 𝐾 must be combined to improve the model's understanding of the database values. The SQL generation stage can be expressed by the following formula, where the function can represent a large language model with parameters 𝜃: .
[0032] The specific steps of S202 include: comparing the preliminarily generated SQL query statement with the predefined test cases in the database management system. If the results are consistent, it indicates that the generated SQL query statement has passed the verification; if not, the feedback information is re-input into the model for further correction. Among them, the feedback information for SQL execution verification is divided into two categories: one is the situation where the SQLite execution fails due to SQL syntax errors. In this case, the error message of SQLite is fed back to the SQL correction stage, and the LLM makes targeted modifications. The other is that the SQL is successfully executed but the query result is inconsistent with the expected output of the test case. In this case, the incorrect query result and the expected correct result are both fed back to the SQL correction stage, and the LLM makes corresponding adjustments and corrections.
[0033] The specific steps of S203 include: correcting the possible errors in the SQL query statement by using the feedback information of the test cases to generate a SQL query statement that meets the expectations. Among them, the following function is used for SQL query statement correction: ; Among them, the function represents a large language model with parameter 𝜃. The feedback information includes error messages, expected correct results, and existing SQL execution results, which can be expressed as .
[0034] The self-correcting Text-to-SQL method of test cases of the present invention is further elaborated in detail below.
[0035] Aiming at the problem that it is difficult to correct the semantic errors in the SQL query statement due to the semantic understanding deviation of the original problem, the present invention innovatively proposes a self-correcting Text-to-SQL method of test cases based on a large model multi-agent. It not only designs an interactive framework of three large model agents to generate reliable test cases including test data and test code, but also constructs an iterative framework of "SQL generation - test case feedback - SQL correction" to detect and correct the semantic errors existing in the SQL query statement. Figure 2Shows the overall implementation process of the present invention, mainly including the following two modules: SQL test case generation based on large language models and SQL generation based on test cases. The part of SQL test case generation based on large language models utilizes the powerful natural language processing ability of large language models to automatically generate diverse SQL test cases. The core of this part lies in how to interpret natural language descriptions through large language models and convert them into effective SQL test cases, so as to cover various SQL query scenarios that may occur. The part of SQL generation based on test cases utilizes the test cases generated in the previous stage to drive the generation of SQL queries. The core of this part lies in how to deeply analyze the feedback information of test cases and generate more accurate and efficient SQL queries through iterative correction.
[0036] 1. SQL test case generation based on large language models The overall design framework of the method for generating SQL test cases based on large language models is as Figure 3 shown, mainly including three stages: context retrieval, relationship schema selection, and test data and code segment generation.
[0037] The main goal of the context retrieval stage is to extract context information related to the SQL statement to be generated. The data content values (values) in the database will be indexed using the Locality Sensitive Hashing (LSH) algorithm, while the database description information (i.e., column descriptions, column aliases, value descriptions, etc.) will be constructed into a vector database. The large language model first extracts key query words from the input natural language question, and these keywords will serve as the basis for subsequent query generation. Using the keywords extracted by the LLM in the question, relevant values similar to the keywords are retrieved from the database content, and relevant information is retrieved from the database description information in the vector database.
[0038] The goal of the relationship schema selection stage is to select the tables and fields most relevant to the current query from the database schema. The large language model uses natural language questions and knowledge to identify the tables that may be involved in the current query, ranks the relevance of these tables, and finally selects the tables related to the question, excluding irrelevant information; after determining the tables, the fields related to the query will be further identified and selected. This step simplifies the complex database schema, reduces irrelevant information, and improves the accuracy of code generation.
[0039] In the test data and code segment generation phase, three large model intelligent agent interaction frameworks are designed to generate test cases. Among them, the data generation intelligent agent generates test data (test data) for verifying the execution results of SQL statements based on the original data analysis problem and the database schema. The code generation intelligent agent and the code inspection intelligent agent generate test code (test code) for obtaining the correct results based on the test data. Based on the selected context and relationship patterns mentioned above, the test data and preliminary test code generated through multi-agent interaction in this phase will be used to construct SQL test cases. The generated code is then executed through a Python interpreter to verify its correctness. If an error occurs during execution or the query result is empty, the system will adjust according to the feedback information and regenerate the code until the generated code can be executed correctly; if the execution is error-free and the query result is obtained, then the query result will be used as the expected output data of the test case.
[0040] Through the above process, the system will be able to generate SQL test cases with high accuracy, which can effectively improve the accuracy of subsequent SQL generation tasks.
[0041] 2. SQL Generation Based on Test Cases The overall design framework of the SQL generation method based on test cases is as Figure 4 shown, mainly including three stages: SQL generation, execution verification, and SQL correction.
[0042] First, preliminary SQL generation is performed through a large language model. The model receives the question, database relationship schema, and relevant knowledge as inputs and generates a preliminary SQL query statement. The goal of this stage is to quickly generate a possibly correct SQL statement for verification and correction in subsequent stages. Specifically, in the SQL generation stage, the natural language question 𝑄 (Question) is converted into an SQL query 𝑌, which can retrieve relevant data from the database. The database (Database) can be represented as 𝐷 = ⟨𝐶, 𝑇⟩. Where 𝐶 and 𝑇 refer to the information of columns (Column) and tables (Table) respectively. When dealing with complex database values, external knowledge 𝐾 (Knowledge) must be combined to improve the model's understanding of database values. The SQL generation stage can be expressed by the following formula, where the function can represent a large language model with parameters 𝜃: .
[0043] Then, the generated SQL statement enters the execution verification phase. In this phase, the SQL statement will be executed in a database management system (such as SQLite) and its correctness will be verified by comparing it with pre-defined test cases. The query result after executing the SQL will be compared with the expected output of the test case. If the results are consistent, it indicates that the generated SQL has passed the verification; if not, the feedback information will be re-input into the model for further correction. The feedback information for SQL execution verification is divided into two categories: one is the situation where the SQLite execution fails due to SQL syntax errors. In this case, the execution module will feedback the error message of SQLite to the SQL correction phase, and the LLM will make targeted modifications. The other is that the SQL is successfully executed but the query result is inconsistent with the expected output of the test case. In this case, the execution module will feedback the incorrect query result and the expected correct result to the SQL correction phase, and the LLM will make corresponding adjustments and corrections.
[0044] Next, the SQL statement that has completed the execution verification will enter the correction phase. The goal of this phase is to optimize and correct the generated SQL based on the feedback information from the test cases. To improve the accuracy of the SQL, the model will use the feedback information to correct possible errors in the SQL, thus generating an SQL query statement that meets the expectations. SQL correction can be described by the following function representation: .
[0045] Among them, the function can represent a large language model with parameter 𝜃. The feedback information (Feedback) includes error messages, expected correct results, and existing SQL execution results, and can be expressed as . By processing this information, the model can generate a more accurate SQL query statement to pass the subsequent verification phase.
[0046] Through continuous iteration of generation, correction, and verification, the entire process finally generates an SQL statement that meets the requirements and can pass all test cases.
[0047] The present invention is applicable to real business scenarios for data analysis problems. In this scenario, data analysis business personnel need to write SQL query statements for natural language problems in data analysis to obtain stored content related to the data analysis problem from the database for subsequent data analysis tasks. The key point of the present invention is to detect and self-correct semantic errors in SQL query statements caused by semantic understanding deviations in the original problem through test cases.
[0048] The SQL test case generation method based on large language models proposed by the present invention designs three large model agent interaction frameworks to generate test cases including test data and test code based on the context semantic learning ability and multi-agent interaction and iteration ability of large models. Among them, the data generation agent generates test data (test data) for verifying the execution result of SQL statements according to the original data analysis problem and database schema. The code generation agent and the code inspection agent generate test code (test code) for obtaining correct results according to the test data through multiple interactions of verification and correction. This method does not need to construct a complete generation process framework from the problem to the SQL query statement, but checks the possible errors in the SQL query statement by constructing as accurate test cases as possible, so as to improve the accuracy of generating SQL query statements.
[0049] The SQL generation method based on test cases proposed by the present invention enables the large model to generate a version of SQL query statement for the original problem according to the test cases (test data and corresponding test code) obtained in the previous link. Then, by executing the SQL and the test code on the test data respectively, comparing the feedback results, and iteratively correcting the SQL in combination with the test code logic. This method uses an iterative framework of "SQL generation - test case feedback - SQL correction" to check and correct the subtle semantic errors in the SQL query statement, so as to adjust the semantic understanding deviation of the original problem and improve the execution accuracy of the generated SQL query statement.
[0050] The beneficial effects of the self-correcting Text-to-SQL method of test cases of the present invention are: Related technologies involving text-to-SQL (Text-to-SQL) convert problems into SQL query statements that can run successfully in the database and obtain the results required by the problems by analyzing natural language problems and database-related information. Such technologies can improve the efficiency of writing SQL and enable non-professional field personnel without basic programming skills to directly search for data related to natural language problems in the database. These data contain potential information required to answer natural language questions. By analyzing these data, it can help relevant personnel adjust strategies and design plans in a timely manner according to the characteristics, states or situations reflected by the data, and give play to the advantages brought by the data value.
[0051] Existing methods for text-to-SQL mainly include text-to-SQL using pre-trained language models and text-to-SQL using large language models. The former, while training a deep learning model to parse natural language questions and generate SQL, further obtains relevant information in natural language questions and databases in combination with technologies such as neural networks. Although such methods have achieved good results in realizing the conversion of natural language questions to SQL, they ignore the fact that the generated SQL query statements are prone to errors that need to be corrected, because the end-to-end pre-trained model can only output the SQL statement corresponding to the problem according to the requirements of the model training stage. The text-to-SQL method using large language has achieved results far exceeding other methods in the text-to-SQL task based on the super context learning ability and knowledge emergence ability of the large language model. However, such methods focus more on how to design processes to improve the accuracy of SQL generation, and ignore the design scheme to detect errors in the generated SQL. This is different from the present invention, which does not require training and directly uses large models to build test cases to self-correct the SQL query statement results generated by the large model.
[0052] The present invention proposes a method for generating SQL test cases based on a large language model, which can provide test cases that are as accurate as possible for targeted testing and checking of possible errors in SQL query statements, thereby improving the accuracy of generating SQL query statements.
[0053] The present invention proposes a test case-based SQL generation method, which can correct subtle semantic errors in SQL query statements through the design framework of "SQL generation-test case feedback-SQL correction", thereby avoiding semantic understanding deviations of the original problem as much as possible and ensuring that the query results of the generated SQL query statements will not affect subsequent data analysis links due to inconsistency with the requirements of the original problem.
[0054] To achieve the above object, the present invention also proposes a test case self-correcting Text-to-SQL system, such as Figure 5As shown, the system includes a processor 1001, a CPU, a network interface 1004, a user interface 1003, a memory 1005, a communication bus 1002, and a test case self-correcting Text-to-SQL program stored on the processor. Among them, the communication bus 1002 is used to implement the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0055] Those skilled in the art can understand that Figure 5 the system structure shown in does not constitute a limitation on the system, and it may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0056] As Figure 5 shown, the memory 1005, as a computer storage medium, may include an operating device, a network communication module, a user interface module, and a test case self-correcting Text-to-SQL program.
[0057] In Figure 5 the system shown, the network interface 1004 is mainly used to connect to a network server and communicate with the network server for data; the user interface 1003 is mainly used to interact with a user terminal and receive instructions input by the user; and the processor 1001 may be used to call the test case self-correcting Text-to-SQL program stored in the memory 1005.
[0058] To achieve the above object, the present invention also proposes a computer-readable storage medium storing a test case self-correcting Text-to-SQL program. When the test case self-correcting Text-to-SQL program is run by a processor, it executes the steps of the method described above, which will not be elaborated here.
[0059] The above are only the preferred embodiments of the present invention, and do not limit the scope of the present invention. Any equivalent structural or process transformation made using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the scope of the invention protection of the present invention by the same token.
Claims
1. A test case self-correcting Text-to-SQL method, characterized in that, The method includes the following steps: Step S10, generating diverse SQL test cases using a large language model; Step S20, analyzing the feedback information of the SQL test cases and generating SQL query statements through iterative correction.
2. The self-correcting Text-to-SQL method for test cases according to claim 1, wherein The step S10 includes: Step S101, context retrieval: using a large language model to extract context information related to the SQL query statement to be generated from the database; Step S102, relationship schema selection: using a large language model to select the tables and fields most relevant to the current query from the database schema; Step S103, test data and code segment generation: generating test data and test code based on the context information and the relationship schema selection to construct SQL test cases; Among them, the step S101 includes: using a large language model to extract key query words from the input natural language question, retrieving data content values similar to the key query words in the database content, and retrieving relevant information in the database description information in the vector database; The step S102 includes: using a large language model to utilize natural language questions and knowledge to identify the tables that the current query may involve, sorting the relevance of these tables, finally selecting the tables related to the question, and excluding irrelevant information; after determining the tables, further identifying the fields related to the query and selecting them; The step S103 includes: Generating test data for verifying the execution result of the SQL statement according to the original data analysis problem and the database schema; Generating test code for obtaining the correct result according to the test data.
3. The test case self-correcting Text-to-SQL method according to claim 2, wherein The step of generating test code for obtaining the correct result according to the test data includes: Executing the generated code through a Python interpreter to verify its correctness. If the execution fails or the query result is empty, the system will adjust according to the feedback information and regenerate the code until the generated code can be executed correctly; if the execution is correct and a query result is obtained, then the query result will be used as the expected output data of the test case.
4. The test case self-correcting Text-to-SQL method according to claim 3, characterized in that, The step S20 includes: Step S201, generating a preliminary SQL query statement based on the SQL test cases; Step S202, performing verification on the preliminarily generated SQL query statement by combining predefined test cases; Step S203, optimizing and correcting the SQL query statement that has completed the execution verification according to the feedback information of the test cases.
5. The test case self-correcting Text-to-SQL method according to claim 4, wherein The step S201 includes: converting the natural language question 𝑄 into an SQL query 𝑌, which can retrieve relevant data from the database. The database can be represented as 𝐷 = ⟨𝐶, 𝑇⟩, where 𝐶 and 𝑇 refer to the information of columns and tables respectively. When dealing with complex database values, external knowledge 𝐾 must be incorporated to improve the model's understanding of database values. The SQL generation stage can be expressed by the following formula, where the function can represent a large language model with parameters 𝜃: 。 6. The test case self-correcting Text-to-SQL method according to claim 5, wherein, The step S202 includes: comparing the preliminarily generated SQL query statement with predefined test cases in a database management system. If the results are consistent, it indicates that the generated SQL query statement has passed the verification; if they are inconsistent, the feedback information is re-input into the model for further correction. Among them, the feedback information for SQL execution verification is divided into two categories: one is the situation where the SQLite execution fails due to SQL syntax errors. In this case, the error message of SQLite is fed back to the SQL correction stage, and the LLM makes targeted modifications. The other is that the SQL is successfully executed but the query result is inconsistent with the expected output of the test case. In this case, the incorrect query result and the expected correct result are both fed back to the SQL correction stage, and the LLM makes corresponding adjustments and corrections.
7. The test case self-correcting Text-to-SQL method according to claim 6, characterized in that, The step S203 includes: correcting possible errors in the SQL query statement by using the feedback information of the test case to generate a SQL query statement that meets the expectations. Among them, the following function is used for correcting the SQL query statement: ; Among them, the function can represent a large language model with parameter 𝜃. The feedback information includes error messages, expected correct results, and existing SQL execution results, and can be expressed as .
8. A test case self-correcting Text-to-SQL system, characterized in that, The system includes a memory, a processor, and a test case self-correcting Text-to-SQL program stored on the processor. When the test case self-correcting Text-to-SQL program is run by the processor, it executes the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a test case self-correcting Text-to-SQL program. When the test case self-correcting Text-to-SQL program is run by a processor, it executes the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Large language model SQL (Structured Query Language) generation method based on high-quality context sample and self-correction
CN118568127A
Intelligent SQL query generation method and system based on large language model
CN118861081A
Method for correcting nonstandard SQL (Structured Query Language) statement based on large language model
CN119003563A
Cypher query statement generation optimization method, device and system based on large language model
CN119807232A
Cited By
SQL statement generation method and device and electronic equipment
CN121070970A
SQL (Structured Query Language) conversion method, system and equipment based on large model and storage medium
CN121277962A
Text-to-structured query statement generation method and system based on large language model
CN122332419A