Automatic error detection and repair method and system for translation from natural language to SQL (Structured Query Language)
By generating candidate SQL queries and performing consistency verification using SQL equivalence verification and large language models, errors in the Text-to-SQL dataset are automatically detected and fixed, and the dataset quality problem is solved and the translation accuracy of the model is improved.
Patent Information
- Application Number
- CN202510313661.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
AI Technical Summary
There are a large number of error mapping relationships in the existing Text-to-SQL dataset, which affects the training effect and accuracy of the model and lacks automatic detection and repair methods.
By generating multiple candidate SQL queries, perform consistency verification using SQL equivalence verification and large language models, automatically detecting and fixing error mappings.
It realizes automated error detection and repair, improves the quality of the data set, and improves the translation accuracy of the Text-to-SQL model.
Smart Images

Figure CN120277093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing and databases. Specifically, it relates to a method and system for automatically detecting and repairing errors in natural language to SQL translation, and more specifically, to a method for automatically detecting and repairing errors in natural language to SQL translation based on execution consistency. Background Art
[0002] In the Internet era, databases are an important way to store and query data. Databases are not only used to store data, but also provide efficient data retrieval, update, and management functions, supporting enterprises and organizations in making data-driven decisions. SQL (Structured Query Language) plays an irreplaceable role in database management. As a standardized query language, SQL allows users to interact with databases in a concise and intuitive manner. Through SQL, users can perform various operations such as insert, delete, update, and query. Text-to-SQL is a technology based on natural language processing that automatically converts natural language queries into SQL queries. This technology not only allows non-technical users to directly interact with databases using natural language, but also reduces the burden on software developers in writing SQL.
[0003] The current mainstream Text-to-SQL technologies are mainly divided into two categories: one is fine-tuning based on pre-trained models, which performs supervised learning on a labeled Text-to-SQL dataset to change the model parameters to make them suitable for the natural language to SQL translation task; the other is prompt engineering based on large language models, which uses few-shot learning or zero-shot learning and in-context learning, and directly guides the model to generate SQL queries by designing prompts without additional fine-tuning. The former relies on high-quality datasets for model training optimization, and the latter improves the model performance through strategies such as examples and Chain-of-Thought. Both of these technologies rely on high-quality training data.
[0004] Existing Text-to-SQL datasets are mainly manually written and organized, inevitably containing some incorrect translations from natural language to SQL. Taking Spider and BIRD as examples, which are two important datasets in the Text-to-SQL field and are widely used for model training and evaluation. These two datasets are collected and written by professional teams, containing more than 10,000 mappings from natural language to SQL and covering multiple database instances, and are used by multiple models in Text-to-SQL research. However, even with a large amount of manpower spent on data collection, cleaning, writing, and error correction, these two datasets still contain many incorrect mappings between natural language and SQL, which affects the training effect and accuracy measurement of Text-to-SQL models.
[0005] Currently, there is no tool that can automatically detect and repair errors in Text-to-SQL datasets. Therefore, previous work has not delved deeply into the automatic detection and repair of Text-to-SQL translation errors and the research on dataset quality assurance, and there are still certain gaps. Summary of the Invention
[0006] Aiming at the defects in the prior art, the purpose of the present invention is to provide an automatic error detection and repair method and system for natural language to SQL translation.
[0007] An automatic error detection and repair method for natural language to SQL translation provided by the present invention includes:
[0008] Step S1: Generate multiple candidate SQL queries based on the natural language query, and construct an SQL set based on the multiple candidate SQL queries and the original SQL query;
[0009] Step S2: Use SQL equivalence verification technology to verify the equivalence of the SQL queries in the SQL set and generate counterexamples;
[0010] Step S3: Use a large language model to execute the natural language query on the generated counterexamples to obtain the first execution result;
[0011] Step S4: Use a database engine to execute the original SQL query on the generated counterexamples to obtain the second execution result;
[0012] Step S5: Compare the execution results of the natural language and the original SQL based on execution consistency. If the execution results are inconsistent on a certain counterexample, it is determined that an error is detected;
[0013] Step S6: Score each SQL query based on the comparison result;
[0014] Step S7: Select the candidate SQL with a score higher than the highest-scoring candidate SQL in the original SQL query as the repaired SQL.
[0015] Preferably, step S1 includes: using a base model to generate multiple candidate SQL queries for a natural language query, and adding the original SQL query thereto to construct an SQL set;
[0016] The base model is a large language model with the highest accuracy in using natural language queries among existing models;
[0017] The large language model generates multiple candidate SQL queries based on prompt engineering; wherein, the prompt includes a natural language query, a database schema definition, and optionally, an example;
[0018] The prompt further includes: requiring the large language model to generate N candidate SQLs with different syntactic structures and keywords, and restricting the number to be less than or equal to N while ensuring the diversity of candidate SQLs; and simultaneously filtering out ungrammatical queries among them in advance using a database execution engine.
[0019] Preferably, step S2 includes: using SQLSolver to filter equivalent SQL query pairs within a set timeout; if the SQL query pair is not equivalent, then using VeriEQL to generate a counterexample for the SQL query pair, and if the counterexample is successfully generated, retaining the counterexample and processing the next query pair, otherwise skipping the query pair.
[0020] Preferably, step S2 further includes: if VeriEQL fails to successfully generate a counterexample, generating a counterexample for the SQL query pair by means of database instance generation;
[0021] The method of generating a counterexample for the SQL query pair by means of database instance generation includes:
[0022] Using database fuzz testing technology to generate database instances for candidate SQL queries, where the database instances include the values of each table and each column, and each table contains several rows of data; if the execution results of the SQL query pair are inconsistent on a certain database instance, then the database instance is a counterexample.
[0023] Preferably, step S2 further includes: for all generated counterexamples, reducing the number of data rows included in the counterexamples by means of row-by-row deletion;
[0024] The method of reducing the number of data rows included in the counterexample by means of row-by-row deletion includes: starting from a table without foreign key constraints, deleting data rows row by row in a cascading manner, and if the execution results of the SQL query pair are still inconsistent after deleting a certain row of data, then the deletion is successful, otherwise restoring the row of data; repeating the above operation for each data row of each table until no data row can be deleted.
[0025] Preferably, step S3 includes: providing a prompt for the large language model, where the prompt includes a natural language query, database table names, column names, integrity constraints, data rows of counterexamples, a chain of thought guidance for executing a natural language query task on a database instance, and requirements for the structured output format; using a large language model to execute the prompt, and parsing and retaining the execution result.
[0026] Preferably, when comparing the execution results of the natural language and the original SQL, if the following conditions are simultaneously satisfied, it is considered that the execution result meets the execution consistency condition;
[0027] Condition 1: Columns sorted in any order are regarded as equivalent;
[0028] Condition 2: The row order depends on whether sorting is required in the natural language and the original SQL query;
[0029] Condition 3: The same data rows appear the same number of times in the result.
[0030] Preferably, step S7 includes: if there are multiple SQL queries with the same highest score, if the given original SQL query is the highest score, then select the original SQL query; otherwise, randomly select one SQL query from the SQL queries with the highest score as the final SQL.
[0031] An automatic error detection and repair system for natural language to SQL translation provided by the present invention includes:
[0032] Module M1: generating multiple candidate SQL queries based on a natural language query, and constructing an SQL set based on the multiple candidate SQL queries and the original SQL query;
[0033] Module M2: verifying the equivalence of the SQL queries in the SQL set using SQL equivalence verification technology, and generating counterexamples;
[0034] Module M3: using a large language model to execute a natural language query on the generated counterexamples to obtain a first execution result;
[0035] Module M4: using a database engine to execute the original SQL query on the generated counterexamples to obtain a second execution result;
[0036] Module M5: comparing the execution results of the natural language and the original SQL based on execution consistency, and if the execution results are inconsistent on a certain counterexample, it is determined that an error is detected;
[0037] Module M6: scoring each SQL query based on the comparison result;
[0038] Module M7: Select the candidate SQL with a score higher than the highest-scoring candidate SQL in the original SQL query as the repaired SQL.
[0039] Preferably, the module M1 includes: generating multiple candidate SQL queries for the natural language query using a base model, and adding the original SQL query thereto to construct an SQL set;
[0040] The base model is a large language model with the highest accuracy rate in using natural language queries among existing models;
[0041] The large language model generates multiple candidate SQL queries based on prompt engineering; wherein, the prompt includes a natural language query, a database schema definition, and optional examples;
[0042] The prompt further includes: requiring the large language model to generate N candidate SQLs with different syntactic structures and keywords, and restricting the number to be less than or equal to N while ensuring the diversity of candidate SQLs; at the same time, using the database execution engine in advance to filter out the queries that do not conform to the syntax;
[0043] The module M2 includes: using SQLSolver to filter equivalent SQL query pairs within a set timeout; if the SQL query pair is not equivalent, using VeriEQL to generate a counterexample for the SQL query pair, and if a counterexample is successfully generated, retaining the counterexample and processing the next query pair, otherwise skipping the query pair;
[0044] The module M2 further includes: if VeriEQL fails to successfully generate a counterexample, generating a counterexample for the SQL query pair by means of database instance generation;
[0045] The generating a counterexample for the SQL query pair by means of database instance generation includes:
[0046] Using database fuzz testing technology to generate database instances for candidate SQL queries, where the database instances include the values of each table and each column, and each table contains several rows of data; if the execution results of the SQL query pair are inconsistent on a certain database instance, then the database instance is a counterexample;
[0047] The module M3 includes: providing a prompt for the large language model, where the prompt includes a natural language query, database table names, column names, integrity constraints, the data rows of counterexamples, a thought chain guidance for executing a natural language query task on a database instance, and structured output format requirements; using a large language model to execute the prompt, and parsing and retaining the execution results;
[0048] When comparing the execution results of the natural language and the original SQL, if the following conditions are simultaneously met, it is considered that the execution results meet the execution consistency conditions;
[0049] Condition 1: Columns sorted in any order are considered equivalent;
[0050] Condition 2: The row order depends on whether sorting is required in the natural language and the original SQL query;
[0051] Condition 3: The same data rows appear the same number of times in the result.
[0052] Compared with the prior art, the present invention has the following beneficial effects:
[0053] 1. The present invention realizes automatic error detection and repair for Text-to-SQL translation, which can effectively reduce the burden of manually writing and reviewing Text-to-SQL datasets, and improve the quality of the dataset by automatically detecting and repairing incorrect translations in the dataset;
[0054] 2. The present invention is the first invention that clearly proposes execution consistency and applies it to automatic error detection and repair for Text-to-SQL translation; by converting the problem of judging whether a natural language query matches an SQL query into the problem of judging whether the execution results of the natural language and the SQL query are consistent on a specific database instance, the problem of the lack of formal semantics in natural language is solved; and the semantics of the natural language is visualized on a specific instance by executing the natural language on a large model in a specific instance;
[0055] 3. The present invention increases the possibility of having an SQL that matches the natural language in the candidate SQL query set by using a large language model to generate multiple candidate SQL queries with different syntactic structures and keywords, and can maximize the probability of repairing incorrect Text-to-SQL translations;
[0056] 4. The present invention can also improve the accuracy of the Text-to-SQL model translation in the actual application scenario by allowing the model to generate multiple candidate SQL queries and selecting the highest-scoring SQL in the Text-to-SQL application. Description of the Drawings
[0057] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes, and advantages of the present invention will become more obvious:
[0058] Figure 1 It is the overall flowchart for automatic error detection and repair of Text-to-SQL translation.
[0059] Figure 2 It is an example diagram of the specific implementation of each step in the automatic error detection and repair of Text-to-SQL translation.
[0060] Figure 3Partial screenshot of the execution process of the automatic error detection and repair method for natural language to SQL translation. Detailed implementation
[0061] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all belong to the protection scope of the present invention.
[0062] Embodiment 1
[0063] In view of the deficiencies in existing research, the purpose of the present invention is to provide an automatic error detection and repair method and system for natural language to SQL translation. It first proposes execution consistency, which is the correctness condition for the matching of natural language queries and SQL queries. Its core idea is to detect potential errors in the mapping by comparing the execution results of natural language queries and SQL queries on a given database instance. If the expected result of the natural language query is inconsistent with the actual result of the SQL query, it indicates that there is an error.
[0064] The automatic error detection and repair method for natural language to SQL translation includes:
[0065] Step 1: Generate candidate SQLs, generate multiple candidate SQL queries using a replaceable base model, and add the original SQL query to them.
[0066] Step 2: Counterexample generation, verify the equivalence of the SQL queries generated in Step 1 using SQL equivalence verification technology and generate counterexamples.
[0067] Step 3: Natural language query execution, use a large language model to execute the natural language query on the counterexamples generated in Step 2 and obtain the results.
[0068] Step 4: SQL query execution, execute the SQL query in the database engine and obtain the execution results on the counterexamples generated in Step 2.
[0069] Step 5: Error detection, based on the execution results of Steps 3 and 4, compare the execution results of the natural language and the original SQL based on execution consistency. If the execution results are inconsistent on a certain counterexample, it is determined that an error is detected and go to Step 6; otherwise, no error is detected and the process ends.
[0070] Step 6: SQL query scoring, score each SQL query based on the comparison results.
[0071] Step 7: SQL query repair. Select the candidate SQL with a score higher than the highest-scoring candidate SQL in the original SQL in Step 6 as the repaired SQL.
[0072] Specifically, in Step 1, the base model used is the large language model based on prompt engineering with the highest accuracy using the original dataset. The prompt needs to include natural language queries, database schema definitions (schemas), and optional examples. The large language model is required to generate N candidate SQLs with different syntactic structures and keywords. While ensuring the diversity of the candidate SQLs, limit their number to be less than or equal to N to avoid a large amount of additional generation overhead; at the same time, use the database execution engine in advance to filter out the queries that do not conform to the syntax.
[0073] Specifically, in Step 2, any two SQL queries need to be paired up, and an attempt is made to generate a counterexample for this pair of SQL queries.
[0074] Specifically, in Step 2, first use SQLSolver, a SQL equivalence verification tool, to filter out equivalent SQL query pairs within a set timeout. If the SQL query pair is not equivalent, use VeriEQL, a SQL equivalence verification and counterexample generation tool, to attempt to generate a counterexample for this pair of SQL queries. If a counterexample is successfully generated, retain the counterexample and process the next query pair; otherwise, skip this query pair.
[0075] Specifically, in Step 2, if there is already a generated counterexample that can prove that the current SQL query pair is not equivalent, directly skip the current query pair and do not perform subsequent verification and counterexample generation for the current SQL query pair.
[0076] Specifically, in Step 2, when using SQLSolver to filter out equivalent SQL query pairs, while ensuring the verification ability of SQLSolver, limit its verification time not to exceed G minutes to avoid the program from stalling due to too long a verification time of SQLSolver.
[0077] Specifically, in Step 2, if VeriEQL fails to successfully generate a counterexample, the method of generating through the database instance can be tried to generate a counterexample for the SQL query pair. Specifically, use database fuzz testing technology to generate a large number of database instances for the candidate SQL queries in Step 1. The database instances include the values of each table and each column, and each table contains several rows of data. If the execution results of the SQL query pair are inconsistent on a certain database instance, then this database instance is a counterexample.
[0078] Specifically, in step 2, for all the generated counterexamples, try to minimize the number of data rows included in the counterexample by attempting to delete rows one by one. Specifically, start from a table without foreign key constraints and try to delete data rows one by one in a cascading manner. If the execution result of the SQL query pair remains inconsistent after deleting a certain row of data, then the deletion is successful; otherwise, restore the row of data. Repeat the above operation for each data row of each table until no data row can be deleted.
[0079] Specifically, in step 3, the prompts provided to the large language model include natural language queries, database table names, column names, integrity constraints, data rows of counterexamples, a thought chain guidance for executing a natural language query task on a database instance, and requirements for structured output formats. Use a large language model to execute the prompt, and parse and retain the execution result.
[0080] Specifically, in steps 5 and 6, when comparing natural language and SQL query results, row and column reordering should be considered. Columns sorted in any order are considered equivalent, and the row order depends on whether sorting is required in the natural language and the original SQL query. The bag semantics of SQL also needs to be considered, that is, the same data rows are considered equivalent only when they appear the same number of times in the result. If the execution results are the same on any counterexample, add one point to the corresponding SQL query.
[0081] Specifically, in step 7, if there are multiple SQL queries with the same highest score, if the given original SQL query is the highest score, then select the original SQL query; otherwise, randomly select one SQL query from the SQL queries with the highest score as the final SQL.
[0082] The present invention also provides an automatic error detection and repair system for natural language to SQL translation. The automatic error detection and repair system for natural language to SQL translation can be implemented by executing the process steps of the automatic error detection and repair method for natural language to SQL translation. That is, those skilled in the art can understand the automatic error detection and repair method for natural language to SQL translation as a preferred implementation manner of the automatic error detection and repair system for natural language to SQL translation.
[0083] Embodiment 2
[0084] Embodiment 2 is a preferred example of Embodiment 1
[0085] According to the automatic error detection and repair method for natural language to SQL translation provided by the present invention, taking natural language query nlq, database schema definition schema, and original SQL query query as examples, the execution process of the method is as Figure 3 shown.
[0086] The automatic error detection and repair method for natural language to SQL translation includes:
[0087] Step 1: Generate candidate SQLs. Add nlq and schema to the prompts, and use the base model - GPT-4 to generate multiple candidate SQL queries corresponding to nlq. At the same time, add the original query as SQL1 to the candidate SQL queries. The generated candidate SQL queries are as Figure 2 shown in the candidate SQLs.
[0088] Step 2: Counterexample generation. Use SQLSolver to verify that SQL1 and SQL2 are not equivalent, and use VeriEQL to generate counterexample 1; SQLSolver verifies that SQL1 and SQL3 are not equivalent, but VeriEQL fails to generate a counterexample, so SQL1 and SQL3 are skipped. The generated counterexample 1 can distinguish that SQL2 and SQL3 are not equivalent, so SQL2 and SQL3 are skipped. The counterexample generated in this step is counterexample 1.
[0089] Step 3: Natural language query execution. Use a large language model to execute natural language queries on the counterexamples generated in Step 2 and obtain the results. The execution result of using GPT-4 on counterexample 1 is that the count of 'Boeing' is zero times and the count of 'Airbus' is one time.
[0090] Step 4: SQL query execution. Execute the SQL query in the database engine and obtain the execution results on the counterexamples generated in Step 2. In the SQLITE database engine, the execution result of SQL1 is that the count of 'Airbus' is one time; the execution result of SQL2 is that the count of 'Boeing' is zero times and the count of 'Airbus' is one time; the execution result of SQL3 is that the count of 'Airbus' is one time.
[0091] Step 5: Error detection. Based on the execution results of Steps 3 and 4, the execution results of the original SQL query SQL1 and the natural language nlq are inconsistent, an error is detected and Step 6 is entered.
[0092] Step 6: SQL query scoring. Based on execution consistency, compare the execution results of Steps 3 and 4, and score each candidate SQL query. The execution result of the natural language query is consistent with the execution result of SQL2, so SQL2 is counted as 1 point; it is inconsistent with the execution results of both SQL1 and SQL3, so SQL1 and SQL3 are recorded as 0 points.
[0093] Step 7: SQL query repair. Select the candidate SQL with a score higher than the highest score in the original SQL in Step 6 as the repaired SQL. Since SQL2 has the highest score, SQL2 is selected as the final SQL to replace and repair the original query.
[0094] Those skilled in the art know that in addition to implementing the system, device, and their respective modules provided by the present invention in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the system, device, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same program. Therefore, the system, device, and their respective modules provided by the present invention can be considered as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structure within the hardware component.
[0095] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. An automatic error detection and repair method for natural language to SQL translation, characterized in that Including: Step S1: Generate multiple candidate SQL queries based on a natural language query, and construct an SQL set based on the multiple candidate SQL queries and the original SQL query; Step S2: Use SQL equivalence verification technology to verify the equivalence of SQL queries in the SQL set and generate counterexamples; Step S3: Use a large language model to execute a natural language query on the generated counterexample to obtain a first execution result; Step S4: Use a database engine to execute the original SQL query on the generated counterexample to obtain a second execution result; Step S5: Compare the execution results of the natural language and the original SQL based on execution consistency. If the execution results are inconsistent for a certain counterexample, it is determined that an error has been detected; Step S6: Score each SQL query based on the comparison result; Step S7: Select the candidate SQL with a score higher than the highest-scoring candidate SQL in the original SQL query as the repaired SQL.
2. The automatic error detection and repair method for natural language to SQL translation according to claim 1, characterized in that The said Step S1 includes: Using a base model to generate multiple candidate SQL queries for a natural language query, and adding the original SQL query thereto to construct an SQL set; The said base model is the large language model with the highest accuracy rate for natural language queries; The said large language model generates multiple candidate SQL queries based on prompt engineering; wherein, the said prompt includes a natural language query, a database schema definition, and optional examples; The said prompt further includes: Requiring the large language model to generate N candidate SQLs with different syntactic structures and keywords, and restricting the number thereof to be less than or equal to N while ensuring the diversity of candidate SQLs; and simultaneously filtering out the queries that do not conform to the syntax in advance using a database execution engine.
3. The automatic error detection and repair method for natural language to SQL translation according to claim 1, characterized in that, The said Step S2 includes: Using SQLSolver to filter equivalent SQL query pairs within a set timeout. If the SQL query pair is not equivalent, use VeriEQL to generate a counterexample for the SQL query pair. If a counterexample is successfully generated, retain the counterexample and process the next query pair; otherwise, skip the query pair.
4. The automatic error detection and repair method for natural language to SQL translation according to claim 3, characterized in that The said Step S2 further includes: If VeriEQL fails to successfully generate a counterexample, generate a counterexample for the SQL query pair by means of database instance generation; The method of generating a counterexample for the SQL query pair by means of database instance generation includes: Using database fuzz testing technology to generate database instances for candidate SQL queries. The database instances include the values of each table and each column, and each table contains several rows of data; if the execution results of the SQL query pair are inconsistent on a certain database instance, then the database instance is a counterexample.
5. The automatic error detection and repair method for natural language to SQL translation according to claim 1, characterized in that The said Step S2 further includes: For all generated counterexamples, reduce the number of data rows included in the counterexample by means of row-by-row deletion; The method of reducing the number of data rows included in the counterexample by means of row-by-row deletion includes: Starting from a table without foreign key constraints, deleting data rows row by row in a cascading manner. If the execution results of the SQL query pair are still inconsistent after deleting a certain row of data, then the deletion is successful; otherwise, restore the row of data; repeat the above operation for each data row of each table until no data row can be deleted.
6. The automatic error detection and repair method for natural language to SQL translation according to claim 1, characterized in that, Step S3 includes: providing a prompt for the large language model, where the prompt includes a natural language query, database table names, column names, integrity constraints, data rows of counterexamples, a chain of thought guidance for executing a natural language query task on a database instance, and structured output format requirements; using a large language model to execute the prompt, and parsing and retaining the execution result.
7. The automatic error detection and repair method for natural language to SQL translation according to claim 1, characterized in that When comparing the execution results of the natural language and the original SQL, if the following conditions are simultaneously met, it is considered that the execution result meets the execution consistency condition; Condition 1: Columns sorted in any order are regarded as equivalent; Condition 2: The row order depends on whether sorting is required in the natural language and the original SQL query; Condition 3: The same data rows appear the same number of times in the result.
8. The automatic error detection and repair method for natural language to SQL translation according to claim 1, characterized in that Step S7 includes: If there are multiple SQL queries with the same highest score, if the given original SQL query is the highest score, then select the original SQL query; otherwise, randomly select one SQL query from the SQL queries with the highest score as the final SQL.
9. An automatic error detection and repair system for natural language to SQL translation, characterized in that, Includes: Module M1: Generate multiple candidate SQL queries based on the natural language query, and construct an SQL set based on the multiple candidate SQL queries and the original SQL query; Module M2: Use SQL equivalence verification technology to verify the equivalence of the SQL queries in the SQL set and generate counterexamples; Module M3: Use a large language model to execute the natural language query on the generated counterexamples to obtain the first execution result; Module M4: Use the database engine to execute the original SQL query on the generated counterexamples to obtain the second execution result; Module M5: Compare the execution results of the natural language and the original SQL based on execution consistency. If the execution results are inconsistent on a certain counterexample, it is determined that an error is detected; Module M6: Score each SQL query based on the comparison result; Module M7: Select the candidate SQL with a score higher than the highest score among the original SQL queries as the repaired SQL.
10. The automatic error detection and repair system for natural language to SQL translation according to claim 9, characterized in that, Module M1 includes: Using a base model to generate multiple candidate SQL queries for the natural language query, and adding the original SQL query to it to construct an SQL set; The base model is the large language model with the highest accuracy rate for using natural language queries; The large language model generates multiple candidate SQL queries based on prompt engineering; among them, the prompt includes a natural language query, database schema definition, and optional examples; The prompt also includes: Requiring the large language model to generate N candidate SQLs with different syntactic structures and keywords, and restricting the number to be less than or equal to N while ensuring the diversity of candidate SQLs; at the same time, using the database execution engine in advance to filter out the queries that do not conform to the syntax; Module M2 includes: Using SQLSolver to filter equivalent SQL query pairs within a set timeout. If the SQL query pairs are not equivalent, use VeriEQL to generate counterexamples for the SQL query pairs. If counterexamples are successfully generated, retain the counterexamples and process the next query pair; otherwise, skip the query pair; Module M2 also includes: If VeriEQL fails to successfully generate counterexamples, generate counterexamples for the SQL query pairs by means of database instance generation; The method of generating by means of a database instance is to generate counterexamples for SQL query pairs, including: Using the method of database fuzz testing technology to generate database instances for candidate SQL query pairs. The database instance includes the values of each table and each column, and each table contains several rows of data. If the execution results of the SQL query pair are inconsistent on a certain database instance, then this database instance is a counterexample. The module M3 includes: providing prompts for the large language model, where the prompts include natural language queries, database table names, column names, integrity constraints, data rows of counterexamples, the thought chain guidance for executing a natural language query task on a database instance, and the requirements for the structured output format; using a large language model to execute this prompt, and parsing and retaining the execution results. When comparing the execution results of natural language and the original SQL, if the following conditions are met simultaneously, it is considered that the execution results meet the execution consistency conditions. Condition 1: Columns sorted in any order are regarded as equivalent. Condition 2: The row order depends on whether sorting is required in the natural language and the original SQL query. Condition 3: The same data rows appear the same number of times in the results.