Text2SQL (Structured Query Language) generation method based on self error correction
By introducing meta-learning and self-supervised learning mechanisms into Text2SQL technology, the Text2SQL generation method for self-correction is realized, which solves the problems of poor data adaptability and insufficient error correction capabilities in the existing technology, and improves the accuracy and reliability of generating SQL query statements.
Patent Information
- Application Number
- CN202510059379.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The existing Text2SQL technology has training data dependence in practical applications, is difficult to adapt to new data sets, and generates SQL query statements in complex query scenarios that are prone to errors and lacks automatic error correction capabilities, resulting in low generation efficiency and insufficient flexibility.
A Text2SQL generation method based on self-correction is proposed. The structured query statement is trained to generate a task model by using the meta-learning mechanism and the self-supervised learning mechanism. Through multiple verification and correction of the generated SQL query statement, it improves its accuracy and reliability.
Through multiple verification and correction optimization, the accuracy and reliability of the generated SQL query statements are significantly improved, and the ability to meet complex query needs and generation efficiency are enhanced.
Smart Images

Figure CN120030035A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a Text2SQL generation method based on self-correction. Background Art
[0002] With the rapid development of natural language processing technology, the technology of converting natural language queries into Structured Query Language (SQL for short, full English name Structured Query Language), that is, Text2SQL, which allows users to interact with databases in the form of natural language, is widely used in scenarios such as intelligent question-answering systems, business intelligence analysis, and automated data queries. However, the existing Text2SQL technology still faces the following key problems in practical applications: Traditional models rely on data in specific domains or with fixed structures for training, making it difficult to quickly adapt to new data sets, which limits its wide application; in complex query scenarios (such as multi-table joins, nested queries), it is easy to make mistakes when generating SQL query statements, lacking the ability of automatic error correction and only relying on manual correction, with low error correction efficiency; using static templates or manually designed modes is difficult to meet dynamic and complex query requirements, resulting in low efficiency and insufficient flexibility in generating SQL query statements.
[0003] Therefore, based on the above problems, there is an urgent need to propose a structured query statement generation method to improve the accuracy and reliability of generating SQL query statements. Summary of the Invention
[0004] In order to solve at least one or more of the above-mentioned technical problems, the present invention proposes a Text2SQL generation solution based on self-correction in multiple aspects.
[0005] In a first aspect, an embodiment of the present invention provides a Text2SQL generation method based on self-correction, the method comprising: obtaining a first natural language query text; testing the first natural language query text using a trained structured query statement generation task model to obtain a first structured query statement corresponding to the first natural language query text, wherein the structured query statement generation task model is obtained by training and learning according to a meta-learning mechanism and a self-supervised learning mechanism; parsing the first natural language query text using a pre-trained large language model to obtain a query operation mode corresponding to the first natural language query text; when it is determined that the first structured query statement does not meet the query requirements corresponding to the query operation mode, dynamically optimizing the first structured query statement to obtain a second structured query statement; executing the second structured query statement to obtain a first query result; when it is determined that the first query result does not meet the expected result, correcting and optimizing the second structured query statement according to the first query result to obtain a third structured query statement, and the third structured query statement is output as a target structured query statement corresponding to the first natural language query text.
[0006] In some embodiments, the method further includes: when it is determined that the first structured query statement meets the query requirements corresponding to the query operation mode, executing the first structured query statement to obtain a second query result; when it is determined that the second query result does not meet the expected result, correcting and optimizing the first structured query statement according to the second query result to obtain a fourth structured query statement, and the fourth structured query statement is output as a target structured query statement corresponding to the first natural language query text.
[0007] In some embodiments, dynamically optimizing the first structured query statement includes: selecting a query template related to the first natural language query text from a preconfigured template library; optimizing the first structured query statement according to the query template to obtain a second structured query statement.
[0008] In some embodiments, the first query result includes an error type indicating that the second structured query statement exists, and the second structured query statement is corrected and optimized according to the first query result, including: determining an adjustment strategy corresponding to the error type; and correcting the second structured query statement according to the adjustment strategy corresponding to the error type.
[0009] In some embodiments, a structured query statement generation task model obtained by training and learning according to a meta-learning mechanism and a self-supervised learning mechanism includes: constructing a labeled first training data set, the first training data set including multiple training data pairs, each training data pair including a training natural language query text corresponding to an actual application scenario, and a training structured query statement corresponding to the training natural query text; obtaining a second training data set corresponding to each open source data set from multiple different open source data sets respectively; according to a meta-learning algorithm, using multiple second training data sets and the first training data set, pre-training the structured query statement generation task model to be trained to obtain a preliminary structured query statement generation task model; fine-tuning the preliminary structured query statement generation task model to obtain a current structured query statement generation task model; according to a self-supervised learning algorithm, supervising and detecting the output of the current structured query statement generation task model to obtain a trained structured query statement generation task model.
[0010] In some embodiments, before pre-training the structured query statement generation task model to be trained using multiple second training data sets and a first training data set according to a meta-learning algorithm, the method also includes: using a pre-trained large language model to perform semantic expansion on each second natural language query text to obtain multiple third natural language query texts with the same semantics and different text expressions corresponding to each second natural language query text; constructing a third training data set based on the correspondence between the multiple third natural language query texts, the second natural language query texts and the training structured query statements; and updating the third training data set to the first training data set.
[0011] In some embodiments, after outputting a target structured query statement corresponding to the first natural language query text, the method further includes: using the first natural language query text and the target structured query statement as a new training data pair; and adding the new training data pair to the first training data set to obtain an updated first training data set.
[0012] The present invention provides a Text2SQL generation method based on self-correction. The method first obtains a first natural language query text; then uses a trained structured query statement generation task model to test the first natural language query text to obtain a first structured query statement corresponding to the first natural language query text, and the structured query statement generation task model is trained and learned according to a meta-learning mechanism and a self-supervised learning mechanism; then, uses a pre-trained large language model to parse the first natural language query text to obtain a query operation mode corresponding to the first natural language query text; then, when it is determined that the first structured query statement does not meet the query requirement corresponding to the query operation mode, the first structured query statement is dynamically optimized to obtain a second structured query statement; the second structured query statement is executed to obtain a first query result; finally, when it is determined that the first query result does not meet the expected result, the second structured query statement is corrected and optimized according to the first query result to obtain a third structured query statement, and the third structured query statement is output as a target structured query statement corresponding to the first natural language query text. Compared with the related art, the technical solution provided by the embodiment of the present invention verifies the SQL query statements generated by model prediction multiple times, automatically discovers and corrects potential errors, and effectively improves the accuracy and reliability of the generated SQL query statements. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0014] Figure 1 is an exemplary schematic diagram of a Text2SQL generation method 100 based on self-correction according to an embodiment of the present invention;
[0015] Figure 2 An exemplary flow chart of a self-correcting Text2SQL generation method 200 according to some other embodiments of the present invention is shown;
[0016] Figure 3 An exemplary flow chart of a structured query statement generation task model training method 300 according to other embodiments of the present invention is shown;
[0017] Figure 4 is a structural diagram of a Text2SQL generation device 400 based on self-correction according to an embodiment of the present invention;
[0018] Figure 5 A structural schematic diagram showing a structured query statement generation task model training device 500 according to some other embodiments of the present invention;
[0019] Figure 6 A schematic block diagram of an electronic device 600 according to an embodiment of the present invention is shown.
[0020] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION
[0021] In order to better understand and explain the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings. The present invention is not limited to these specific embodiments. On the contrary, modifications or equivalent substitutions made to the present invention should all be included in the scope of the claims of the present invention.
[0022] It should be noted that numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present invention can also be implemented without these specific details. In the multiple specific embodiments given below, the principles, structures and components well known in the art are not described in detail in order to highlight the main purpose of the present invention.
[0023] The self-correcting Text2SQL generation method provided by the present invention can be executed by a computer device, wherein the computer device can be a terminal or a server. The terminal can be a terminal device such as a smart phone, a tablet computer, a laptop computer, a touch screen, a personal computer (PC), a personal digital assistant (PDA), etc., and the terminal can also include a client. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0024] Please refer to Figure 1 , Figure 1 FIG. 1 is an exemplary schematic diagram of a Text2SQL generation method 100 based on self-correction according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0025] In step S101, a first natural language query text is obtained.
[0026] In the above steps, the natural language query text (NLPT) refers to a query request for a database proposed by a user in natural language. The natural language query text can be represented as an NLP query text. The first natural language query text is a natural language query text to be processed that is received by the electronic device after completing the model training.
[0027] In step S102, the first natural language query text is tested using the trained structured query statement generation task model to obtain a first structured query statement corresponding to the first natural language query text. The structured query statement generation task model is trained and learned according to a meta-learning mechanism and a self-supervised learning mechanism.
[0028] In the above steps, the structured query statement generation task model is trained and learned based on the meta-learning mechanism and the self-supervised learning mechanism. It can be a pre-trained large language model (such as Generative Pretrained Transformer, GPT) or a deep learning model optimized for the Text2SQL task. Meta-learning algorithms, such as MAML (Model-Agnostic Meta-Learning) and Repti le, update model parameters through multiple iterations, so that it can quickly adapt to new tasks on a small amount of sample data.
[0029] A structured query statement is a command used to retrieve specific information from a database. SQL (Structured Query Language) is a programming language specifically used to manage relational database systems. Through SQL query statements, users can specify the data they want to obtain from the database, including which columns to select, which rows to filter, and how to sort and group the results. The basic structure of an SQL query statement usually includes the following main parts: The SELECT keyword is followed by the column name (or expression) to be retrieved, which specifies which columns of data should be included in the query results. The FROM keyword is followed by the name of the data table to be queried. This part specifies which table the data comes from. The WHERE clause is used to specify the filtering conditions, and only rows that meet these conditions will be included in the query results. The GROUP BY clause is used to group the query results by one or more columns. This clause is usually used with aggregate functions (such as COUNT, SUM, AVG, etc.) to perform aggregate operations on each group. The ORDER BY clause is used to specify how the query results are sorted. This clause can be followed by one or more column names, as well as the sorting order of each column (ascending or descending).
[0030] Use the trained structured query statement generation task model to test the first natural language query text, and obtain the first structured query statement corresponding to the first natural language query text. Suppose the first natural language query text input by the user is: "Query the order amount, payment method, and order status of customers in the East China region in 2019." Using the trained structured query statement generation task model, the corresponding SQL query can be generated as:
[0031] SELECT order_amount,payment_method FROM orders WHERE region='East' AND year=2019。
[0032] In step S103, use the pre-trained large language model to parse the first natural language query text, and obtain the query operation mode corresponding to the first natural language query text.
[0033] In step S104, when it is determined that the first structured query statement does not meet the query requirements corresponding to the query operation mode, dynamically optimize the first structured query statement to obtain the second structured query statement.
[0034] In step S105, execute the second structured query statement to obtain the first query result.
[0035] In step S106, when it is determined that the first query result does not meet the expected result, correct and optimize the second structured query statement according to the first query result to obtain the third structured query statement, and the third structured query statement is output as the target structured query statement corresponding to the first natural language query text.
[0036] In the above steps, the query operation mode refers to the query operation that needs to be executed in the database determined according to the query intention of the natural language query text. For example, the query operation mode includes but is not limited to multi-table join, nested query, complex aggregation, etc. Suppose in a relational database management system (full English name Relational Database Management System, English abbreviation RDBMS), multi-table join, nested query, and complex aggregation are three common query operations, which are used to retrieve and summarize data from one or more tables.
[0037] A multi-table join is a SQL query operation that allows two or more tables to be combined based on some common columns (usually primary keys and foreign keys), thereby retrieving data from multiple tables in the same query. Types of multi-table joins include INNER JOIN, LEFT OUTER JOIN, RIGHT OUTER JOIN, FULL OUTER JOIN, and CROSS JOIN. For example, if there are two tables, one is an order table (Orders) and the other is a customer table (Customers), you can use a multi-table join to join the two tables based on the customer ID (assuming it is a common column between the two tables) so that you can see the order information and its corresponding customer information in the same query result.
[0038] Nested query, also known as subquery, is a SQL query that is embedded inside another query. The outer query (also known as the main query) depends on the result of the inner query (subquery). Nested queries are often used in situations where you need to determine a condition based on the result of another query. For example, you want to find all orders whose order amount is greater than the average order amount. In this case, write a nested query in which the inner query calculates the average amount of all orders, and the outer query finds all orders greater than this average.
[0039] Complex aggregation refers to the use of aggregate functions (such as SUM, AVG, COUNT, MAX, MIN, etc.) in SQL queries to perform more complex analysis and calculations. Complex aggregation may involve aggregation calculations of multiple columns, grouped aggregation (GROUP BY), having clause filtering, aggregation operations after multiple table joins, etc. For example, you want to calculate the total order amount of each customer and only display those customers whose total order amount exceeds a certain threshold. First, group by customer, then calculate the total order amount of each customer, and finally filter out the records whose total amount exceeds the threshold. These three concepts represent different types of operations in SQL queries, which demonstrate the power of SQL to process and analyze complex data relationships stored in relational databases.
[0040] When it is determined that the first structured query statement does not meet the query requirements corresponding to the query operation mode, the first structured query statement is dynamically optimized to obtain a second structured query statement. For example, the first natural language query text is: "Query the order amount, payment method and order status of customers in East China in 2019". For the field "order status", the first SQL query statement generated does not include this field, nor does it process the associated customer information. Therefore, it is evaluated that the first SQL query statement cannot fully meet the query requirements. In order to solve this problem, a query template related to the first natural language query text is selected from a pre-configured template library; the first SQL query statement is optimized according to the query template to obtain a second structured query statement. Query templates include but are not limited to: join query templates, nested query templates and aggregate query templates. The second SQL query statement can be as follows:
[0041] SELECT o.order_amount,o.payment_method,o.order_status,c.customer_nameFROM orders o
[0042] JOIN customers c ON o.customer_id=c.customer_id
[0043] WHERE o.region='East'AND o.year=2019;
[0044] In the second structured query statement, the grammatical structure of the second SQL query statement is dynamically adjusted to ensure that the second structured query statement includes the "order status" field (order_status) and the customer name (customer_name) that needs to be obtained from the "customers" table using the join query, so that the query result of the second structured query statement can accurately display the customer information. In this way, more complex query requirements can be flexibly met and more accurate SQL queries can be generated.
[0045] The second structured query statement is executed in the database and compared with the expected query result. Through this verification process, potential errors in the generated SQL query can be identified in real time, such as missing fields, table connection errors, query logic errors, etc. When it is found that the SQL query result deviates from the expected result, the generated SQL query statement is corrected and the model is optimized based on the error feedback from the query result, thereby improving the accuracy of the generated result and the reliability of execution.
[0046] The self-correction-based Text2SQL generation method proposed in the present invention first generates SQL query statements using a structured query statement generation task model, then dynamically optimizes the SQL query statements according to the query operation mode corresponding to the natural language query text, and finally corrects the SQL query statements again according to the query results of the SQL query statements. Optimizing SQL query statements according to the query operation mode can better meet complex query requirements and significantly improve the flexibility and efficiency of SQL query statement generation; correcting SQL query statements again according to the query results can improve the accuracy of SQL query statements; the structured query statement generation task model obtained by training and learning according to the meta-learning mechanism and the self-supervised learning mechanism can quickly adjust the learning strategy to complete the test task, and effectively improve the generalization ability of the model in different fields.
[0047] Figure 2 FIG. 2 shows an exemplary flow chart of a Text2SQL generation method 200 based on self-correction according to some other embodiments of the present invention. Figure 2 As shown, the method includes:
[0048] In step S201, a first natural language query text is obtained;
[0049] In step S202, the first natural language query text is tested using the trained structured query statement generation task model to obtain a first structured query statement corresponding to the first natural language query text. The structured query statement generation task model is trained and learned according to a meta-learning mechanism and a self-supervised learning mechanism.
[0050] In step S203, the first natural language query text is parsed using a pre-trained large language model to obtain a query operation mode corresponding to the first natural language query text.
[0051] In step S204, when it is determined that the first structured query statement does not meet the query requirement corresponding to the query operation mode, the first structured query statement is dynamically optimized to obtain a second structured query statement.
[0052] In step S205, the second structured query statement is executed to obtain a first query result.
[0053] In step S206, when it is determined that the first query result does not meet the expected result, the second structured query statement is modified and optimized according to the first query result to obtain a third structured query statement, and the third structured query statement is output as a target structured query statement corresponding to the first natural language query text.
[0054] In step S207, when it is determined that the first query result meets the expected result, a second structured query statement is output as a target structured query statement corresponding to the first natural language query text.
[0055] In step S208, when it is determined that the first structured query statement meets the query requirement corresponding to the query operation mode, the first structured query statement is executed to obtain a second query result.
[0056] In step S209, when it is determined that the second query result meets the expected result, the first structured query statement is output as the target structured query statement.
[0057] In step S210, when it is determined that the second query result does not meet the expected result, the first structured query statement is modified and optimized according to the second query result to obtain a fourth structured query statement, and the fourth structured query statement is output as a target structured query statement corresponding to the first natural language query text.
[0058] In the above steps, the trained structured query statement generation task model is used to generate a first structured query statement corresponding to the first natural language query text, and then the first structured query statement is subjected to a first self-correction check according to the query operation mode corresponding to the first natural language query text. When it is determined that the first structured query statement meets the query requirements corresponding to the query operation mode, it is indicated that the first SQL query statement meets the requirements of the query operation mode, and then, by executing the first SQL query statement, it is determined whether a second self-correction check is required for the first SQL query statement. The first SQL query statement is executed to obtain a second query result, and when the second query result meets the expected result, the first SQL query statement is output, and the first SQL query statement is the target SQL query statement corresponding to the first natural language query text.
[0059] When it is determined that the first structured query statement does not meet the query requirements corresponding to the query operation mode, it means that the first structured query statement still has errors that need to be corrected in a complex query scenario. The first SQL query statement is corrected by dynamically optimizing the first SQL query statement to obtain a second SQL query statement. Then, the second SQL query statement is executed, and a second self-correction check is performed on the second SQL query statement through feedback of the query result. The second SQL query statement is executed to obtain the first query result. Then, when it is determined that the first query result meets the expected result, it means that the second SQL query statement has overcome the relevant problems after the first self-correction check and can be directly output. When it is determined that the first query result does not meet the expected result, the second SQL query statement is corrected and optimized according to the first query result to obtain a third SQL query statement, which is a SQL query statement that has been corrected twice.
[0060] In some embodiments, the second structured query statement is modified and optimized according to the first query result, including determining an adjustment strategy corresponding to the error type; and modifying the second structured query statement according to the adjustment strategy corresponding to the error type.
[0061] After executing SQL queries in professional databases, the query results are compared with the expected results to identify errors in SQL query statements, and then the SQL query statements are corrected based on these errors. Through multiple rounds of feedback loops, the SQL generation strategy is continuously improved, and common errors in SQL query statements, such as field mismatches, missing fields, logical errors, etc., are gradually eliminated. After each correction is completed, the corrected SQL query statement and its corresponding natural language query text are used as new training data pairs and fed back to the training data set used to train the structured query statement generation task model. The updated training data set is used to adjust the generation strategy of the structured query statement generation task model.
[0062] When the query result of the SQL query statement indicates that there are some missing fields, the missing fields are automatically identified and the missing fields are supplemented into the generated SQL query statement through a feedback mechanism.
[0063] When the query result of the SQL query statement indicates that the query information is incomplete or inaccurate, the SQL query statement is adjusted according to the difference between the actual data and the expected result. For example, the query conditions are adjusted or a join query operation is added to the SQL query statement.
[0064] When the query result of the SQL query statement indicates that there is a table join error in the SQL query statement, the data related to the table join error in the SQL query statement is adjusted according to the query result of the SQL query statement. The table join error may be, for example, a missing join condition, an incorrect join condition order, or a join condition that cannot correctly associate multiple data tables. The data related to the table join error may be, for example, a field or condition missing in the supplementary join condition, or ensuring the accuracy of the association between different tables.
[0065] When the query result of the SQL query statement indicates that there is a logical error in the SQL query statement, add aggregation operations, grouping conditions, or filtering conditions. Ensure the accuracy of the query result of the SQL query statement by adding aggregation operations, grouping conditions, or filtering conditions. Logical errors refer to errors such as failure to add aggregation operations or omission of conditions.
[0066] When the query result of an SQL query statement differs greatly from the expected result, the query fields or conditions in the SQL query statement are adjusted to optimize the query data range and query logic so that the query result of the SQL query statement approaches the expected result.
[0067] When the query result of the SQL query statement indicates that there is a syntax error in the SQL query statement, the SQL query statement is adjusted according to the syntax rule to obtain an SQL query statement that complies with the SQL syntax specification. The syntax error includes but is not limited to spelling errors, mismatched brackets, and keyword errors.
[0068] When the query result of the SQL query statement indicates that the SQL query statement has a data reference error, the data reference error in the SQL query statement is adjusted according to the context and feedback information of the first natural language query text. The data reference error includes but is not limited to an invalid column name, table name or constant value.
[0069] Assume that the natural language query text entered by the user is: "Show the order amounts of all customers in East China in 2019 and the orders in the 'paid' status". The initially generated SQL query statement is as follows:
[0070] SELECT order_amount,order_status FROM orders WHEREregion='East'ANDyear=2019;
[0071] After executing the SQL query statement, compare the query result of the SQL query statement with the expected result to obtain the comparison result. When the query result does not match the expected result, after analysis, it can be identified that the data range of the "Order Status" field in the SQL query statement is not accurately defined. After executing the SQL query statement, only the status of all orders is returned. However, the user's requirement is to display the order amount of all customers and orders in the "paid" status. Therefore, the initially generated SQL query statement does not contain the filter condition for the "paid" status. According to the execution feedback, it is identified that the data range does not meet the expected results, and the "data range adjustment" strategy is adopted to correct the SQL query statement. Add a filter condition to the SQL query statement to limit the order status to "paid". The final generated SQL query is:
[0072] SELECT order_amount,order_status
[0073] FROM orders WHEREregion='East'AND year=2019AND order_status='Paid';
[0074] Through the above adjustments, the revised SQL query statement is executed to return only the amount and status of the "paid" order, thus meeting the user's query needs. In this process, by automatically identifying that the data range does not meet expectations, adjusting the filter conditions in the SQL query statement, and optimizing the query logic of the SQL query statement, the revised SQL query statement can accurately return the required data results.
[0075] The self-correction-based Text2SQL generation method provided by the present invention, through the above-mentioned two self-correction SQL query statement processing methods, continuously feeds back and corrects when relevant conditions are not met, and gradually optimizes the SQL query statement generation strategy, which significantly improves the accuracy of query results and the query statement generation efficiency.
[0076] Figure 3 FIG. 3 is an exemplary flow chart of a structured query statement generation task model training method 300 according to some other embodiments of the present invention. Figure 3 As shown, the method includes:
[0077] In step S301, a first labeled training data set is constructed, the first training data set includes multiple training data pairs, each training data pair includes a second natural language query text corresponding to an actual application scenario, and a third structured query statement corresponding to the second natural language query text.
[0078] In some embodiments, before pre-training the structured query statement generation task model to be trained using multiple second training data sets and a first training data set according to a meta-learning algorithm, the method also includes: using a pre-trained large language model to perform semantic expansion on each second natural language query text to obtain multiple third natural language query texts with the same semantics and different text expressions corresponding to each second natural language query text; constructing a third training data set based on the correspondence between the multiple third natural language query texts, the second natural language query texts and the training structured query statements; and updating the third training data set to the first training data set.
[0079] Assume that in the order query application, the expert team annotates the natural language query text "Query the total order amount of customers in East China in 2019" from the actual sales data, and obtains the SQL query corresponding to the natural language query text as follows:
[0080] SELECT SUM(order_amount)FROM orders WHEREregion='East'AND year=2019;
[0081] Based on the annotated natural language query text "Query the total order amount of customers in East China in 2019", the pre-trained large language model is used to generate multiple natural language query texts with different expressions, such as: "Count the order amount of customers in East China in 2019", "Calculate the total order amount of all customers in East China in 2019", "View the total order amount of customers in East China in 2019", etc.
[0082] In this way, different forms of query inputs can be processed, effectively improving the robustness of the model.
[0083] In step S302, second training data sets corresponding to each open source data set are obtained from a plurality of different open source data sets.
[0084] After the first training data set is constructed, the structured query statement generation task model to be trained is pre-trained using multiple second training data sets and the first training data set to obtain a preliminary structured query statement generation task model. The multiple second training data sets may be, for example, open source Text2SQL data sets such as WikiSQL, SPIDER, and BI RD.
[0085] The first training data set is a training data set in a professional field, for example, a training data set in an order query application.
[0086] In some embodiments, after outputting a target structured query statement corresponding to the first natural language query text, the method includes: constructing the first natural language query text and the target structured query statement into a new training data pair; and appending the new training data pair to the first training data set to obtain an updated first training data set.
[0087] In step S303, according to the meta-learning algorithm, the structured query statement generation task model to be trained is pre-trained using multiple second training data sets and the first training data set to obtain a preliminary structured query statement generation task model.
[0088] In the above steps, a large language model is pre-trained using a first training data set and a second training data set through a meta-learning algorithm (such as Model-Agnostic Meta-Learning, MAML), and a processing strategy for a common SQL query structure and a common query operation mode can be obtained through meta-learning. The query operation mode is, for example, a simple query, a join query, an aggregate query, etc. The introduction of the meta-learning mechanism enables the structured query statement generation task model to be quickly adjusted according to a small amount of new domain-specific data, thereby optimizing the SQL query statement generation strategy.
[0089] In step S304, the preliminary structured query statement generation task model is fine-tuned to obtain a current structured query statement generation task model.
[0090] In step S305, supervised detection is performed on the output of the current structured query statement generation task model according to the self-supervised learning algorithm to obtain a trained structured query statement generation task model.
[0091] Assume that according to the meta-learning algorithm, after pre-training on the WikiSQL, SPIDER, and BI RD datasets, the large language model has learned the basic structure of SQL queries in general fields and common complex query patterns in general fields. When the pre-trained structured query statement generation task model is applied to process a specific domain task (such as the customer order data analysis domain), the initial parameters of the pre-trained structured query statement generation task model are quickly fine-tuned using a small amount of training data in the customer order data analysis domain according to the meta-learning mechanism to obtain the optimal query generation strategy to generate SQL query statements that meet the query requirements in the customer order data analysis domain.
[0092] For example, the natural language query text for customer order data is "Query the total order amount of customers in East China in 2019". Based on the existing pre-trained knowledge (such as aggregation query, join query, etc.), a small amount of data in the field of customer order data analysis is used for rapid fine-tuning to generate the correct SQL query statement.
[0093] The transfer efficiency is significantly improved by the meta - learning mechanism without retraining the entire model.
[0094] In the training stage, through the self - supervised learning mechanism, the SQL query statements generated by the pre - trained model can be compared with the labeled SQL query statements in the labeled training data pairs, automatically checking and correcting the errors in the generated SQL query statements, thus optimizing the SQL query generation process. By comparing with the labeled SQL query statements, the deviation between the generated SQL query statements and the labeled SQL query statements can be found. When there are deviations (such as missing fields, incorrect conditions, etc.) between the generated SQL query statements and the labeled SQL query statements, these errors can be corrected in a timely manner through the self - supervised feedback mechanism. After multiple rounds of self - supervised training, learning from the errors continuously, the quality of SQL query result generation in the training stage can be gradually improved.
[0095] Suppose, in the order data analysis, the natural language query text entered by the user is "Query the total order amount of each customer in the East China region in 2019". The SQL query statement generated by the pre - trained model is:
[0096] SELECT SUM(order_amount)FROM orders WHERE region='East' AND year=2019 GROUP BY customer_id;
[0097] The labeled SQL query statement is:
[0098] SELECT customer_id,SUM(order_amount)FROM orders WHERE region='East' AND year=2019 GROUP BY customer_id;
[0099] Comparing the generated SQL query statement with the annotated SQL query statement, the generated SQL query statement lacks the "customer_id" field, which is the key information for achieving the user's query purpose. Although the "SUM(order_amount)" part correctly calculates the total order amount of each customer, the generated SQL query statement does not include the "customer_id" field, and the query result of the generated SQL query statement will not be able to identify the total order amount of each customer. Therefore, after the generated SQL query statement is checked by the self-supervision mechanism, the generation strategy of the generated SQL query statement can be adjusted through the feedback of the inspection result, so that the generated SQL query statement contains the "customer_id" field to meet the user's query needs, and the final generated SQL query statement is consistent with the annotated SQL query statement.
[0100] The structured query statement generation task model training method proposed in the present invention solves the problem of poor cross-dataset adaptability in the existing Text2SQL technology by introducing a meta-learning method. Adaptive meta-learning enables the model to quickly adjust the learning strategy according to the data of the new data set, optimize the generalization ability of the model in different fields, and flexibly respond to different fields and tasks, and achieve rapid migration. Furthermore, by introducing a self-supervised optimization mechanism, potential problems in the training data set are automatically learned and discovered, reducing the need for manual labeling, and the SQL query statement generated by prediction is compared with the standard SQL query statement, and self-correction is performed according to the results, thereby improving the accuracy of the generated SQL query statement.
[0101] Please refer to Figure 4 , Figure 4 FIG. 4 is a schematic diagram of a structure of a Text2SQL generating device 400 based on self-correction according to an embodiment of the present invention. Figure 4 As shown, the structured query statement generating device 400 includes:
[0102] The natural language query text acquisition module 401 is configured to acquire a first natural language query text.
[0103] The structured query statement generation module 402 is configured to use a trained structured query statement generation task model to test the first natural language query text to obtain a first structured query statement corresponding to the first natural language query text. The structured query statement generation task model is trained and learned according to a meta-learning mechanism and a self-supervised learning mechanism.
[0104] The query operation mode parsing module 403 is configured to parse the first natural language query text using a pre-trained large language model to obtain a query operation mode corresponding to the first natural language query text.
[0105] The structured query statement dynamic optimization module 404 is configured to dynamically optimize the first structured query statement to obtain a second structured query statement when it is determined that the first structured query statement does not meet the query requirement corresponding to the query operation mode.
[0106] The structured query statement execution module 405 is configured to execute the second structured query statement to obtain the first query result;
[0107] The structured query statement correction module 406 is configured to correct and optimize the second structured query statement according to the first query result when it is determined that the first query result does not meet the expected result, to obtain a third structured query statement, and the third structured query statement is output as a target structured query statement corresponding to the first natural language query text.
[0108] In some embodiments, the structured query statement execution module 405 is configured to execute the first structured query statement to obtain a second query result when it is determined that the first structured query statement meets the query requirement corresponding to the query operation mode;
[0109] The structured query statement correction module 406 is configured to correct and optimize the first structured query statement according to the second query result when it is determined that the second query result does not meet the expected result, to obtain a fourth structured query statement, and the fourth structured query statement is output as a target structured query statement corresponding to the first natural language query text.
[0110] In some embodiments, the structured query statement dynamic optimization module 404 is configured to select a query template related to the first natural language query text from a pre-configured template library; optimize the first structured query statement according to the query template to obtain a second structured query statement.
[0111] In some embodiments, the first query result includes an error type indicating that the second structured query statement exists, and the structured query statement correction module 406 is configured to determine an adjustment strategy corresponding to the error type; and correct the second structured query statement according to the adjustment strategy corresponding to the error type.
[0112] The self-correcting Text2SQL generation device provided by the present invention, through the above-mentioned two self-correcting SQL query statement processing methods, continuously provides feedback and correction when relevant conditions are not met, and gradually optimizes the SQL query statement generation strategy, which significantly improves the accuracy of query results and the generation efficiency of query statements.
[0113] Figure 5FIG. 5 is a schematic diagram showing a structured query statement generation task model training device 500 according to some other embodiments of the present invention. Figure 5 As shown, the device 500 includes:
[0114] The training data construction module 501 is configured to construct a first labeled training data set, the first training data set includes multiple training data pairs, each training data pair includes a training natural language query text corresponding to an actual application scenario, and a training structured query statement corresponding to the training natural language query text.
[0115] The open source data acquisition module 502 is configured to respectively acquire a second training data set corresponding to each open source data set from a plurality of different open source data sets.
[0116] The meta-learning pre-training module 503 is configured to pre-train the structured query statement generation task model to be trained using multiple second training data sets and the first training data set according to the meta-learning algorithm to obtain a preliminary structured query statement generation task model.
[0117] The meta-learning fine-tuning module 504 is configured to fine-tune the preliminary structured query statement generation task model to obtain a current structured query statement generation task model;
[0118] The self-supervised learning module 505 is configured to perform supervised detection on the output of the current structured query statement generation task model according to the self-supervised learning algorithm to obtain a trained structured query statement generation task model.
[0119] In some embodiments, the apparatus 500 further includes:
[0120] A text expansion module is configured to perform semantic expansion on each second natural language query text using a pre-trained large language model to obtain a plurality of third natural language query texts with the same semantics and different text expressions corresponding to each second natural language query text;
[0121] The training data construction module 501 is further configured to construct a third training data set according to the correspondence between the plurality of third natural language query texts, the second natural language query texts and the training structured query statements; and update the third training data set to the first training data set.
[0122] In some embodiments, the apparatus further comprises:
[0123] A training data pair construction module is configured to use the first natural language query text and the target structured query sentence as a new training data pair;
[0124] The training data updating module is configured to add the new training data pair to the first training data set to obtain an updated first training data set.
[0125] The structured query statement generation task model training device proposed in the present invention solves the problem of poor cross-dataset adaptability in the existing Text2SQL technology by introducing a meta-learning method. Adaptive meta-learning enables the model to quickly adjust the learning strategy according to the data of the new data set, optimize the generalization ability of the model in different fields, and flexibly respond to different fields and tasks, and achieve rapid migration. Furthermore, by introducing a self-supervised optimization mechanism, potential problems in the training data set are automatically learned and discovered, reducing the need for manual labeling, and the SQL query statement generated by prediction is compared with the standard SQL query statement, and self-correction is performed according to the results, thereby improving the accuracy of the generated SQL query statement.
[0126] Figure 6 FIG. 6 is a schematic block diagram of an electronic device 600 according to an embodiment of the present invention. Figure 6 As shown, the electronic device 600 may include a processor 601 and a memory 602. The memory 602 stores computer program instructions for executing the Text2SQL generation method based on self-correction. When the computer program instructions are executed by the processor 601, the electronic device 600 executes the method according to the above combined with Figures 1 to 3 For example, in some embodiments, the electronic device 600 is used for the Text2SQL generation method based on self-correction, which can be referred to in the above embodiments and will not be described in detail here.
[0127] An electronic device provided by an embodiment of the present invention includes: a processor; and a memory storing computer program instructions implemented by a computer for executing a Text2SQL generation method based on self-correction. When the computer program instructions are executed by the processor, the electronic device executes the following Figures 1 to 3 The method described is described.
[0128] A computer-readable storage medium provided by an embodiment of the present invention includes computer program instructions implemented by a computer for executing a Text2SQL generation method based on self-correction. When the computer program instructions are executed by a processor, the following steps are implemented: Figures 1 to 3 The methods described are described.
[0129] It is known to those skilled in the art that the embodiments of the present invention may be implemented as a system, method or computer program product. Therefore, the present invention may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit", "module", "unit" or "system". In addition, in some embodiments, the present invention may also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.
[0130] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive examples) of computer-readable storage media may include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
[0131] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0132] It should be understood that each box in the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine, and these computer program instructions are executed by a computer or other programmable data processing device to produce a device that implements the functions / operations specified in the boxes in the flowchart and / or block diagram.
[0133] Although multiple embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art may conceive of many changes, modifications, and alternatives without departing from the thought and spirit of the present invention. It should be understood that in the process of practicing the present invention, various alternatives to the embodiments of the present invention described herein may be adopted. The appended claims are intended to define the scope of protection of the present invention, and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A Text2SQL generation method based on self-correction, characterized in that: The method includes: Obtaining a first natural language query text; Using a trained structured query statement generation task model to test a first natural language query text, to obtain a first structured query statement corresponding to the first natural language query text, wherein the structured query statement generation task model is obtained by training and learning according to a meta-learning mechanism and a self-supervised learning mechanism; Parsing the first natural language query text using a pre-trained large language model to obtain a query operation mode corresponding to the first natural language query text; When it is determined that the first structured query statement does not meet the query requirement corresponding to the query operation mode, dynamically optimizing the first structured query statement to obtain a second structured query statement; Execute the second structured query statement to obtain the first query result; When it is determined that the first query result does not meet the expected result, the second structured query statement is corrected and optimized according to the first query result to obtain a third structured query statement, and the third structured query statement is output as a target structured query statement corresponding to the first natural language query text.
2. The method according to claim 1, characterized in that The method further includes: When it is determined that the first structured query statement meets the query requirement corresponding to the query operation mode, executing the first structured query statement to obtain a second query result; When it is determined that the second query result does not meet the expected result, the first structured query statement is corrected and optimized according to the second query result to obtain a fourth structured query statement, and the fourth structured query statement is output as a target structured query statement corresponding to the first natural language query text.
3. The method according to claim 1 or 2, characterized in that: Dynamically optimizing the first structured query statement includes: Selecting a query template related to the first natural language query text from a preconfigured template library; The first structured query statement is optimized according to the query template to obtain a second structured query statement.
4. The method according to claim 1, characterized in that: The first query result includes an error type indicating that the second structured query statement exists, and the second structured query statement is corrected and optimized according to the first query result, including: Determining an adjustment strategy corresponding to the error type; The second structured query statement is corrected according to an adjustment strategy corresponding to the error type.
5. The method according to claim 1, characterized in that The structured query statement generation task model obtained through training and learning based on the meta-learning mechanism and the self-supervised learning mechanism includes: Constructing a first labeled training data set, wherein the first training data set includes a plurality of training data pairs, each of the training data pairs including a training natural language query text corresponding to an actual application scenario, and a training structured query statement corresponding to the training natural language query text; Obtaining second training data sets corresponding to each open source data set from a plurality of different open source data sets respectively; According to the meta-learning algorithm, using the plurality of the second training data sets and the first training data sets, pre-training the structured query statement generation task model to be trained to obtain a preliminary structured query statement generation task model; Fine-tuning the preliminary structured query statement generation task model to obtain a current structured query statement generation task model; According to the self-supervised learning algorithm, supervised detection is performed on the output of the current structured query statement generation task model to obtain a trained structured query statement generation task model.
6. The method according to claim 5, characterized in that Before pre-training the structured query statement generation task model to be trained using the plurality of the second training data sets and the first training data sets according to the meta-learning algorithm, the method further includes: Using a pre-trained large language model to perform semantic expansion on each of the second natural language query texts, to obtain a plurality of third natural language query texts with the same semantics and different text expressions corresponding to each of the second natural language query texts; constructing a third training data set according to the correspondence between the plurality of third natural language query texts, the second natural language query texts and the training structured query statements; The third training data set is updated to the first training data set.
7. The method according to claim 5, characterized in that After outputting the target structured query sentence corresponding to the first natural language query text, the method further includes: Using the first natural language query text and the target structured query sentence as a new training data pair; The new training data pair is added to the first training data set to obtain an updated first training data set.
Citation Information
Patent Citations
Method for automatically generating database query statement based on NLP language model
CN116991869A
Language query model construction method, query language acquisition method and related devices
CN117271558A
Text-to-SQL (Structured Query Language) conversion method and device based on automatic process supervision
CN118820286A
Device and method for converting natural language query into SQL query
US20230169074A1
Schema-aware encoding of natural language
US20240248896A1
Cited By
Text2SQL self-correction method based on error pattern perception
CN121542289A
Text2SQL self-correction method based on error pattern perception
CN121542289B