Human-in-the-loop Text2SQL iterative data annotation method fused with large model

By integrating the large model to iterate the data annotation method in loop Text2SQL, using natural language query extension model and structured query statement generation model, combining manual annotation and reinforcement learning mechanism, the problem of inefficient data annotation in complex query scenarios is solved, and high-quality data annotation and model generalization capabilities are improved.

CN120030034AActive Publication Date: 2025-05-23JOINT WARFARE COLLEGE NAT DEFENSE UNIV OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510059365.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-23
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

Existing data annotation methods are inefficient in complex query scenarios, and the annotation data lacks diversity and cannot effectively cover various possible complex query scenarios and complex SQL query statements.

Method used

A method of iterating data annotation in loop Text2SQL is proposed by humans who combine large models, train structured query statements to generate models through pre-constructed initial training data sets, and uses natural language query extension models to perform synonymous diversification expansion, dynamically perform manual annotation and model incremental training until the query results meet expectations.

Benefits of technology

Through human-in-loop technology and reinforcement learning mechanism, the data quality of natural language query text generated by natural language query expansion models in complex query scenarios is improved, the generalization ability of the generation model is optimized, and the accuracy and efficiency of data annotation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030034A_ABST
    Figure CN120030034A_ABST
Patent Text Reader

Abstract

The invention provides a human-in-the-loop Text2SQL iterative data labeling method fused with a large model. The method comprises the following steps: firstly, carrying out training and fine tuning on a generative model by utilizing a pre-constructed manual labeling training sample; performing same-semantic diversification expansion on the training sample by using the expansion model; inputting the plurality of extended training samples into a generative model to obtain a first SQL query statement; for each first SQL query statement, reward and punishment assignment is carried out on model parameters of the expansion model according to an expert analysis result and a reinforcement learning mechanism to determine a target training sample; performing manual annotation on the second SQL query statement, and performing incremental training on the generation model by utilizing a manual annotation result to obtain a second SQL query statement; and when the verification result of the second SQL query statement shows that the query expectation is not met, manually marking the target training sample circularly until the verification result meets the query expectation. The efficiency of data annotation is improved through man-machine cooperation and reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular to a human-in-the-loop Text2SQL iterative data labeling method integrating a large model. Background Art

[0002] With the continuous development and application of artificial intelligence technology, data annotation has become an important part of artificial intelligence technology development. In particular, in the task of converting natural language query (NLQ) to structured query language (SQL), the existing data annotation methods face some challenges and problems. For example, most of the training data sets obtained by current data annotation methods are static and cannot be dynamically annotated. Secondly, the lack of diversity in the annotated data means that the annotated data cannot cover various possible complex query scenarios and complex SQL query statements.

[0003] Therefore, based on the above problems, it is urgent to propose a data labeling method based on human-computer collaboration to improve the labeling efficiency in complex query scenarios. Summary of the invention

[0004] In order to at least solve one or more of the technical problems mentioned above, the present invention proposes a human-in-the-loop Text2SQL iterative data labeling solution integrating a large model in multiple aspects.

[0005] In a first aspect, an embodiment of the present invention provides a human-in-the-loop Text2SQL iterative data labeling method integrating a large model, the method comprising using a pre-constructed initial training data set to train and fine-tune a structured query statement generation model to be trained to obtain a trained structured query statement generation model, the initial training data set comprising a plurality of training sample pairs, each training sample pair comprising a natural language query training sample and a labeled structured query statement corresponding to the natural language query training sample, the labeled structured query statement being obtained through manual labeling; using the trained natural language query expansion model to perform semantically diverse expansion on each natural language query training sample to obtain a plurality of expanded training samples corresponding to each natural language query training sample; inputting the plurality of expanded training samples into the structured query statement generation model to obtain a first structured query statement generation model corresponding to each expanded training sample one by one. query statement; for each first structured query statement, assign rewards and punishments to the model parameters of the natural language query expansion model according to the expert analysis results and the reinforcement learning mechanism, so as to determine that the extended training sample corresponding to the first structured query statement assigned with reward feedback is the target training sample; manually annotate the target training sample to obtain an annotated structured query statement corresponding to the target training sample; according to the target training sample and the annotated structured query statement corresponding to the target training sample, incrementally train the structured query statement generation model to obtain a second structured query statement corresponding to the target training sample; when the verification result of the second structured query statement indicates that the query result of executing the second structured query statement does not meet the query expectation, return to the step of manually annotating the target training sample, and continue processing until the verification result indicates that the query result of executing the second structured query statement meets the query expectation and stops the loop.

[0006] In some embodiments, for each first structured query statement, rewards and penalties are assigned to the model parameters of the natural language query expansion model according to the expert analysis results and the reinforcement learning mechanism, including: determining whether the first structured query statement meets the query requirements; when it is determined that the first structured query statement meets the query requirements, assigning a penalty feedback value to the model parameters of the natural language query expansion model; or, when it is determined that the first structured query statement does not meet the query requirements, rewarding feedback values ​​are assigned to the model parameters of the natural language query expansion model to determine that the extended training sample corresponding to the first structured query statement that does not meet the query requirements is the target training sample; using the natural language query expansion model to perform semantically diverse expansion on the target training sample to obtain multiple target training samples to be labeled corresponding to the target training sample; updating the multiple target training samples to be labeled as new target training samples.

[0007] In some embodiments, incremental training is performed on a structured query statement generation model based on target training samples and labeled structured query statements corresponding to the target training samples, including: constructing an incremental training data set based on new target training samples and labeled structured query statements corresponding to the new target training samples; and incrementally training the structured query statement generation model using the incremental training data set to obtain an optimized structured query statement generation model and a second structured query statement corresponding to each new target training sample.

[0008] In some embodiments, when the verification result of the second structured query statement indicates that the query result of executing the second structured query statement does not meet the query expectation, return to the step of manually labeling the target training sample, and continue processing until the verification result indicates that the query result of executing the second structured query statement meets the query expectation and stops the loop, including: using a professional database to query the second structured query statement to obtain a query result corresponding to the second structured query statement; verifying whether the query result meets the query expectation to obtain a verification result; when the verification result indicates that the query result does not meet the query expectation, marking the target training sample corresponding to the second structured query statement as a new target training sample; returning to the step of manually labeling the target training sample and continuing processing until the verification result indicates that the second structured query statement meets the query expectation.

[0009] In some embodiments, manually labeling the target training samples includes: for each target training sample, obtaining a labeled structured query statement corresponding to the target training sample; and establishing a mapping relationship between the target training sample and the labeled structured query statement corresponding to the target training sample.

[0010] In some embodiments, the natural language query expansion model is a deep learning model based on the Transformer structure.

[0011] In some embodiments, the structured query statement generation model is a deep learning model based on the Transformer structure.

[0012] The human-in-the-loop Text2SQL iterative data annotation method for integrating a large model provided by the present invention comprises the following steps: using a pre-constructed initial training data set to train and fine-tune a structured query statement generation model to be trained, so as to obtain a trained structured query statement generation model (i.e., an SQL generation model); then using the trained natural language query expansion model (i.e., an NLQ expansion model) to perform semantically diverse expansion on each natural language query training sample, so as to obtain a plurality of expanded training samples corresponding to each natural language query training sample; inputting the plurality of expanded training samples into the SQL generation model, so as to obtain a first SQL query statement corresponding to each expanded training sample one by one; for each first SQL query statement, performing a plurality of semantically diverse expansions on each natural language query training sample according to an expert analysis result and a reinforcement learning mechanism; The model parameters of the LQ extension model are rewarded and punished to determine that the extended training sample corresponding to the first SQL query statement assigned with reward feedback is the target training sample; the target training sample is manually annotated to obtain the annotated SQL query statement corresponding to the target training sample; according to the target training sample and the annotated SQL query statement corresponding to the target training sample, the SQL generation model is incrementally trained to obtain the second SQL query statement corresponding to the target training sample; when the verification result of the second SQL query statement indicates that the query result of executing the second SQL query statement does not meet the query expectation, the step of manually annotating the target training sample is returned, and the processing is continued until the verification result indicates that the query result of executing the second SQL query statement meets the query expectation and the loop is stopped. The technical solution provided by the present invention dynamically iterates the annotation through the human-in-the-loop technology, and improves the data quality of the natural language query text generated by the natural language query extension model for complex query scenarios according to the expert analysis results and the reinforcement learning mechanism; the extended training samples with the same semantics are manually annotated, and the generation model is incrementally trained using the manual annotation results, so that the generation model can be further optimized and its generalization ability to handle complex queries can be improved. At the same time, the second SQL query statement generated in the incremental training is verified to improve the accuracy of data annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0014] Figure 1 is a schematic flow chart of a human-in-the-loop Text2SQL iterative data annotation method 100 integrating a large model according to an embodiment of the present invention;

[0015] Figure 2 shows a schematic flow chart of step S104 according to some other embodiments of the present invention;

[0016] Figure 3A schematic flow chart of a human-in-the-loop Text2SQL iterative data annotation method 300 integrating a large model is shown in some other embodiments of the present invention;

[0017] Figure 4 is a structural diagram of a human-in-the-loop Text2SQL iterative data annotation device 400 integrating a large model according to an embodiment of the present invention;

[0018] Figure 5 A schematic block diagram of an electronic device 500 according to an embodiment of the present invention is shown.

[0019] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION

[0020] In order to better understand and explain the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings. The present invention is not limited to these specific embodiments. On the contrary, modifications or equivalent substitutions made to the present invention should all be included in the scope of the claims of the present invention.

[0021] It should be noted that numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present invention can also be implemented without these specific details. In the multiple specific embodiments given below, the principles, structures and components well known in the art are not described in detail in order to highlight the main purpose of the present invention.

[0022] The human-in-the-loop Text2SQL iterative data annotation method for integrating a large model provided by the present invention can be executed by a computer device, wherein the computer device can be a terminal or a server. The terminal can be a terminal device such as a smart phone, a tablet computer, a laptop computer, a touch screen, a personal computer (PC), a personal digital assistant (PDA), etc., and the terminal can also include a client. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0023] Please refer to Figure 1 , Figure 1 1 is a schematic flow chart of a human-in-the-loop Text2SQL iterative data annotation method 100 integrating a large model according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0024] In step S101, a structured query statement generation model to be trained is trained and fine-tuned using a pre-constructed initial training data set to obtain a trained structured query statement generation model. The initial training data set includes multiple training sample pairs, each training sample pair includes a natural language query training sample and an annotated structured query statement corresponding to the natural language query training sample, and the annotated structured query statement is obtained through manual annotation.

[0025] In the above steps, the natural language query training sample is a sample of natural language query text. Natural language query text refers to a query request for a database proposed by a user in natural language form. For example, it can represent an NLP sample. The annotated structured query statement corresponding to the natural language query training sample refers to a structured query statement (i.e., SQL query statement) obtained by manually annotating the NLP sample. The training sample pair can be represented as (NLP, SQL). The initial training data set includes multiple training sample pairs, and the initial training data set can be represented as {(NLP 0 , SQL 0 ), (NLP 1 , SQL 1 ), (NLP 2 , SQL 2 ), (NLP 3 , SQL 3 ), ..., (NLP N , SQL N )}. Natural language query training samples are manually selected by domain experts. The SQL query statements corresponding to the natural language query training samples are manually written and annotated. The annotators can become experts in the relevant fields.

[0026] For example, in a database query scenario in a certain scenario, the expert selected the following typical natural language queries: "Query the total number of employees in each department", "Get the comparison of the number of employees in the sales department and the financial department", etc. For these natural language queries, the corresponding SQL query statements are marked. For "Query the total number of employees in each department", the marked SQL query statement is:

[0027] SELECT department_name,COUNT(employee_id)AS total_count

[0028] FROM employees

[0029] GROUP BY department_name;

[0030] For "obtain the comparison of the number of employees in the sales department and the finance department", the SQL query statement is:

[0031] SELECT department_name,COUNT(employee_id)AS total_count

[0032] FROM employees

[0033] WHERE department_name IN('sales','finance')

[0034] GROUP BY department_name;

[0035] These labeled data are organized into a standard format as a preliminary training data set.

[0036] The structured query statement generation model can be a deep learning model based on the Transformer structure. For example, a large language model, an open source large model, etc. Basic open source large models, such as Qwen and Lamma. For example, based on the constructed initial training data set, a standard large language model fine-tuning method is used to train the basic open source large model to learn the mapping rules from NLQ to SQL queries. During the training phase, the model can gradually master the grammatical structure, semantic information, and generation rules of SQL statements of natural language.

[0037] Given a natural language query training sample "query the total number of employees in each department", after training, the structured query statement generation model generates the corresponding SQL query statement as follows:

[0038] SELECT department_name,COUNT(employee_id)AS total_count

[0039] FROM employees

[0040] GROUP BY department_name;

[0041] After preliminary training, the structured query generation model can generate SQL query statements based on the given natural language query training samples. During the training phase, although the structured query generation model can basically generate SQL query statements, due to the limited number of training data sets, some of the generated SQL query statements may contain syntax errors, incomplete query logic, or be inconsistent with actual business needs. For example, for more complex queries, such as "obtain a comparison of the number of employees in the sales department and the finance department", the structured query generation model in training may generate the following erroneous SQL query statements:

[0042] SELECT department_name,employee_count

[0043] FROM employees

[0044] WHERE department_name IN('sales','finance')

[0045] GROUP BY department_name;

[0046] The SQL query statement lacks the COUNT() function, which results in the inability to correctly count the number of employees. This causes the query result of the SQL query statement to not meet the query requirements.

[0047] In step S102, the trained natural language query expansion model is used to perform semantic diversification expansion on each natural language query training sample to obtain a plurality of expanded training samples corresponding to each natural language query training sample.

[0048] In the above steps, the natural language query expansion model is used to perform homosemantic diversification expansion on the natural language query training samples. The natural language query expansion model can be a deep learning model based on the Transformer structure, such as a large language model, an open source large model, etc. The homosemantic diversification expansion is to generate diversified natural language query samples through the natural language query expansion model. These natural language query samples can cover different query modes, such as single-table query, multi-table join, nested query, etc., and involve scenarios of various query complexities. The generated natural language query samples will be sent to the structured query statement generation model trained in the previous step to generate SQL query statements, and then the structured query statement generation model will be optimized through expert review and reinforcement learning mechanism.

[0049] Each natural language query training sample is diversified and expanded to obtain multiple extended training samples corresponding to each natural language query training sample, for example, natural language query (NLQ): "Query the comparison of the number of employees in each department in different quarters between January 1, 2023 and December 31, 2023, and count the difference in the number of employees in each department each quarter." By performing the same semantic diversified expansion, the generated self-extended training samples can cover more similar query requirements, such as: "Query the comparison of the number of employees in all departments in each quarter in 2023, and calculate the difference", "Obtain the changes in employees in each department in each quarter in 2023, and display the difference in the number of people", "Count the number of employees in each department in each quarter according to the time period, and calculate the difference in the number of employees in each quarter".

[0050] In step S103, a plurality of extended training samples are input into a structured query statement generation model to obtain a first structured query statement corresponding to each of the extended training samples.

[0051] In step S104, for each first structured query statement, rewards and punishments are assigned to the model parameters of the natural language query expansion model according to the expert analysis results and the reinforcement learning mechanism, so as to determine that the expanded training sample corresponding to the first structured query statement assigned with reward feedback is the target training sample.

[0052] In the above steps, the expert can review the first structured query statement according to the actual query requirements of the natural language query, and introduce a reinforcement learning mechanism for supervision and recording. The first structured query statement is an SQL query statement corresponding to the extended training sample generated by the structured query statement generation model. If the generated SQL query statement meets the query requirements, the reinforcement learning mechanism will assign a penalty to reduce the probability of generating the SQL query statement, prompting the model to improve its generation ability. If the generated SQL query statement cannot meet the query requirements, the reinforcement learning mechanism will assign a reward to encourage the natural language query expansion model to continue to generate high-quality extended training samples that meet the query requirements. The extended training sample is a natural language query text with the same semantics as the natural language training sample, but with a different language expression.

[0053] The target training samples refer to natural language query texts that contain complex query logic (such as cross-table query, multi-table join, nested query, etc.) and cannot generate correct structured query statements during the training and fine-tuning phase of the structured query statement generation model. The target training samples can also be called "query difficulty samples".

[0054] In step S105, the target training sample is manually labeled to obtain a labeled structured query statement corresponding to the target training sample.

[0055] In step S106, incremental training is performed on a structured query statement generation model according to the target training sample and the annotated structured query statement corresponding to the target training sample to obtain a second structured query statement corresponding to the target training sample.

[0056] In the above steps, the target training samples are manually annotated by annotation experts. For example, the annotation experts can write SQL query statements corresponding to the target training samples according to actual business needs and database structure, that is, obtain SQL query statements corresponding to the target training samples, and then establish a mapping relationship between the target training samples and the SQL query statements, that is, complete the annotation. The SQL query statements corresponding to the target training samples become annotated structured query statements.

[0057] After the labeling experts manually label multiple target training samples one by one, a new training data set can be obtained. The new training data set is used to perform incremental training on the structured query statement generation model to generate SQL query statements corresponding to the multiple target training samples one by one. The SQL query statement corresponding to the target training sample is called the second structured query statement.

[0058] Through the above incremental training process, the processing capability of the structured query statement generation model for complex queries can be effectively improved.

[0059] In step S107, when the verification result of the second structured query statement indicates that the query result of executing the second structured query statement does not meet the query expectation, return to the step of manually labeling the target training sample and continue processing until the verification result indicates that the query result of executing the second structured query statement meets the query expectation and stops the loop.

[0060] In the above steps, the second structured query statement generated during the incremental training process is verified using a business database or a professional database. If the query result of the second structured query statement meets the user's query expectations, the target training sample and the second structured query statement are output as the annotation result. If the query result of the second structured query statement does not meet the user's query expectations, the step of annotating the target training sample is returned, and the annotation expert analyzes and corrects the second generated SQL query statement to obtain a corrected second generated SQL query statement, and then the corrected second generated SQL query statement is verified using a professional database. If the query result of the second generated SQL query statement meets the query expectations, a mapping relationship between the target training sample and the corrected second generated SQL query statement is established to obtain a new training data pair.

[0061] The human-in-the-loop Text2SQL iterative data annotation method that integrates a large model proposed in the present invention uses a pre-constructed initial training data set to train a structured query statement generation model, and uses a natural language query expansion model to perform semantic expansion on natural language query training samples, which can improve the generalization ability of complex queries; then, based on expert analysis results and reinforcement learning mechanisms, the data quality of natural language query texts generated by the natural language query expansion model for complex query scenarios is improved, and manual annotation is performed on extended training samples with the same semantics. Incremental training is performed using extended training samples, which can improve the query statement generation model's processing ability for complex query statements, and the accuracy and consistency of data annotation results can be improved by verifying the second SQL query statement. In summary, through the above-mentioned scheme, based on the human-in-the-loop method in the data annotation process, reinforcement learning is applied to achieve iterative annotation, which can effectively improve the annotation efficiency and accuracy of data annotation.

[0062] Figure 2 Schematic flow charts of step S104 of other embodiments of the present invention are shown. It can be understood that Figure 2 The steps S2041-S2045 shown are a specific implementation of the aforementioned step S104, so the aforementioned Figure 1 The related features described in can be similarly applied here. Figure 2 As shown, the method includes:

[0063] In step S2041, it is determined whether the first structured query statement meets the query requirement.

[0064] In step S2042, when it is determined that the first structured query statement meets the query requirement, a penalty feedback value is assigned to a model parameter of the natural language query expansion model.

[0065] In step S2043, when it is determined that the first structured query statement does not meet the query requirement, a reward feedback value is performed on the model parameter of the natural language query expansion model to determine the expanded training sample corresponding to the first structured query statement that does not meet the query requirement as the target training sample;

[0066] In step S2044, the target training sample is expanded in a homosemantic manner using a natural language query expansion model to obtain a plurality of target training samples to be labeled corresponding to the target training sample.

[0067] In step S2045, a plurality of target training samples to be labeled are updated to be new target training samples.

[0068] In the above steps, it is determined whether the SQL query statement meets the query requirements, and the query requirements can be evaluated from multiple angles. For example, the query requirements may include query targets, filter conditions, aggregate functions, sorting and grouping, and other factors used to implement the query. The query target can be a clear data object that needs to be queried, such as querying sales data within a specific time period, or querying product inventory in a certain area. The filter conditions are conditions used to limit the scope of the query results, such as time range, geographical restrictions, product categories, etc. By narrowing the scope of the query results, the relevance of the query results can be ensured. Whether the aggregate function matches the natural language query, for example, the natural language query text can be analyzed to determine the SQL aggregate function for statistical analysis of the data. SQL aggregate functions include but are not limited to sum, count, average, etc. Sorting and grouping can sort or group the results as needed for easy analysis or display.

[0069] After clarifying the above query requirements, determine whether the SQL query statement meets the query requirements. You can determine whether the various components of the SQL query statement, such as SELECT, FROM, WHERE, GROUP BY, ORDER BY, etc., are used correctly to ensure the accuracy and efficiency of the query results.

[0070] Use the natural language query expansion model to generate multiple natural language query samples with the same semantics (i.e., extended training samples). These natural language query samples can cover a variety of scenarios from simple query scenarios to complex query scenarios (such as cross-table joins, nested queries, etc.). Input the generated multiple natural language query samples with the same semantics into the trained structured query statement generation model to generate SQL query statements corresponding to the extended training samples.

[0071] According to the query requirements of natural language queries, annotation experts review the SQL query statements corresponding to the extended training samples, and introduce a reinforcement learning mechanism for supervision and recording. If the SQL query statements corresponding to the extended training samples meet the query requirements, the reinforcement learning mechanism will impose penalties, which can be to assign penalty feedback values ​​to the model parameters of the natural language query expansion model. If the SQL query statements corresponding to the extended training samples do not meet the query requirements, the reinforcement learning mechanism will reward, that is, to assign reward feedback values ​​to the model parameters of the natural language query expansion model, so as to determine that the extended training samples corresponding to the SQL query statements that do not meet the query requirements are the target training samples. Through the reinforcement learning mechanism, the natural language query expansion model is encouraged to continue to generate natural language query texts that meet the query requirements.

[0072] Through the above-mentioned reinforcement learning mechanism, the ability of the natural language query expansion model to generate natural language query question samples that meet complex query scenarios can be improved.

[0073] Assume that the natural language query text is: "Query the comparison of the number of employees in each department in different quarters between January 1, 2023 and December 31, 2023, and count the difference in the number of employees in each department each quarter". This natural language query text involves the following key points, such as the need to obtain data from multiple tables (such as employees, departments, orders), which involves cross-table query; the query condition involves date range filtering; the number of employees in each quarter needs to be counted and compared, which involves complex aggregation operations.

[0074] The SQL query statement generated by the structured query statement generation model is:

[0075] SELECT department_name,quarter,COUNT(employee_id)AS total_count

[0076] FROM employees

[0077] WHERE hi re_date BETWEEN'2023-01-01'AND'2023-12-31'

[0078] GROUP BY department_name,quarter;

[0079] The above-generated SQL query statement is considered not to meet the query requirements. For example, the SQL query statement does not involve the demand expression parameters corresponding to the "difference in the number of employees in each department"; the SQL query statement does not involve the demand expression parameters for calculating the "difference in the number of employees in each department each quarter".

[0080] The annotation experts found that the generated SQL query statement did not meet the query requirements, especially because it did not provide the requirement expression parameter of "statistics of the difference in the number of people in each department". The reinforcement learning mechanism can reward it and provide reward feedback values ​​so that the natural language query samples generated by the natural language query expansion model can cover more similar query requirements. When the generated SQL query statement does not meet the actual query requirements of the natural language query text, the model parameters of the natural language query expansion model can be rewarded through the reinforcement learning mechanism to encourage the natural language query expansion model to generate multiple natural language query texts to cover the query requirements, thereby improving the adaptability of complex query scenarios.

[0081] The human-in-the-loop Text2SQL iterative data labeling method integrating a large model provided by the present invention dynamically increases training samples through expert review and reinforcement learning mechanism. While ensuring data quality, it reduces the time cost of constructing natural language expression samples of complex queries and the corresponding SQL query statements, and significantly improves data labeling efficiency.

[0082] Figure 3 3 is a schematic flow chart of a human-in-the-loop Text2SQL iterative data annotation method 300 integrating a large model according to an embodiment of the present invention. It can be understood that Figure 3 The steps S3071-S3075 shown are a specific implementation of the aforementioned step S107, so the aforementioned Figure 1 The related features described in can be similarly applied here. Figure 3 As shown, the method includes:

[0083] In step S301, a structured query statement generation model to be trained is trained and fine-tuned using a pre-constructed initial training data set to obtain a trained structured query statement generation model. The initial training data set includes multiple training sample pairs, each training sample pair includes a natural language query training sample and an annotated structured query statement corresponding to the natural language query training sample, and the annotated structured query statement is obtained through manual annotation.

[0084] In step S302, each natural language query training sample is diversifiedly expanded using the trained natural language query expansion model to obtain a plurality of expanded training samples corresponding to each natural language query training sample.

[0085] In step S303, a plurality of extended training samples are input into a structured query statement generation model to obtain a first structured query statement corresponding to each of the extended training samples.

[0086] In step S304, for each first structured query statement, rewards and punishments are assigned to the model parameters of the natural language query expansion model according to the expert analysis results and the reinforcement learning mechanism, so as to determine that the expanded training sample corresponding to the first structured query statement assigned with reward feedback is the target training sample.

[0087] In step S305, the target training sample is manually labeled to obtain a labeled structured query statement corresponding to the target training sample.

[0088] In step S306, incremental training is performed on the structured query statement generation model according to the target training sample and the annotated structured query statement corresponding to the target training sample to obtain a second structured query statement corresponding to the target training sample.

[0089] In step S3071, the second structured query statement is queried using a professional database to obtain a query result corresponding to the second structured query statement.

[0090] In step S3072, it is verified whether the query result meets the query expectation and a verification result is obtained.

[0091] In step S3073, when the verification result indicates that the query result does not meet the query expectation, the target training sample corresponding to the second structured query statement is marked as a new target training sample, and then step S3074 is executed.

[0092] In step S3074, the process returns to step S105 to continue processing until the verification result indicates that the second structured query statement meets the query expectation, and the loop is stopped.

[0093] In step S3075, when the verification result indicates that the query result meets the query expectation, it means that the generation capability of the optimized structured query statement generation model obtained through incremental training does not need to be dynamically adjusted. The optimized structured query statement generation model can be directly applied to the business scenario for continued use.

[0094] In the above steps, query expectation refers to the expected result set according to the query requirements before executing the database query. For example, the result obtained by executing a certain SQL query. Query expectations may include but are not limited to: the expected query results can accurately reflect the information stored in the database, without omissions or errors; the expected query can return all data records that meet the query conditions, no more and no less; the logic of the expected query is consistent with the business logic, for example, when doing statistical aggregation, the aggregation function is used correctly, and the grouping and sorting logic is correct. Verification can be done by comparative analysis, such as comparing the actual query results with the query expectations, analyzing whether there are differences and their causes.

[0095] The SQL query statements generated by the incrementally trained structured query statement generation model are connected to a professional database for actual query verification. If the SQL query statement can be executed correctly and the query results returned meet the expected query results, it means that the generation capability of the model has been improved and can continue to be used. If the SQL query statement can be executed correctly and the query results returned do not meet the expected query results, the target training samples corresponding to the SQL query statement are marked as "difficult samples". These "difficult samples" will be sent back to the step of manually annotating the target training samples, and the target training samples will be corrected and annotated again by the annotation experts. The annotation experts can review the SQL query statement and its corresponding natural language query text, find out the cause of the problem and correct it to generate the correct SQL query statement. Subsequently, these corrected SQL query statements and their corresponding natural language query texts can be used as new incremental data sets and reused for incremental training of the structured query statement generation model. Through repeated corrections and feedback, the ability of the structured query statement generation model to handle complex query scenarios can be improved, meet diverse query requirements, and generate SQL query statements accurately and efficiently.

[0096] For example, the generated SQL query statement is "query the comparison of the number of employees in each department in a certain quarter". After executing the query, the query result returned does not match the actual data, indicating that there is a problem with the query logic. At this time, the generated SQL query statement is marked as a "difficult sample" and returns to the step of manually annotating the target training sample. The annotation expert reviewed the generated SQL query statement again and found that the problem was that the screening condition for quarterly data was incorrectly expressed. The annotation expert corrected the generated SQL query statement and obtained the corrected SQL query statement:

[0097] SELECT department_name,COUNT(employee_id)AS total_count

[0098] FROM employees

[0099] WHERE quarter = 'correct quarter'

[0100] GROUP BY department_name;

[0101] The human-in-the-loop Text2SQL iterative data annotation method that integrates a large model proposed in the present invention overcomes the shortcomings of the prior art in processing complex queries, multi-domain adaptability, and data annotation efficiency through reinforcement learning and dynamic feedback mechanisms, and effectively improves the data quality of natural language query text generated by the natural language query expansion model for complex query scenarios. Manual annotation is performed on extended training samples with the same semantics, and the generation model is incrementally trained using the manual annotation results, which can further optimize the generation model and improve its generalization ability to process complex queries. At the same time, the second SQL query statement generated in the incremental training is verified to improve the accuracy of data annotation.

[0102] In summary, the human-in-the-loop Text2SQL iterative data labeling method that integrates a large model significantly improves the flexibility, adaptability, and accuracy of data labeling.

[0103] Please refer to Figure 4 , Figure 4 4 is a schematic diagram of a human-in-the-loop Text2SQL iterative data annotation device 400 according to an embodiment of the present invention. Figure 4 As shown, the device comprises:

[0104] The model training and fine-tuning module 401 is configured to use a pre-constructed initial training data set to train and fine-tune the structured query statement generation model to be trained to obtain a trained structured query statement generation model. The initial training data set includes multiple training sample pairs, each training sample pair includes a natural language query training sample and an annotated structured query statement corresponding to the natural language query training sample, and the annotated structured query statement is obtained through manual annotation.

[0105] The sample diversification expansion module 402 is configured to perform semantic diversification expansion on each natural language query training sample using the trained natural language query expansion model to obtain multiple expanded training samples corresponding to each natural language query training sample.

[0106] The query statement generation module 403 is configured to input a plurality of extended training samples into the structured query statement generation model to obtain a first structured query statement corresponding to each extended training sample.

[0107] The expert reinforcement learning module 404 is configured to assign rewards and punishments to the model parameters of the natural language query expansion model for each first structured query statement according to the expert analysis results and the reinforcement learning mechanism, so as to determine that the expanded training sample corresponding to the first structured query statement to which the reward feedback is assigned is the target training sample.

[0108] The manual labeling module 405 is configured to manually label the target training sample to obtain a labeled structured query statement corresponding to the target training sample.

[0109] The model incremental training module 406 is configured to perform incremental training on the structured query statement generation model according to the target training sample and the annotated structured query statement corresponding to the target training sample to obtain a second structured query statement corresponding to the target training sample.

[0110] The query statement verification module 407 is configured to return to the step of manually labeling the target training sample when the verification result of the second structured query statement indicates that the query result of executing the second structured query statement does not meet the query expectation, and continue processing until the verification result indicates that the query result of executing the second structured query statement meets the query expectation and stops the loop.

[0111] In some embodiments, the expert reinforcement learning module 404 is configured to determine whether the first structured query statement meets the query requirements; when it is determined that the first structured query statement meets the query requirements, a penalty feedback value is assigned to the model parameters of the natural language query expansion model; or, when it is determined that the first structured query statement does not meet the query requirements, a reward feedback value is assigned to the model parameters of the natural language query expansion model to determine that the extended training sample corresponding to the first structured query statement that does not meet the query requirements is the target training sample; use the natural language query expansion model to perform semantically diverse expansion on the target training sample to obtain multiple target training samples to be labeled corresponding to the target training sample; update the multiple target training samples to be labeled as new target training samples.

[0112] In some embodiments, the manual labeling module 405 is further configured to obtain, for each target training sample, a labeled structured query statement corresponding to the target training sample; and establish a mapping relationship between the target training sample and the labeled structured query statement corresponding to the target training sample.

[0113] In some embodiments, the model incremental training module 406 is configured to construct an incremental training data set based on new target training samples and annotated structured query statements corresponding to the new target training samples; and use the incremental training data set to incrementally train the structured query statement generation model to obtain an optimized structured query statement generation model and a second structured query statement corresponding to each new target training sample.

[0114] In some embodiments, the query statement verification module 407 can also be configured to use a professional database to query the second structured query statement to obtain a query result corresponding to the second structured query statement; verify whether the query result meets the query expectation to obtain a verification result; when the verification result indicates that the query result does not meet the query expectation, mark the target training sample corresponding to the second structured query statement as a new target training sample; return to the step of manually labeling the target training sample and continue processing until the verification result indicates that the second structured query statement meets the query expectation.

[0115] In some embodiments, the natural language query expansion model is a deep learning model based on the Transformer structure.

[0116] In some embodiments, the structured query statement generation model is a deep learning model based on the Transformer structure.

[0117] The human-in-the-loop Text2SQL iterative data labeling device that integrates a large model proposed in the present invention can ensure that the data set after each iteration is reviewed by experts through the reasoning ability of the model combined with expert labeling, thereby effectively improving the accuracy and reliability of SQL query statements. Combined with reinforcement learning and dynamic iterative labeling methods, the structured query statement generation model can be continuously optimized when processing new complex query patterns, significantly improving the adaptability of the structured query statement generation model in changing scenarios.

[0118] Figure 5 FIG. 5 is a schematic block diagram of an electronic device 500 according to an embodiment of the present invention. Figure 5 As shown, the electronic device 500 may include a processor 501 and a memory 502. The memory 502 stores computer program instructions for executing the human-in-the-loop Text2SQL iterative data labeling method for integrating a large model. When the computer program instructions are executed by the processor 501, the electronic device 500 executes the method according to the above combined method. Figures 1 to 3 For example, in some embodiments, the electronic device 500 is used to implement the human-in-the-loop Text2SQL iterative data annotation method integrating a large model, which can be referred to in the above embodiments and will not be described in detail here.

[0119] An electronic device provided by an embodiment of the present invention includes: a processor; and a memory storing computer program instructions implemented by a computer for executing a human-in-the-loop Text2SQL iterative data labeling method for integrating a large model. When the computer program instructions are executed by the processor, the electronic device executes the following Figures 1 to 3 Describe the method.

[0120] A computer-readable storage medium provided by an embodiment of the present invention includes computer program instructions implemented by a computer for executing a human-in-the-loop Text2SQL iterative data labeling method for integrating a large model. When the computer program instructions are executed by a processor, the following is implemented: Figures 1 to 3 Describe the method.

[0121] It is known to those skilled in the art that the embodiments of the present invention may be implemented as a system, method or computer program product. Therefore, the present invention may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit", "module", "unit" or "system". In addition, in some embodiments, the present invention may also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.

[0122] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive examples) of computer-readable storage media may include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.

[0123] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0124] It should be understood that each box in the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine, and these computer program instructions are executed by a computer or other programmable data processing device to produce a device that implements the functions / operations specified in the boxes in the flowchart and / or block diagram.

[0125] Although multiple embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art may conceive of many changes, modifications, and alternatives without departing from the thought and spirit of the present invention. It should be understood that in the process of practicing the present invention, various alternatives to the embodiments of the present invention described herein may be adopted. The appended claims are intended to define the scope of protection of the present invention, and therefore cover equivalents or alternatives within the scope of these claims.

Claims

1. A human-in-the-loop Text2SQL iterative data annotation method integrating a large model, characterized in that: The method includes: Using a pre-constructed initial training data set to train and fine-tune a structured query statement generation model to be trained, to obtain a trained structured query statement generation model, wherein the initial training data set includes a plurality of training sample pairs, each training sample pair includes a natural language query training sample and an annotated structured query statement corresponding to the natural language query training sample, wherein the annotated structured query statement is obtained through manual annotation; Using the trained natural language query expansion model, each of the natural language query training samples is expanded with the same semantic diversity to obtain a plurality of expanded training samples corresponding to each of the natural language query training samples; Inputting the plurality of the extended training samples into the structured query statement generation model to obtain a first structured query statement corresponding to each of the extended training samples; For each of the first structured query statements, rewards and punishments are assigned to the model parameters of the natural language query expansion model according to the expert analysis results and the reinforcement learning mechanism, so as to determine the expanded training samples corresponding to the first structured query statements assigned with reward feedback as the target training samples; Manually labeling the target training sample to obtain a labeled structured query statement corresponding to the target training sample; According to the target training sample and the annotated structured query statement corresponding to the target training sample, incrementally training the structured query statement generation model to obtain a second structured query statement corresponding to the target training sample; When the verification result of the second structured query statement indicates that the query result of executing the second structured query statement does not meet the query expectation, return to the step of manually labeling the target training sample, and continue processing until the verification result indicates that the query result of executing the second structured query statement meets the query expectation and stops the loop.

2. The method according to claim 1, characterized in that For each of the first structured query statements, rewards and penalties are assigned to the model parameters of the natural language query expansion model according to the expert analysis results and the reinforcement learning mechanism, including: Determining whether the first structured query statement meets the query requirement; When it is determined that the first structured query statement meets the query requirement, a penalty feedback value is assigned to a model parameter of the natural language query expansion model; or, When it is determined that the first structured query statement does not meet the query requirement, a reward feedback value is given to a model parameter of the natural language query expansion model to determine an expanded training sample corresponding to the first structured query statement that does not meet the query requirement as a target training sample; Using the natural language query expansion model to perform semantically diverse expansion on the target training sample, to obtain a plurality of to-be-annotated target training samples corresponding to the target training sample; The multiple target training samples to be labeled are updated to be new target training samples.

3. The method according to claim 2, characterized in that Incrementally training the structured query statement generation model according to the target training sample and the annotated structured query statement corresponding to the target training sample, including: Constructing an incremental training data set according to the new target training sample and the annotated structured query statement corresponding to the new target training sample; The structured query statement generation model is incrementally trained using the incremental training data set to obtain an optimized structured query statement generation model and a second structured query statement corresponding to each of the new target training samples.

4. The method according to claim 1, characterized in that: When the verification result of the second structured query statement indicates that the query result of executing the second structured query statement does not meet the query expectation, returning to the step of manually labeling the target training sample, and continuing the process until the verification result indicates that the query result of executing the second structured query statement meets the query expectation, the loop is stopped, including: Using a professional database to query the second structured query statement to obtain a query result corresponding to the second structured query statement; Verify whether the query result meets the query expectation, and obtain a verification result; When the verification result indicates that the query result does not meet the query expectation, marking the target training sample corresponding to the second structured query statement as a new target training sample; Return to the step of manually labeling the target training sample, and continue processing until the verification result indicates that the query result of executing the second structured query statement meets the query expectation and then stops the loop.

5. The method according to claim 1, characterized in that Manually labeling the target training sample includes: For each of the target training samples, a labeled structured query statement corresponding to the target training sample is obtained; and a mapping relationship between the target training sample and the labeled structured query statement corresponding to the target training sample is established.

6. The method according to claim 1, characterized in that The natural language query expansion model is a deep learning model based on the Transformer structure.

7. The method according to claim 1, characterized in that The structured query statement generation model is a deep learning model based on the Transformer structure.

Citation Information

Patent Citations

  • Training method of statement generation model, statement generation method, system and equipment

    CN117992791A

  • Method and system for generating nl2sql training set based on large model and suitable for all industries

    CN118964579A

  • Device and method for converting natural language query into SQL query

    US20230169074A1

  • System, method, and computer program for augmenting multi-turn text-to-SQL datasets with self-play

    US20240078230A1

  • Database Query Generation from Natural Language Statements

    US20240273090A1