Training data generation method and device, equipment, storage medium and program product

By synthesising data augmentation and error type attribution reply data of the Text2SQL evaluation model through question-and-answer data enhancement and error type attribution reply data, the problem of inefficient training data generation in the prior art is solved, and training data generation is efficiently covered by real business scenarios.

CN120067685APending Publication Date: 2025-05-30BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510134286.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is inefficient, time-consuming and labor-intensive when building training data of Text2SQL evaluation model, and it is difficult to cover the coverage of real online business scenarios.

Method used

By obtaining the Q&A pair data, the problem data is enhanced based on the Q&A pair attribute information of the seed data, and the first target Q&A pair data is synthesized; based on the error types of the Q&A pair data and/or the first target Q&A pair data, attribution reply data is generated, and the second target Q&A pair data is synthesized to form the target training data.

Benefits of technology

The automatic generation of target training data is achieved, the efficiency of training data generation is improved, the real online business scenarios can be better covered, and the detection ability of the evaluation model under different error types is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067685A_ABST
    Figure CN120067685A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a training data generation method and device, equipment, a storage medium and a program product, and the method comprises the steps: obtaining question and answer pair data; on the basis of attribute information of question and answer pair seed data in the question and answer pair data, question data in the question and answer pair seed data is enhanced to obtain question enhanced data, and first target question and answer pair data is synthesized by using the question enhanced data and reply data in the question and answer pair seed data; on the basis of the error type of the question-answer pair data and / or the first target question-answer pair data, generating attribution reply data matched with the error type, and synthesizing second target question-answer pair data by utilizing the attribution reply data and question data and / or question enhancement data in the question-answer pair data; and determining the first target question and answer pair data and the second target question and answer pair data as target training data. By implementing the technical scheme, the generation quality, the generation efficiency, the diversity and the data volume of the target training data are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and particularly to a method, apparatus, device, storage medium, and program product for generating training data. Background Art

[0002] With the increasing development of computer technology, it has become possible to automatically generate query statements using a model. For example, a Text2SQL model can generate corresponding query statements according to text description information. In order to evaluate the query statements generated by the Text2SQL model, an evaluation model is required to evaluate whether the query statements are correct. For the evaluation model, not only the data of the Text2SQL model but also the evaluation data of the Text2SQL model are needed.

[0003] Currently, when constructing the training data for the evaluation model of Text2SQL, different proportions of error types are mainly constructed manually, and the explanatory information of the incorrect query statements in the negative samples is manually labeled. However, this method is inefficient, time-consuming and laborious, and the manually labeled data is difficult to cover the coverage of the real online business scenarios. Summary of the Invention

[0004] In view of this, the present disclosure provides a method, apparatus, device, storage medium, and program product for generating training data to solve the problem of poor generation effect of the training data for the evaluation model.

[0005] In a first aspect, the present disclosure provides a method for generating training data, including: obtaining question-and-answer pair data, where the question-and-answer pair data includes labeled question-and-answer pair seed data; enhancing the question data in the question-and-answer pair seed data based on the attribute information of the question-and-answer pair seed data to obtain enhanced question data, and synthesizing first target question-and-answer pair data using the enhanced question data and the answer data in the question-and-answer pair seed data; generating attribution answer data that matches the error type based on the error type of the question-and-answer pair data and / or the first target question-and-answer pair data, and synthesizing second target question-and-answer pair data using the attribution answer data and the question data and / or the enhanced question data in the question-and-answer pair data; and determining the first target question-and-answer pair data and the second target question-and-answer pair data as target training data.

[0006] Second aspect, the present disclosure provides a training data generation device, including: an acquisition module configured to acquire question-and-answer pair data, the question-and-answer pair data including labeled question-and-answer pair seed data; a data enhancement module configured to enhance the question data in the question-and-answer pair seed data based on the attribute information of the question-and-answer pair seed data to obtain enhanced question data, and synthesize first target question-and-answer pair data by using the enhanced question data and the answer data in the question-and-answer pair seed data; an attribution data generation module configured to generate attribution answer data matching the error type based on the error type of the question-and-answer pair data and / or the first target question-and-answer pair data, and synthesize second target question-and-answer pair data by using the attribution answer data and the question data and / or the enhanced question data in the question-and-answer pair data; a data synthesis module configured to determine the first target question-and-answer pair data and the second target question-and-answer pair data as target training data.

[0007] Third aspect, the present disclosure provides a computer device, including: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the training data generation method of the first aspect or any corresponding implementation manner thereof.

[0008] Fourth aspect, the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the training data generation method of the first aspect or any corresponding implementation manner thereof.

[0009] Fifth aspect, the present disclosure provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the training data generation method of the first aspect or any corresponding implementation manner thereof.

[0010] The method, apparatus, device, storage medium and program product for generating training data provided by the present disclosure enhance the question data in the question-and-answer pair seed data according to the attribute information of the labeled question-and-answer pair seed data, so as to ensure that the generated question-enhanced data has high quality, improve the diversity and data volume of the question-enhanced data, and synthesize the question-enhanced data and the answer data in the question-and-answer pair seed data into the first target question-and-answer pair data, thereby improving the diversity of the first target question-and-answer pair data and the quality of question-and-answer pair construction. By analyzing the error types of the answer data in the question-and-answer pair data, generating attribution answer data that matches them, and synthesizing the attribution answer data with the question data and / or question-enhanced data in the question-and-answer pair data, corresponding second target question-and-answer pair data is generated to balance the distribution of negative sample data under different error types. Further, using the first target question-and-answer pair data and the second target question-and-answer pair data as target training data, the automatic generation of the target training data is realized, the generation efficiency of the target training data is improved, and it helps to improve the detection ability of the evaluation model under different error types. At the same time, the generation of attribution answer data can be combined with the real error types in the business scenario to ensure that the target training data can cover the online real business scenario, thereby realizing the efficient generation of the target training data, and ensuring that the evaluation model trained according to the target training data can accurately identify and evaluate the error types and error reasons of the query statement. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the related art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 is a schematic flowchart of a method for generating training data according to an embodiment of the present disclosure;

[0013] Figure 2 is a schematic flowchart of another method for generating training data according to an embodiment of the present disclosure;

[0014] Figure 3 is a schematic flowchart of data enhancement according to an embodiment of the present disclosure;

[0015] Figure 4 is a schematic flowchart of yet another method for generating training data according to an embodiment of the present disclosure;

[0016] Figure 5 is a schematic diagram of generating question-and-answer pair training data according to an embodiment of the present disclosure;

[0017] Figure 6 It is a schematic diagram for generating another question-and-answer pair training data according to an embodiment of the present disclosure;

[0018] Figure 7 It is a structural block diagram of a generating device for training data according to an embodiment of the present disclosure;

[0019] Figure 8 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0021] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0022] For example, when receiving a user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0023] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0024] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present disclosure.

[0025] It can be understood that the data involved in the technical solution of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations, and related provisions.

[0026] In the digital age, data has become a key factor in many scenarios such as enterprise decision-making. Therefore, it is crucial to efficiently extract valuable information from massive data. With the continuous progress of related technologies such as natural language processing and machine learning, especially the emergence and development of pre-trained language models based on the Transformer architecture, such as BERT and GPT, the model's ability to understand and generate natural language has been significantly improved. This enables Text2SQL to handle more complex natural language queries and generate more accurate SQL statements, facilitating non-technical personnel to interact with the database through natural language and meeting the growing demand for rapid data query and analysis.

[0027] To accurately measure the performance of the Text2SQL system in converting natural language queries into correct and effective SQL statements, thereby improving the system's accuracy and reliability and enhancing the user experience, an evaluation model for the SQL query statements generated by the Text2SQL model has emerged. However, the evaluation of the SQL query statements generated by the Text2SQL model not only involves determining whether the generated SQL query statements are correct but, more crucially, locating the error types, error causes, and even error positions of incorrect SQL query statements. It evaluates the query statements generated by the Text2SQL model in an interpretable manner and provides reliable auxiliary decision-making information to provide an optimization direction for the iteration of the Text2SQL model.

[0028] Based on the above business scenario, the Critic Model based on fine-tuning of large language models is proposed as an evaluation model for Text2SQL. To support the training of the Critic Model, the construction of high-quality Critic training data is very important. Most critically, the construction of negative sample data for Critic training involves the description of incorrect SQL, including whether the SQL is correct, the error type (errortype), error cause (errorinfo), and error location (errorinserted) corresponding to the incorrect SQL. Among them, the error type is a special error taxonomy formulated based on business understanding, and the error cause also needs to follow a certain format for easy user understanding. For the training data required by the evaluation model, the following several methods are mainly used currently:

[0029] (1) Manually construct data, that is, manually write natural language questions and corresponding SQL query statements. Its advantage is high data quality, and it can precisely control the difficulty, semantics of the questions, and the structure of SQL query statements to meet specific evaluation requirements. For example, when studying the processing ability of the Text2SQL model for complex multi-table join queries under a specific database schema, natural language questions containing multiple table join operations and corresponding SQL query statements can be manually constructed. However, it can only be applied to small-scale experimental scenarios with high-precision requirements. For example, in the initial test stage of a new Text2SQL model architecture, variables and data characteristics need to be strictly controlled to observe the performance changes of the model. Moreover, its efficiency is low, time-consuming and laborious, and it cannot meet the needs of large-scale data. In addition, manual construction is easily affected by subjective factors, the diversity and coverage of the data are limited, and some special cases may be omitted or it is difficult to consider all possible combinations.

[0030] (2) Template-based data construction, that is, generate data by defining natural language question templates and SQL templates, and quickly generate a large number of paired data by replacing different parameters. This method can efficiently generate a large number of data with certain structural similarities. However, the data generated in this way is relatively similar in structure, may lack sufficient flexibility and diversity, and is difficult to cover all complex situations and changes in real scenarios. If the template design is unreasonable, it may lead to biases in the generated data or non-conformance to actual requirements.

[0031] (3) Extract data from existing database logs. The data extracted in this way may be relatively messy and need to be cleaned and sorted, such as removing incomplete records, incorrect queries, etc. In addition, there may be privacy and security issues with this data, and strict desensitization processing is required, otherwise sensitive information may be leaked.

[0032] (4) Automatically generate data through programs. For example, use random generation algorithms to create table structures, generate random data to fill the tables, and then automatically generate query questions and corresponding SQL based on these table structures and data. This method can quickly generate large-scale data, but the quality of the generated data may need to be further verified and adjusted, that is, the quality of the generated data is difficult to guarantee, and there may be situations that do not conform to grammar rules, logical errors, or do not match the actual application scenarios. Moreover, due to the lack of in-depth understanding of business logic and semantics, the generated data may appear rather rigid and unnatural, and cannot well meet the requirements of the evaluation model for data authenticity and practicality.

[0033] (5) Crowdsourcing data construction: Through a crowdsourcing platform, multiple participants are allowed to write natural language questions and SQL query statements. This approach can collect data from people with different backgrounds and expertise, and the data has diverse styles and ideas. However, the data quality is uneven, and a large amount of manpower and time need to be invested in auditing and screening to ensure the correctness and rationality of the data. The professional levels and backgrounds of the participants vary greatly, which may make it difficult to guarantee the consistency and standardization of the data, increasing the difficulty of data processing and integration.

[0034] For the construction of the training data for the evaluation model, the data is constructed in the form of Query-Response Pair. Specifically, the positive data sample is a Query and the correct SQL statement, and the Response is mainly reflected in the information that the SQL is correct, that is, {"result":"true","errorinfo":"","errortype": "","errorinserted":""}. The negative data sample is a Query, an incorrect SQL statement, and the explanatory information of the incorrect SQL (including error type, error reason, and error location). For the positive sample, the data construction method is similar to the above. For the negative sample, especially the acquisition of error information is difficult. The most commonly used method is manual construction, that is, different proportions of error types are manually constructed, and the explanatory information of the incorrect SQL in the negative sample is manually labeled. The obvious defect of this method is low efficiency, time-consuming and laborious. In addition, the high-quality data manually labeled far from meets the coverage of the real business scenario online. For such special data requirements, facing the problem that it is impossible to obtain from real business data, although some data can be manually labeled, due to the limitations of labor cost and labeling efficiency, it cannot meet the requirements of the scale of the training data for the evaluation model.

[0035] Based on this, the technical solution of the present disclosure enhances the question data in the question-and-answer pair seed data according to the attribute information of the labeled question-and-answer pair seed data, so as to ensure that the generated question-enhanced data has high quality and improve the diversity and quantity of the question-enhanced data. By analyzing the error types of the reply data in the question-and-answer pair data, the corresponding attributed reply data is generated, and the question data is synthesized with the attributed reply data to generate the corresponding target training data, thereby realizing the automatic generation of the target training data and improving the generation efficiency of the target training data.

[0036] According to an embodiment of the present disclosure, an embodiment of a method for generating training data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from here.

[0037] In this embodiment, a method for generating training data is provided, which can be used in computer devices such as computers, tablets, etc. Figure 1 It is a flowchart of the method for generating training data according to an embodiment of the present disclosure, as Figure 1 shown, and the process includes the following steps:

[0038] Step S101, obtain question-and-answer pair data, where the question-and-answer pair data includes labeled question-and-answer pair seed data.

[0039] The question-and-answer pair data is sample training data composed of question data and answer data; the labeled question-and-answer pair seed data is the initial basic sample data when generating training data, and the question-and-answer pair seed data is composed of question data and its corresponding answer data, where the answer data represents whether the query statement generated according to the question data is correct.

[0040] Specifically, the question-and-answer pair (Query-Response Pair) data includes positive sample question-and-answer pair data and negative sample question-and-answer pair data. Among them, the positive sample question-and-answer pair data is composed of question data Query and correct answer data, and the correct answer data Response is composed of a correct SQL statement and description information of the correct SQL statement, mainly reflected in the information that the SQL statement is correct, that is, {"result":"true","errorinfo":"","errortype": "","errorinserted":""}.

[0041] The negative sample question-and-answer pair data is composed of question data Query and incorrect answer data, and the incorrect answer data is composed of an incorrect SQL statement and description information of the incorrect SQL statement (including error type, error reason, error location).

[0042] The question-and-answer pair seed data is the corresponding accurate answer or incorrect answer pre-labeled for each question. The positive sample question-and-answer pair seed data is composed of a question and its corresponding accurate answer, and the negative sample question-and-answer pair seed data is composed of a question and its corresponding incorrect answer.

[0043] In another specific example, an available question-and-answer pair data set disclosed in the business-related field is used to extract labeled question-and-answer pair seed data from the question-and-answer pair data set.

[0044] In another specific example, a data acquisition tool can also be used to automatically capture and label question-and-answer pair seed data from the compliant data of the business scenario.

[0045] Of course, the question-and-answer pair seed data can also be obtained through other means, and specific limitations are not made here.

[0046] Step S102: Based on the attribute information of the Q&A pair seed data, enhance the question data in the Q&A pair seed data to obtain enhanced question data, and synthesize the first target Q&A pair data by using the enhanced question data and the answer data in the Q&A pair seed data.

[0047] The attribute information is the unique attribute possessed by the Q&A pair seed data, such as the semantic information of the question data, the metrics related to the question data, etc. The enhanced question data is multiple question data obtained by performing data enhancement on the Q&A data in the Q&A pair seed data.

[0048] Specifically, extract the original Q&A data from each labeled Q&A pair seed data, rewrite each original Q&A data by using the attribute information of the Q&A pair seed data to generate multiple enhanced question data related to the original Q&A data and multiple enhanced question data related to the business. Here, only the question data is enhanced, and the answer data corresponding to the question data remains unchanged. Then, form the first target Q&A pair data by combining the enhanced question data with its corresponding answer data.

[0049] Step S103: Based on the error types of the Q&A pair data and / or the first target Q&A pair data, generate attribution answer data that matches the error types, and synthesize the second target Q&A pair data by using the attribution answer data and the question data and / or the enhanced question data in the Q&A pair data.

[0050] The error type is the error form of the negative sample Q&A pair data (this negative sample Q&A pair data can come from the initial Q&A pair data or the first target Q&A pair data), such as syntax errors, filtering condition errors, etc. The attribution answer data is the error reason description information generated for the error type, specifically including the description information of the error location, the description information of the error reason, and the description information of the error type.

[0051] In order to adapt to the answer data form of the Q&A pair data and balance the distribution of the negative sample Q&A pair data under different error types, at this time, based on the pre-trained dialogue model, generate the attribution answer data that matches the error type of the negative sample Q&A pair data in the Q&A pair data, and use this attribution answer data to represent the error information of the negative sample Q&A data. Then, synthesize the attribution answer data with the corresponding question data to form a negative sample Q&A pair data with a specific error type, that is, the second target Q&A pair data.

[0052] Similarly, in order to adapt to the reply data form of the first target Q&A pair data, attribution reply data matching the error type of the negative sample Q&A pair data in the first target Q&A pair data can be generated. Subsequently, the attribution reply data is combined with the corresponding question enhancement data to form Q&A pairs, obtaining negative sample Q&A pair data with specific error types, that is, the second target Q&A pair data. Among them, the dialogue model can be trained based on the model structure of a large language model or the model architecture of a neural network, and no specific limitation is made here.

[0053] Step S104: Determine the first target Q&A pair data and the second target Q&A pair data as the target training data.

[0054] By combining the first target Q&A pair data and the second target Q&A pair data, positive sample question data and their corresponding correct information can be obtained to form positive sample Q&A pair data, and negative sample question data and their corresponding error information can be obtained to form negative sample Q&A pair data. Therefore, the first target Q&A pair data and the second target Q&A pair data are used as the target training data, and the evaluation model is trained using this target training data so that the evaluation model can effectively evaluate the query statements generated by the query statement generation model.

[0055] The method for generating training data provided in this embodiment enhances the question data in the Q&A pair seed data according to the attribute information of the labeled Q&A pair seed data to ensure that the generated question enhancement data has high quality, improves the diversity and data volume of the question enhancement data, and combines the question enhancement data with the reply data in the Q&A pair seed data to form the first target Q&A pair data, improving the diversity and Q&A pair construction quality of the first target Q&A pair data. By analyzing the error types of the reply data in the Q&A pair data, attribution reply data matching them is generated, and the attribution reply data is combined with the question data and / or question enhancement data in the Q&A pair data to generate the corresponding second target Q&A pair data to balance the distribution of negative sample data under different error types. Further, using the first target Q&A pair data and the second target Q&A pair data as the target training data, the automatic generation of the target training data is realized, improving the generation efficiency of the target training data, which helps to improve the detection ability of the evaluation model under different error types. At the same time, it is possible to generate attribution reply data in combination with the real error types in the business scenario to ensure that the target training data can cover the online real business scenario, thereby realizing the efficient generation of the target training data, and ensuring that the evaluation model trained according to the target training data can accurately identify and evaluate the error types and error reasons of the query statements.

[0056] In this embodiment, a method for generating training data is provided, which can be used in computer devices such as computers and tablets.Figure 2 is a flowchart of a method for generating training data according to an embodiment of the present disclosure. As Figure 2 shown, the process includes the following steps:

[0057] Step S201, obtain question-and-answer pair data, where the question-and-answer pair data includes labeled question-and-answer pair seed data. For details, please refer to the relevant descriptions of the corresponding steps in the above-mentioned embodiments, which will not be elaborated here.

[0058] Step S202, based on the attribute information of the question-and-answer pair seed data, enhance the question data in the question-and-answer pair seed data to obtain question-enhanced data, and synthesize first target question-and-answer pair data using the question-enhanced data and the answer data in the question-and-answer pair seed data.

[0059] Specifically, the attribute information includes semantic similarity. The above step S202 includes:

[0060] Step S2021, based on the semantic similarity of the question data, rewrite the question data according to a preset positive and negative enhancement ratio to obtain question-enhanced data.

[0061] The preset positive and negative enhancement ratio is the ratio of positive samples and negative samples set in advance, and this preset positive and negative enhancement ratio can be determined according to actual needs, and no specific limitation is made here. As Figure 3 shown, use a question rewriting model to rewrite the question data according to the semantic similarity of the question data, and the answer data corresponding to the question data and the attribute information corresponding to the question data are not rewritten, thereby obtaining question-enhanced data that is semantically consistent with the question data. Among them, the question rewriting model can be trained based on the model structure of a large language model.

[0062] The above rewriting method only needs to modify the question data according to a preset ratio, thereby avoiding uneven data distribution and ensuring that the question-enhanced data has high quality.

[0063] Specifically, the above step S202 further includes:

[0064] Step S2022, obtain the question type to be enhanced.

[0065] The question type is a special question for the business scenario. For example, questions related to "index calculation error", questions related to "store query", questions related to "live broadcast data", etc. Specifically, collect the business question data generated for each business in the business scenario, analyze the business question data, and identify the question types that need to be enhanced according to the data analysis results, so as to generate special questions for the question types.

[0066] Step S2023, extract the attribute indicators matching the question type from the attribute information.

[0067] Attribute indicators are indicators specific to the problem type, such as aggregated indicators that satisfy calculations between indicators. Specifically, by analyzing the problem content of the problem type, attributes directly related to the problem type can be identified from the text of the attribute information according to predefined rules to extract specific attribute indicators that match the problem type; entity recognition technology can also be used to identify attribute entities that match the problem type from the text of the attribute information and determine them as attribute indicators; or a pre-trained problem enhancement model can be used to identify attribute indicators that match the problem type from the attribute information. The problem enhancement model can be trained based on the model structure of the large language model, and no specific limitation is made here.

[0068] Step S2024: Generate problem enhancement data corresponding to the problem type according to the attribute indicators.

[0069] Use the attribute indicators to guide the generation of problem enhancement data with a more specific problem type. For example, when a batch of problem enhancement data related to "indicator calculation error" needs to be generated, attribute indicators that match this problem type (i.e., some aggregated indicators that satisfy calculations between indicators) can be sampled from the attribute information (schema information), and based on this attribute indicator, guide the generation of problem enhancement data that can reflect the intention of indicator calculation.

[0070] In some alternative embodiments, after the above step S2024, it further includes: rewriting the problem enhancement data to obtain the rewritten problem enhancement data.

[0071] Rewrite the generated problem enhancement data so that the rewritten problem enhancement data can better fit the actual user's question style. For example, the questions raised by users in the actual business scenario are sometimes relatively vague and do not give a complete description of the fields or enumerated values. At this time, the problem enhancement data can be rewritten according to the vague question style.

[0072] Such as Figure 3 As shown, input the above-generated problem enhancement data into a query statement generation model (such as the Text2SQL model), use the query statement generation model to generate a query statement corresponding to the problem enhancement data, and then use the annotation model to automatically annotate the query statement or manually annotate it by technicians to generate response data corresponding to the problem enhancement data. Subsequently, use the problem enhancement data and its corresponding response data to construct Q&A pair data for training.

[0073] Step S2025: Synthesize the first target Q&A pair data using the problem enhancement data and the response data in the Q&A pair seed data. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, and will not be elaborated here.

[0074] Step S203: Generate attribution response data that matches the error type based on the Q&A pair data and / or the error type of the first target Q&A pair data, and synthesize the second target Q&A pair data by using the attribution response data and the question data and / or question enhancement data in the Q&A pair data. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be elaborated here.

[0075] Step S204: Determine the first target Q&A pair data and the second target Q&A pair data as the target training data. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be elaborated here.

[0076] The method for generating training data provided in this embodiment enhances the question data in the Q&A pair seed data through different data enhancement methods, ensuring that the question enhancement data has high quality. At the same time, it is beneficial to generate diverse question enhancement data and increase the data volume of the question enhancement data.

[0077] In this embodiment, a method for generating training data is provided, which can be used in computer devices such as computers and tablets. Figure 4 It is a flowchart of the method for generating training data according to an embodiment of the present disclosure, as Figure 4 shown. The process includes the following steps:

[0078] Step S301: Obtain Q&A pair data, where the Q&A pair data includes labeled Q&A pair seed data. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be elaborated here.

[0079] Step S302: Enhance the question data in the Q&A pair seed data based on the attribute information of the Q&A pair seed data to obtain question enhancement data, and synthesize the first target Q&A pair data by using the question enhancement data and the answer data in the Q&A pair seed data. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be elaborated here.

[0080] Step S303: Generate attribution response data that matches the error type based on the Q&A pair data and / or the error type of the first target Q&A pair data, and synthesize the second target Q&A pair data by using the attribution response data and the question data and / or question enhancement data in the Q&A pair data.

[0081] Specifically, the above Step S303 includes:

[0082] Step S3031: Obtain the predicted query statement corresponding to the Q&A pair data and / or the first target Q&A pair data.

[0083] The predicted query statement is a query statement generated based on the question. Specifically, it can be generated based on the question data in the question-answer pair data, or can also be generated based on the question enhancement data in the first target question-answer pair data.

[0084] In a specific example, the question data in the question-answer pair data or the question enhancement data in the first target question-answer pair data is input into a query statement generation model (such as a Text2SQL model). The query statement generation model analyzes the question description information of the question data in the question-answer pair data or the question enhancement data in the first target question-answer pair data, and generates a corresponding predicted query statement according to the question description information.

[0085] Step S3032, if the predicted query statement is incorrect, generate error attribution information for the predicted query statement.

[0086] The error attribution information is the text description information of the error reason of the predicted query statement, specifically including syntax error description, query error description, condition error description, etc. Input the question enhancement data, attribute information, and the predicted query statement into a pre-trained detection model. Use the detection model to detect whether the predicted query statement is correct. If the predicted query statement is correct, use the question enhancement data and its corresponding answer data as positive sample question-answer pairs. If the predicted query statement is incorrect, use the detection model to determine the error reason of the predicted query statement, and generate the text description information of this error reason, that is, the error attribution information of the predicted query statement. Similarly, input the question data, attribute information, and the predicted query statement in the question-answer pair data into a pre-trained detection model, and use the detection model to detect whether the predicted query statement is correct, so as to generate corresponding error attribution information when the predicted query statement is incorrect.

[0087] Step S3033, parse the error attribution information to determine the error location and error type corresponding to the predicted query statement.

[0088] The error location is the location where the error occurs in the predicted query statement; the error type is the type to which the error of the predicted query statement belongs, such as syntax error, query condition error, etc. Since the error attribution information includes the reason for the error occurrence, by analyzing the error attribution information, the reason for the error occurrence can be determined. According to the reason for the error occurrence, the location where the error occurs can be located, and the type to which the error belongs can be determined.

[0089] Step S3034, synthesize the error type, error location, and error attribution information into attribution reply data.

[0090] Combine the error type corresponding to the error attribution information, and merge the error type, error location, and error attribution information to obtain complete attribution reply data. Subsequently, based on the correspondence between the question enhancement data and the error attribution information, and the correspondence between the question data in the Q&A pair data and the error attribution information, negative sample Q&A pair data for evaluating a specific error type can be constructed, facilitating the use of this negative sample Q&A pair data to train an evaluation model to evaluate the correctness of the query statement generated by the text-to-query statement conversion, and when the query statement is incorrect, its corresponding error type can be determined.

[0091] In some alternative embodiments, before step S3033, the above method further includes:

[0092] Step a1, obtain the true query statement corresponding to the Q&A pair data and / or the first target Q&A pair data.

[0093] Step a2, correct the predicted query statement to generate a corrected query statement.

[0094] Step a3, execute the true query statement and the corrected query statement, and obtain the first execution result of the true query statement and the second execution result of the corrected query statement.

[0095] Step a4, if the first execution result and the second execution result are consistent, then parse the error attribution information to determine the error location and error type corresponding to the predicted query statement.

[0096] The true query statement is an accurate query statement marked from the online real data for the question data in the Q&A pair data, or an accurate query statement for the question enhancement data in the first target Q&A pair data. The detection model has pre-written correction prompt information, which is used to guide the detection model to correct errors.

[0097] Specifically, as Figure 5 shown, when it is determined that the predicted query statement is incorrect, the detection model uses the correction prompt information to correct the predicted query statement according to the error attribution information to obtain the corrected corrected query statement. Subsequently, the corrected query statement obtained after correction and the true query statement are respectively executed in the dataset to obtain the corresponding execution results.

[0098] Compare the first execution result of the true query statement with the second execution result of the corrected query statement. If the two are consistent, generate a scoring result with Score equal to 1; if the two are consistent, generate a scoring result with Score not equal to 1, as Figure 5 shown.

[0099] Since the detection model can obtain the correct corrected query statement due to reasonable error reasons. Therefore, when it is determined that the first execution result is consistent with the second execution result, it can be considered that the obtained error attribution information is reasonable. By screening the incorrect predicted query statements, high-quality negative sample training data can be constructed. The error reason of the predicted query statement is the error attribution information in the attribution reply data. At this time, the incorrect predicted query statement can be labeled as 'false'. Then, through the detection model, the error type and error location are generated, and thus complete negative sample Q&A pair data can be obtained.

[0100] In the above embodiment, by combining the execution results of the real query statement and the corrected query statement, the predicted query statement corresponding to the corrected query statement is screened, so as to determine the error type and error location corresponding to the predicted query statement, making the determination of the error type and error location more conform to the distribution of the online real data.

[0101] In some alternative embodiments, the above method further includes: if the first execution result and the second execution result are inconsistent, re-determine whether the predicted query statement is correct, so as to regenerate the error attribution information of the predicted query statement when the predicted query statement is incorrect.

[0102] If the first execution result and the second execution result are inconsistent, that is, Score is not 1, it means that an incorrect result is obtained in either the generation of the error reason or the correction of the predicted query statement. At this time, it is possible to re-determine whether the predicted query statement is correct, and when the predicted query statement is incorrect, regenerate the error attribution information of the predicted query statement. Try multiple times in sequence to determine whether a correct corrected query statement can be obtained. If the correct corrected query statement still cannot be obtained after the retry count reaches the preset number (such as 5 times), the problem enhancement data corresponding to the corrected query statement will be deleted and will no longer be used as the generation of training data.

[0103] In the above embodiment, by trying the incorrect predicted query statement multiple times in sequence to determine whether it can be used as negative sample training data, the probability of high-quality training data is improved.

[0104] Specifically, the above step S303 may further include:

[0105] Step S3035, obtaining the reference query statement corresponding to the Q&A pair data and / or the first target Q&A pair data.

[0106] Considering that corresponding real query statements cannot be collected in all business scenarios, in order to make up for the problems of insufficient data and unbalanced data distribution, a query statement generation model (such as the Text2SQL model) can be used to generate predicted query statements based on the question data in the question-answer pair data as reference query statements, or predicted query statements generated based on the question enhancement data in the first target question-answer pair data can be used as reference query statements.

[0107] Step S3036, rewrite the reference query statement according to a preset target error type to generate a rewritten target query statement.

[0108] The target error type is one or several pre-specified error types. As Figure 6 shown, take the reference query statement, question enhancement data, and attribute information as the input of the detection model, and pre-write corresponding prompt information for the target error type, so that the detection model can rewrite the reference query statement according to the prompt information to rewrite it into a target query statement for the question enhancement data that matches the target error type.

[0109] Similarly, the reference query statement, the question data in the question-answer pair data, and the attribute information corresponding to the question data can be used as the input of the detection model, and corresponding prompt information can be pre-written for the target error type, so that the detection model can rewrite the reference query statement according to the prompt information to rewrite it into a target query statement for the question data in the question-answer pair data that matches the target error type.

[0110] Step S3037, perform error attribution on the target query statement to generate the error location and error attribution information corresponding to the target query statement.

[0111] Use the detection model to perform error attribution on the target query statement to generate error attribution information for the target error type. By parsing the error attribution information, determine the location where the target query statement has an error under the current target error type, as Figure 6 shown.

[0112] Step S3038, synthesize the target error type, error location, and error attribution information into attribution reply data.

[0113] Merge the target error type, the error location corresponding to the target error type, and the error attribution information to obtain complete attribution reply data. Thus, combined with the correspondence between the question data or question enhancement data and the error attribution information, construct negative sample question-answer pair data for evaluating the target error type, which is convenient for training the evaluation model using the negative sample question-answer pair data, so that when the query statement generation model outputs a query statement with the target error type, the evaluation model can accurately identify and judge the query statement of the target error type.

[0114] Step S304: Determine the first target Q&A pair data and the second target Q&A pair data as target training data. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be elaborated here.

[0115] The method for generating training data provided in this embodiment obtains predicted query statements to perform error attribution when they are incorrect, and then determines the corresponding error types and error positions according to the error attribution information. It can adapt to the specific attribution response data form of question data and question augmentation data, balance the negative sample data distribution under different error types, and improve the detection ability of the subsequent evaluation model for query statements of different error types. By obtaining reference query statements, rewriting them into target query statements with specific target error types, and performing error attribution on the target query statements to determine the corresponding error attribution information and error positions, negative sample data with specific error types can be constructed, which can make up for the data deficiency problem and data distribution imbalance problem in business scenarios without real query statements, facilitate adapting to different business scenarios, adjust negative sample data according to the iterative requirements of the evaluation model, balance the negative sample data distribution under different error types, and further improve the detection ability of the evaluation model for query statements of different error types.

[0116] In this embodiment, a device for generating training data is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be elaborated again. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0117] This embodiment provides a device for generating training data, as Figure 7 shown, including:

[0118] An acquisition module 401, configured to acquire Q&A pair data, where the Q&A pair data includes labeled Q&A pair seed data.

[0119] A data augmentation module 402, configured to augment the question data in the Q&A pair seed data based on the attribute information of the Q&A pair seed data to obtain question augmentation data, and synthesize first target Q&A pair data by using the question augmentation data and the answer data in the Q&A pair seed data.

[0120] An attribution data generation module 403, configured to generate attribution response data matching the error type based on the error type of the Q&A pair data and / or the first target Q&A pair data, and synthesize second target Q&A pair data by using the attribution response data and the question data and / or question augmentation data in the Q&A pair data.

[0121] The data synthesis module 404 is configured to determine the first target Q&A pair data and the second target Q&A pair data as target training data.

[0122] In some alternative embodiments, the attribute information includes semantic similarity. Correspondingly, the data augmentation module 402 includes:

[0123] The first rewriting unit is configured to rewrite the question data based on the semantic similarity of the question data according to a preset positive and negative augmentation ratio to obtain enhanced question data.

[0124] In some alternative embodiments, the data augmentation module 402 includes:

[0125] The type acquisition unit is configured to acquire the type of the question to be augmented.

[0126] The metric extraction unit is configured to extract an attribute metric matching the question type from the attribute information.

[0127] The data generation unit is configured to generate enhanced question data corresponding to the question type according to the attribute metric.

[0128] In some alternative embodiments, the data augmentation module 402 may further include:

[0129] The second rewriting unit is configured to rewrite the enhanced question data to obtain rewritten enhanced question data.

[0130] In some alternative embodiments, the attribution data generation module 403 includes:

[0131] The predicted statement acquisition unit is configured to acquire a predicted query statement corresponding to the Q&A pair data and / or the first target Q&A pair data.

[0132] The first error attribution unit is configured to generate error attribution information for the predicted query statement if the predicted query statement is incorrect.

[0133] The error parsing unit is configured to parse the error attribution information to determine the error location and error type corresponding to the predicted query statement.

[0134] The first attribution synthesis unit is configured to synthesize the error type, error location, and error attribution information into attribution reply data.

[0135] In some alternative embodiments, the above device further includes:

[0136] The true statement acquisition module is configured to acquire a true query statement corresponding to the Q&A pair data and / or the first target Q&A pair data.

[0137] A correction module, configured to correct a predicted query statement and generate a corrected query statement.

[0138] A statement execution module, configured to execute a real query statement and a corrected query statement, and obtain a first execution result of the real query statement and a second execution result of the corrected query statement.

[0139] An error parsing module, configured to, if the first execution result is consistent with the second execution result, parse error attribution information and determine an error position and an error type corresponding to the predicted query statement.

[0140] In some alternative embodiments, the above device further includes:

[0141] A retry module, configured to, if the first execution result is inconsistent with the second execution result, re-determine whether the predicted query statement is correct, so as to regenerate error attribution information of the predicted query statement when the predicted query statement is incorrect.

[0142] In some alternative embodiments, the attribution data generation module 403 may further include:

[0143] A reference statement acquisition unit, configured to acquire a reference query statement corresponding to question-answer pair data and / or first target question-answer pair data.

[0144] A third rewriting unit, configured to rewrite the reference query statement according to a preset target error type to generate a rewritten target query statement.

[0145] A second error attribution unit, configured to perform error attribution on the target query statement to generate an error position and error attribution information corresponding to the target query statement.

[0146] A second attribution synthesis unit, configured to synthesize the target error type, the error position, and the error attribution information into attribution reply data.

[0147] The further function descriptions of the above modules and units are the same as those in the corresponding foregoing embodiments, and will not be elaborated herein.

[0148] The training data generation device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0149] The training data generation device provided in this embodiment enhances the question data in the question-and-answer pair seed data according to the attribute information of the labeled question-and-answer pair seed data, so as to ensure that the generated question-enhanced data has high quality, improve the diversity and data volume of the question-enhanced data, and synthesize the question-enhanced data and the answer data in the question-and-answer pair seed data into the first target question-and-answer pair data, improving the diversity of the first target question-and-answer pair data and the quality of question-and-answer pair construction. By analyzing the error types of the answer data in the question-and-answer pair data, attribution answer data matching them is generated, and the attribution answer data is synthesized with the question data and / or question-enhanced data in the question-and-answer pair data to generate corresponding second target question-and-answer pair data, so as to balance the negative sample data distribution under different error types. Further, using the first target question-and-answer pair data and the second target question-and-answer pair data as target training data, the automatic generation of the target training data is realized, and the generation efficiency of the target training data is improved. At the same time, it is possible to generate attribution answer data in combination with the real error types in the business scenario to ensure that the target training data can cover the online real business scenario, thereby realizing the efficient generation of the target training data, and ensuring that the evaluation model trained according to the target training data can accurately identify and evaluate the error types and error reasons of the query statement.

[0150] This embodiment of the disclosure also provides a computer device having the above Figure 7 shown training data generation device.

[0151] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present disclosure. As Figure 8 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 8 In

[0152] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device may be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.

[0153] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0154] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. The above-mentioned networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0155] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.

[0156] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means, Figure 8 Taking connection through a bus as an example.

[0157] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (such as an LED), and a tactile feedback device (such as a vibration motor), etc. The above-mentioned display device includes, but is not limited to, a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.

[0158] The computer device further includes a communication interface for the computer device to communicate with other devices or communication networks.

[0159] Embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be processed by such software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.

[0160] A part of the present disclosure can be applied as a computer program product, such as computer program instructions. When executed by a computer, through the operation of the computer, the methods and / or technical solutions according to the present disclosure can be invoked or provided. Those skilled in the art should be able to understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0161] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for generating training data, characterized in that: The method comprises: Acquire question-answer pair data, wherein the question-answer pair data includes labeled question-answer pair seed data; Based on the attribute information of the question-answer pair seed data, the question data in the question-answer pair seed data is enhanced to obtain question-enhanced data, and the question-enhanced data and the answer data in the question-answer pair seed data are used to synthesize the first target question-answer pair data; Based on the error type of the question-answer pair data and / or the first target question-answer pair data, generate attribution reply data matching the error type, and synthesize second target question-answer pair data using the attribution reply data and the question data and / or the question enhancement data in the question-answer pair data; The first target question-answer pair data and the second target question-answer pair data are determined as target training data.

2. The method according to claim 1, characterized in that The step of enhancing the question data in the question-answer pair seed data based on the attribute information of the question-answer pair seed data to obtain enhanced question data includes: Based on the semantic similarity of the question data, rewriting the question data according to a preset positive and negative enhancement ratio to obtain the question enhanced data; The attribute information includes semantic similarity.

3. The method according to claim 1 or 2, characterized in that: The step of enhancing the question data in the question-answer pair seed data based on the attribute information of the question-answer pair seed data to obtain enhanced question data includes: Get the type of question to be enhanced; Extracting attribute indicators matching the problem type from the attribute information; The question enhancement data corresponding to the question type is generated according to the attribute indicator.

4. The method according to claim 3, characterized in that After generating the question enhancement data corresponding to the question type according to the attribute index, the method further includes: The question enhancement data is rewritten to obtain the rewritten question enhancement data.

5. The method according to claim 1, characterized in that The generating attribution reply data matching the error type based on the error type of the question-answer pair data and / or the first target question-answer pair data comprises: Obtaining a predicted query statement corresponding to the question-answer pair data and / or the first target question-answer pair data; If the predicted query statement is wrong, generating error attribution information of the predicted query statement; Parsing the error attribution information to determine the error location and error type corresponding to the predicted query statement; The error type, the error location and the error attribution information are combined into the attribution reply data.

6. The method according to claim 5, characterized in that Before parsing the error attribution information and determining the error location and error type corresponding to the predicted query statement, the method further includes: Obtaining a real query statement corresponding to the question-answer pair data and / or the first target question-answer pair data; Correcting the predicted query statement to generate a corrected query statement; Executing the real query statement and the corrected query statement to obtain a first execution result of the real query statement and a second execution result of the corrected query statement; If the first execution result is consistent with the second execution result, the error attribution information is parsed to determine the error location and error type corresponding to the predicted query statement.

7. The method according to claim 6, characterized in that Also includes: If the first execution result and the second execution result are inconsistent, whether the predicted query statement is correct is re-determined, so as to regenerate error attribution information of the predicted query statement when the predicted query statement is wrong.

8. The method according to claim 1 or 5, characterized in that: The generating attribution reply data matching the error type based on the error type of the question-answer pair data and / or the first target question-answer pair data comprises: Obtaining a reference query statement corresponding to the question-answer pair data and / or the first target question-answer pair data; Rewrite the reference query statement according to a preset target error type to generate a rewritten target query statement; Performing error attribution on the target query statement, and generating error position and error attribution information corresponding to the target query statement; The target error type, the error location, and the error attribution information are combined into the attribution reply data.

9. A training data generating device, characterized in that: The device comprises: An acquisition module, configured to acquire question-answer pair data, wherein the question-answer pair data includes labeled question-answer pair seed data; A data enhancement module, configured to enhance the question data in the question-answer pair seed data based on the attribute information of the question-answer pair seed data to obtain enhanced question data, and synthesize first target question-answer pair data using the enhanced question data and the answer data in the question-answer pair seed data; an attribution data generation module, configured to generate attribution reply data matching the error type based on the question-answer pair data and / or the first target question-answer pair data, and synthesize the second target question-answer pair data using the attribution reply data and the question data and / or the question enhancement data in the question-answer pair data; A data synthesis module is used to determine the first target question-answer pair data and the second target question-answer pair data as target training data.

10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for generating training data according to any one of claims 1 to 8 by executing the computer instructions.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for generating training data according to any one of claims 1 to 8.

12. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the method for generating training data according to any one of claims 1 to 8.

Citation Information

Cited By

  • Method and device for expanding performance evaluation data of large language model

    CN120929840A

  • Test set construction method and device based on multi-modal knowledge tree, equipment and medium

    CN121278391A

  • Methods, apparatus, equipment, and media for constructing test sets based on multimodal knowledge trees.

    CN121278391B