Training corpus generation method and device based on large language model, medium and equipment
By using the second largest language model to generate training corpus, the problem of insufficient training data of large language models is solved, the automatic generation and quality improvement of training data is achieved, and the training needs of large models are met.
Patent Information
- Application Number
- CN202510990670.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-17
AI Technical Summary
In the prior art, the training corpus of large language models is limited, and there is a risk of training data exhaustion, manual labeling yield is low and quality is unstable, resulting in low training efficiency and uneven quality.
By obtaining the first error sample, execution environment and natural language problem samples, the second largest language model is used to generate the first natural language problem and structured query statements, based on these, the training corpus is generated to expand the training data, reduce human participation, and improve generation efficiency and quality stability.
The automatic generation of training corpus is realized, the number of training data is increased, the risk of data exhaustion is reduced, the training efficiency and quality stability is improved, and the training needs of large models are met.
Smart Images

Figure CN120509494A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of large models, intelligent agents, and artificial intelligence, and specifically to a method, apparatus, medium, and equipment for generating training corpus based on a large language model. Background Art
[0002] With the continuous development of artificial intelligence technology, large language models are increasingly being used in daily life. For example, in data query scenarios, to quickly retrieve target data from a database, a large language model is often used to generate a corresponding query statement based on the data query requirements entered by the user in natural language. This query statement is then used to retrieve data from the database. Large language models can be trained using training corpus.
[0003] However, the existing corpus in related technologies is limited, making it difficult to meet the training needs of large models, and there is a risk of gradual depletion of the training corpus. Furthermore, because the training corpus in related technologies is generally obtained through manual annotation, the efficiency of training corpus generation is relatively low. Furthermore, due to differences in the professional level and / or comprehension ability of different annotators, the annotation quality varies, making it difficult to effectively guarantee the overall quality of the training corpus. Summary of the Invention
[0004] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] In a first aspect, the present disclosure provides a method for generating training corpus based on a large language model, the method comprising: Obtaining a first error sample, an execution environment, and a natural language question sample for data query, wherein the first error sample includes a sample natural language question for data query, an erroneous structured query statement generated by a first large language model based on the sample natural language question, and a correct structured query statement corresponding to the erroneous structured query statement, and the execution environment is used to represent a query data table obtained based on the first error sample; Inputting the first error sample, the execution environment, and the natural language question example into a second large language model to obtain a first natural language question for data query and a first structured query statement corresponding to the first natural language question, wherein the second large language model is used to understand the first error sample, obtain a first query keyword that causes the first large language model to generate the erroneous structured query statement, rewrite the natural language question example according to the first query keyword and the execution environment, output the first natural language question, and output the first structured query statement based on at least the first natural language question and the execution environment; Based on the first natural language question and the first structured query statement, a training corpus for training the first language model is obtained.
[0006] In a second aspect, the present disclosure provides a training corpus generation device based on a large language model, the training corpus generation device based on a large language model comprising: an acquisition module, configured to acquire a first error sample, an execution environment, and a sample natural language question for data query, wherein the first error sample includes a sample natural language question for data query, an erroneous structured query statement generated by a first large language model based on the sample natural language question, and a correct structured query statement corresponding to the erroneous structured query statement; and the execution environment is used to represent a query data table obtained based on the first error sample; a first processing module configured to input the first error sample, the execution environment, and the natural language question example into a second large language model to obtain a first natural language question for data query and a first structured query statement corresponding to the first natural language question, wherein the second large language model is configured to understand the first error sample, obtain a first query keyword that causes the first large language model to generate the erroneous structured query statement, rewrite the natural language question example based on the first query keyword and the execution environment, output the first natural language question, and output the first structured query statement based at least on the first natural language question and the execution environment; The second processing module is used to obtain training corpus for training the first language model based on the first natural language question and the first structured query statement.
[0007] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when executed by a processing device.
[0008] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the method in the first aspect.
[0009] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.
[0010] Through the above technical solution, the first error sample, the execution environment, and the natural language question sample for data query can be input into the second large language model to obtain the first natural language question for data query and the first structured query statement corresponding to the first natural language question. Based on the first natural language question and the first structured query statement, the training corpus for training the first large language model is obtained. Since the training corpus is expanded based on the generation of the large language model during the training corpus generation process, the automatic generation of the training corpus can be achieved. On the one hand, the amount of training corpus can be increased, the risk of training corpus exhaustion can be reduced, and the model training needs can be better met when the demand for model training increases. On the other hand, human participation can be reduced, and the efficiency of training corpus generation can be improved. In addition, during the generation of the training corpus, the natural language question sample is rewritten with reference to the error sample. This can make the error information in the training corpus cover the key error features in the error sample while reducing the interference and redundancy of the remaining noise data, thereby improving the quality stability of the training corpus to a certain extent.
[0011] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 is a flowchart of a method for generating training corpus based on a large language model according to an exemplary embodiment of the present disclosure; Figure 2 is a schematic diagram of a process for generating a first natural language question by using a second large language model according to an exemplary embodiment of the present disclosure; Figure 3 This is a schematic diagram of a process for generating a first structured query statement by using a second language model according to an exemplary embodiment of the present disclosure; Figure 4is a flowchart of another method for generating training corpus based on a large language model according to an exemplary embodiment of the present disclosure; Figure 5 is a structural block diagram of a training corpus generation device based on a large language model according to an exemplary embodiment of the present disclosure; Figure 6 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0013] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0014] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0015] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0017] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0019] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0020] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0021] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0022] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0023] At the same time, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0024] In data query scenarios, to quickly retrieve target data from a database, a large language model is often used. Based on the user's natural language input, a query statement is generated, and then the query statement is used to retrieve data from the database. For example, advertisers may need to query advertising data to promptly understand the effectiveness of their advertising campaigns. In this scenario, advertisers can query advertising data using a natural language interactive database query product based on a large language model. The large language model can be trained using training data.
[0025] However, in related technologies, large language model training solutions often rely on data culled from the internet and low-level, open corpora within the domain. However, as the demand for model training continues to rise, existing corpus resources are gradually failing to meet the higher requirements of model training. Specifically, the following problems exist: Training data exhaustion: As large models continue to grow in size, their parameter size typically grows exponentially. Each model parameter requires sufficient data to adjust its weight to ensure accurate learning and generalization. However, the limited existing corpus in related technologies is no longer sufficient to train larger models, leading to an increasing risk of data exhaustion. Low manual labeling output and low labeling efficiency: As the model size of large models continues to expand, the demand for labeled data is also growing exponentially. However, the output of manual labeling is limited and it is difficult to meet the needs of large-scale model training. In addition, during the manual labeling process, the labelers need to analyze, understand and label the data one by one, which is not only time-consuming, but also increases the difficulty of labeling when faced with complex data or highly professional content, resulting in relatively low efficiency in generating training corpus; Training quality is difficult to guarantee: Public corpora on the Internet, in-domain corpora, and manually annotated corpora may contain a large amount of noise data, such as grammatical errors, spelling errors, and irrelevant content. This leads to uneven quality of training corpora, making it difficult to ensure the quality stability of the training process.
[0026] In view of this, the present disclosure provides a method, apparatus, medium and device for generating training corpus based on a large language model to solve the above technical problems.
[0027] The following further explains the embodiments of the present disclosure with reference to the accompanying drawings.
[0028] Figure 1 This is a flowchart of a method for generating training corpus based on a large language model according to an exemplary embodiment of the present disclosure, with reference to Figure 1 , the training corpus generation method based on the large language model can include the following steps: S101: Obtain a first error sample, an execution environment, and a natural language question sample for data query, wherein the first error sample includes a sample natural language question for data query, an erroneous structured query statement generated by a first large language model based on the sample natural language question, and a correct structured query statement corresponding to the erroneous structured query statement, and the execution environment is used to characterize a query data table obtained based on the first error sample.
[0029] In this embodiment, the execution environment can be obtained by extracting the query keywords in the first error sample. For example, if the first error sample includes the sample natural language question "overall rate of return", the incorrectly structured query statement "avg(`revenue` / `income`)|sum(`revenue` / `income`)", and the correctly structured query statement "sum(`revenue`) / sum(`income`)", the query keywords in the first error sample can be extracted to obtain the execution environment "`revenue`,int, \n `income`,int".
[0030] For example, if the first error sample includes the sample natural language question "average of rate of return (revenue / income)", the incorrectly structured query statement "sum(`revenue`) / sum(`income`)|sum(`revenue` / `income`)", and the correctly structured query statement "avg(`revenue` / `income`)", the query keywords in the first error sample can be extracted to obtain the execution environment "`revenue`,int, \n `income`,int,".
[0031] It should be understood that if the first error samples include multiple ones, then when determining the execution environment, the error type of each first error sample can be determined first, and then clustered based on the error type to obtain error sets of different error types. For each error set, the query keywords in each first error sample under the error set can be extracted to obtain the execution environment corresponding to the different error types.
[0032] In this embodiment, the first error sample can be obtained in the following manner: Obtain an error definition for the first language model, a sample natural language question for data query, and a third structured query statement generated by the first language model based on the sample natural language question, wherein the error definition includes the processing steps required for the first language model to generate the structured query statement, the error type under each processing step, and the error details corresponding to each error type; input the sample natural language question, the third structured query statement, and the error definition into the second language model to obtain a first error sample, wherein the second language model is also used to understand the sample natural language question and the third structured query statement based on the error definition to obtain an error structured query statement, generate a correct structured query statement based on the error structured query statement and the sample natural language question corresponding to the error structured query statement, and output the first error sample based on the error structured query statement, the sample question corresponding to the error structured query statement, and the correct structured query statement.
[0033] It should be understood that when generating structured query statements based on a large language model, the required processing steps generally include two steps: question understanding and logical processing. Therefore, the processing steps in this embodiment can include question understanding and logical processing. The question understanding step is used to identify user intent based on the input natural language question, and the logical processing step is used to perform logical reasoning based on the identified user intent. Through a large number of experiments, it was found that the error types corresponding to the question understanding step and the logical processing step, as well as the error details corresponding to each error type, can be shown in Table 1.
[0034] Table 1. Error definitions for the top language model
[0035] It should be understood that this is merely an illustrative description and does not constitute a limitation on the solution. Limit and sample are commonly used clauses in structured query statements to limit the query result set.
[0036] In addition, it should be understood that when the first language model generates the third structured query statement based on the sample natural language question, the generated third structured query statement may be correct or incorrect. Therefore, in order to obtain a more accurate error sample, after obtaining the third structured query statement corresponding to the sample natural language question through the first language model, the sample natural language question, the third structured query statement and the error definition can be input into the second language model, so that the second language model can identify whether the third structured query statement is correct based on the error definition and the sample natural language question. If the third structured query statement is incorrect, it will be treated as an erroneous structured query statement, and it can be corrected in combination with the corresponding sample natural language question to generate a correct structured query statement, so that the first error sample can be obtained based on the erroneous structured query statement, the sample question corresponding to the erroneous structured query statement and the correct structured query statement.
[0037] S102: Input the first error sample, the execution environment, and the natural language question sample into the second largest language model to obtain a first natural language question for data query and a first structured query statement corresponding to the first natural language question, wherein the second largest language model is used to understand the first error sample, obtain a first query keyword that causes the first largest language model to generate an erroneous structured query statement, rewrite the natural language question sample according to the first query keyword and the execution environment, output the first natural language question, and output the first structured query statement at least based on the first natural language question and the execution environment.
[0038] For example, a first prompt word template and a second prompt word template can be pre-set, wherein the first prompt word template is used to generate a first natural language question, and the second prompt word template is used to generate a first structured query statement. Thus, after obtaining a first error sample, an execution environment, and a natural language question example, the first error sample, the execution environment, and the natural language question example can be filled into the first prompt word template to obtain a first prompt word, and the first prompt word is input into the second large language model to obtain a first natural language question with the error in the first error sample, such as Figure 2 Then, the first natural language question, the execution environment, and the first error sample can be filled into the second prompt word template to obtain the second prompt word, and the second prompt word is input into the second large language model to obtain the corresponding first structured query statement with the error in the first error sample, as shown in FIG. Figure 3 The first prompt word template and the second prompt word template can be determined according to actual conditions, and the embodiment of the present disclosure does not impose any limitation on this.
[0039] S103: Based on the first natural language question and the first structured query statement, obtain training corpus for training a first large language model.
[0040] Through the above technical solution, the first error sample, the execution environment, and the natural language question sample for data query can be input into the second large language model to obtain the first natural language question for data query and the first structured query statement corresponding to the first natural language question. Based on the first natural language question and the first structured query statement, the training corpus for training the first large language model is obtained. Since the training corpus is expanded based on the generation of the large language model during the training corpus generation process, the automatic generation of the training corpus can be achieved. On the one hand, the amount of training corpus can be increased, the risk of training corpus exhaustion can be reduced, and the model training needs can be better met when the demand for model training increases. On the other hand, human participation can be reduced, and the efficiency of training corpus generation can be improved. In addition, during the generation of the training corpus, the natural language question sample is rewritten with reference to the error sample. This can make the error information in the training corpus cover the key error features in the error sample while reducing the interference and redundancy of the remaining noise data, thereby improving the quality stability of the training corpus to a certain extent.
[0041] In addition, in the process of generating training corpus, the thinking boundary of the second largest language model can be defined through the execution environment, so that the generated training corpus can meet actual business needs and further improve the quality of the training corpus.
[0042] To facilitate understanding of the training corpus generation method based on a large language model provided by the present disclosure, possible implementation methods of the present disclosure are described below.
[0043] In a possible manner, the first error sample may further include an error label, where the error label represents the cause of the error in the incorrectly structured query statement. Accordingly, the second largest language model is used to output the first natural language question in the following manner: Based on the error label, a target error type of the erroneous structured query statement is determined; according to the target error type, a second execution environment is determined in multiple preset first execution environments, wherein each first execution environment is pre-associated with the error type of the erroneous structured query statement generated by the first large language model, and the error type associated with the second execution environment is the same as the target error type; the natural language question sample is rewritten according to the first query keyword and the second execution environment, and the first natural language question is output.
[0044] For example, multiple error labels can be pre-set, and each error label includes at least an error type. Thus, after obtaining an erroneous structured query statement corresponding to a sample natural language question, the erroneous structured query statement can be analyzed in combination with the multiple pre-set error labels and the sample natural language question to obtain an error label for the erroneous structured query statement. Alternatively, after obtaining an erroneous structured query statement corresponding to a sample natural language question, the sample natural language question, the erroneous structured query statement, and the multiple pre-set error labels can be input into a large model, and the large model outputs an error label for the erroneous structured query statement.
[0045] Among them, the content included in the error label can be determined according to the actual situation, and the embodiment of the present disclosure does not impose any restrictions on this. For example, in order to accurately indicate the error cause of the erroneous structured query statement, the content included in the error label may include the processing link required for the large language model to generate the structured query statement, the error type corresponding to the processing link, and the error details corresponding to the error type. For example, if the first large language model has a problem in the mean calculation process under the question understanding link during the process of generating a structured query statement based on a sample natural language question, the error label can be: question understanding-calculation method-mean calculation.
[0046] It should be understood that different error types generally correspond to different execution environments. Therefore, an execution environment can be pre-built for each error type. For example, multiple preset first execution environments can be obtained in the following manner: Acquire multiple different first error samples, and for each first error sample, determine the error type of the erroneous structured query statement in the first error sample based on the error label included in the first error sample; perform field recognition on the first error samples with the same error type to obtain the initial data field corresponding to the error type, and perform deduplication processing on the initial data field to obtain the deduplication data field corresponding to the error type; construct a data table based on the deduplication data field corresponding to the error type and the field type corresponding to the deduplication data field to obtain the first execution environment corresponding to the error type.
[0047] After obtaining multiple preset first execution environments and target error types, an error type match can be performed based on the target error type and the error type associated with the first execution environment to obtain a second execution environment with the same error type as the target error type. After obtaining the second execution environment, the natural language question sample can be rewritten based on the first query keyword and the second execution environment to obtain a first natural language question that contains the errors in the first error sample.
[0048] Through the above method, multiple first execution environments can be pre-built, and different first execution environments can correspond to different error types. Therefore, according to the error type of the erroneous structured query statement, the corresponding first execution environment can be selected for subsequent processing. Compared with building a universal execution environment that includes different error types, on the one hand, the data processing volume of the second language model can be reduced, and data processing efficiency can be improved; on the other hand, because the selected execution environment is more closely matched with the current natural language question sample, in the process of rewriting the natural language question sample based on the first query keyword and the second execution environment, the generated first natural language question can meet the actual business needs, that is, data query can be performed in the second execution environment, thereby further improving the quality of the training corpus.
[0049] In a possible manner, the second largest language model may output the first structured query statement in the following manner: The sample natural language question and the erroneous structured query statement in the first error sample are understood to obtain first knowledge for clarifying the erroneous structured query statement, and the sample natural language question and the correct structured query statement in the first error sample are understood to obtain second knowledge for generating the correct structured query statement; based on the first knowledge, the second knowledge, the first natural language question and the execution environment, a first structured query statement corresponding to the first natural language question is generated; and the first structured query statement is output.
[0050] It should be understood that the second largest language model can clarify the knowledge of the incorrect structured query statement by understanding the sample natural language question and the incorrect structured query statement in the first error sample. At the same time, the second largest language model can clarify the knowledge of generating the correct structured query statement by understanding the sample natural language question and the correct structured query statement in the first error sample. As a result, when the second largest language model generates the first structured query statement corresponding to the first natural language question based on the first knowledge, the second knowledge, the first natural language question and the execution environment, the errors contained in the first structured query statement can be considered between the correct knowledge and the existing incorrect knowledge, thereby further improving the diversity and richness of the training corpus.
[0051] It should also be understood that related technologies generally improve training quality through model distillation. However, due to the lack of targeted training for specific problems and error correction, this method is not very targeted. However, in this embodiment, by utilizing the first and second knowledge, structured query statements are generated for known problems in a targeted manner. This can make the generated training corpus more in line with actual needs. Therefore, when training a large language model based on the training corpus, it can help the large language model better learn the characteristics and laws of specific tasks, thereby improving the training effect of the large language model.
[0052] In a possible manner, obtaining training corpus for training the first language model based on the first natural language question and the first structured query statement may include: The first natural language question and the first structured query statement are combined into a first training corpus; the first large language model is trained based on the first training corpus to obtain a third structured query statement generated by the first large language model based on the first natural language question during the training process; when the third structured query statement is an erroneous structured query statement, a second error sample is obtained based on the first natural language question, the first structured query statement and the third structured query statement; the second error sample, the execution environment and the natural language question sample are input into the second large language model to obtain a second natural language question for data query output by the second large language model and a second structured query statement corresponding to the second natural language question; based on the second natural language question and the second structured query statement, a second training corpus for training the first large language model is obtained.
[0053] It should be understood that the third structured query statement generated by the first large language model based on the first natural language question during the training process may be correct or incorrect. Therefore, in order to further improve the diversity and richness of the training corpus, when the generated third structured query statement is incorrect, a new error sample, that is, a second error sample, can be constructed based on the first natural language question, the first structured query statement and the third structured query statement. Then, the execution environment of the second error sample and the natural language question example are input into the second large language model to obtain a new second training corpus for training the first large language model. If there are still errors in the output structured query statement during the training of the large model based on the second training corpus, the above steps can be re-executed, such as Figure 4 shown.
[0054] This approach creates a closed-loop production loop for training data, providing an effective solution for expanding it. Specifically, this closed-loop production process allows for continuous production of training data to address any issues that arise during large model training. This allows for targeted optimization of large model performance, gradually improving its accuracy and reliability when generating structured queries for natural language problems.
[0055] In a possible manner, obtaining training corpus for training the first language model based on the first natural language question and the first structured query statement may include: A syntax check is performed on the first structured query statement to obtain a syntax check result, and based on the syntax check result, a grammatically compliant structured query statement is screened from the first structured query statement; a spot check is performed on the compliant structured query statement to obtain a spot-checked structured query statement, and a matching result between the spot-checked structured query statement and the corresponding first natural language question is determined, wherein the matching result is used to characterize whether the data required for the corresponding first natural language question can be queried through the spot-checked structured query statement; when the matching result meets the preset conditions, a compliant structured query statement and the corresponding first natural language question are combined into a training corpus for training the first large language model.
[0056] For example, if the first structured query statement includes multiple statements, a syntax check can be performed on each first structured query statement. If the syntax check result indicates that the first structured query statement has a syntax problem, the first natural language question and the first structured query statement can be discarded. Alternatively, the first structured query statement can be modified and the modified first structured query statement can be used as a compliant structured query statement. If the syntax check result indicates that the first structured query statement does not have a syntax problem, it can be used as a compliant structured query statement. After obtaining all compliant structured query statements, 10% of the compliant structured query statements can be extracted as spot-checked structured query statements, and for each spot-checked structured query statement, it can be determined whether the spot-checked structured query statement can query the data required by the corresponding first natural language question. If the sampled structured query statement can query the data required for the corresponding first natural language question, then the sampled structured query statement and the corresponding first natural language question are combined into a training corpus for training the first large language model; if the sampled structured query statement cannot query the data required for the corresponding first natural language question, then the sampled structured query statement and the corresponding first natural language question can be discarded, or the sampled structured query statement can be modified accordingly, and then the modified sampled structured query statement and the corresponding first natural language question are combined into a training corpus for training the first large language model.
[0057] Through the above method, a compliant structured query statement can be obtained by grammatically checking the first structured query statement, and a compliant structured query statement can be obtained by spot-checking the compliant structured query statement to obtain a compliant structured query statement that can query the data required for the corresponding first natural language question, thereby obtaining a training corpus for training the first large language model. This can reduce errors in the training corpus and improve the quality of the training corpus. Furthermore, when the large model is trained based on this training corpus, the training effect of the large model can be improved.
[0058] Based on the same concept, the embodiment of the present disclosure also provides a training corpus generation device based on a large language model, such as Figure 5 As shown, the training corpus generation device 500 based on the large language model may include: Acquisition module 501 is configured to acquire a first error sample, an execution environment, and a sample natural language question for data query, wherein the first error sample includes a sample natural language question for data query, an erroneous structured query statement generated by the first language model based on the sample natural language question, and a correct structured query statement corresponding to the erroneous structured query statement, and the execution environment is configured to represent a query data table obtained based on the first error sample; A first processing module 502 is configured to input the first error sample, the execution environment, and the natural language question sample into the second large language model to obtain a first natural language question for data query and a first structured query statement corresponding to the first natural language question, wherein the second large language model is configured to understand the first error sample, obtain a first query keyword that causes the first large language model to generate an erroneous structured query statement, rewrite the natural language question sample based on the first query keyword and the execution environment, output the first natural language question, and output the first structured query statement based at least on the first natural language question and the execution environment; The second processing module 503 is configured to obtain training corpus for training a first large language model based on the first natural language question and the first structured query statement.
[0059] Through the above-mentioned large language model-based training corpus generation device 500, the first error sample, the execution environment, and the natural language question sample for data query can be input into the second large language model to obtain the first natural language question for data query and the first structured query statement corresponding to the first natural language question. Based on the first natural language question and the first structured query statement, the training corpus for training the first large language model is obtained. Since the training corpus is expanded based on the generation of the large language model during the training corpus generation process, the automatic generation of the training corpus can be achieved. On the one hand, the amount of training corpus can be increased, the risk of training corpus exhaustion can be reduced, and the model training demand can be better met when the demand for model training increases; on the other hand, human participation can be reduced, and the efficiency of training corpus generation can be improved. In addition, since the natural language question sample is rewritten based on the error sample during the training corpus generation process, the error information in the training corpus can reduce the interference and redundancy of the remaining noise data while covering the key error features in the error sample, thereby improving the quality stability of the training corpus to a certain extent.
[0060] In addition, since the thinking boundary of the second language model can be defined through the execution environment during the process of generating training corpus, the generated training corpus can meet actual business needs and further improve the quality of the training corpus.
[0061] In a possible manner, the first error sample further includes an error label, and the error label represents the error cause of the incorrectly structured query statement. Accordingly, the first processing module 502 may include: a determination submodule, configured to determine, based on a target error type, a second execution environment from a plurality of preset first execution environments, wherein each first execution environment is pre-associated with an error type of an erroneous structured query statement generated by the first large language model, and the error type associated with the second execution environment is the same as the target error type; The first processing submodule is configured to rewrite the natural language question sample according to the first query keyword and the second execution environment, and output a first natural language question.
[0062] In a possible manner, the plurality of preset first execution environments are obtained in the following manner: Acquire a plurality of different first error samples, and for each first error sample, determine an error type of the erroneous structured query statement in the first error sample based on an error label included in the first error sample; Performing field recognition on the first error samples having the same error type to obtain an initial data field corresponding to the error type, and performing deduplication processing on the initial data field to obtain a deduplication data field corresponding to the error type; A data table is constructed based on the deduplicated data fields corresponding to the error types and the field types corresponding to the deduplicated data fields to obtain a first execution environment corresponding to the error types.
[0063] In a possible embodiment, the first processing module 502 may include: An understanding submodule is configured to understand the sample natural language question and the incorrect structured query statement in the first error sample to obtain first knowledge for clarifying the incorrect structured query statement, and to understand the sample natural language question and the correct structured query statement in the first error sample to obtain second knowledge for generating the correct structured query statement; A generation submodule, configured to generate a first structured query statement corresponding to the first natural language question based on the first knowledge, the second knowledge, the first natural language question, and the execution environment; The output submodule is used to output the first structured query statement.
[0064] In a possible manner, the second processing module 503 may include: A second processing submodule is configured to combine the first natural language question and the first structured query statement into a first training corpus; A training submodule, configured to train the first language model based on the first training corpus to obtain a third structured query statement generated by the first language model based on the first natural language question during the training process; A third processing submodule is configured to obtain a second error sample based on the first natural language question, the first structured query statement, and the third structured query statement when the third structured query statement is an error structured query statement; a fourth processing submodule, configured to input the second error sample, the execution environment, and the natural language question sample into the second large language model, and obtain a second natural language question for data query and a second structured query statement corresponding to the second natural language question output by the second large language model; The fifth processing submodule is configured to obtain a second training corpus for training the first large language model based on the second natural language question and the second structured query statement.
[0065] In a possible manner, the first error sample is obtained by: Obtaining an error definition for the first language model, a sample natural language question for data query, and a third structured query statement generated by the first language model based on the sample natural language question, wherein the error definition includes processing steps required by the first language model to generate the structured query statement, error types in each processing step, and error details corresponding to each error type; The sample natural language question, the third structured query statement and the error definition are input into the second largest language model to obtain a first error sample, wherein the second largest language model is also used to understand the sample natural language question and the third structured query statement based on the error definition to obtain an error structured query statement, generate a correct structured query statement based on the error structured query statement and the sample natural language question corresponding to the error structured query statement, and output the first error sample based on the error structured query statement, the sample question corresponding to the error structured query statement and the correct structured query statement.
[0066] In a possible manner, the second processing module 503 may include: A check submodule, configured to perform syntax check on the first structured query statement to obtain a syntax check result, and select a grammatically compliant structured query statement from the first structured query statement based on the syntax check result; a sixth processing submodule, configured to perform a spot check on the compliant structured query statements to obtain a spot-checked structured query statement, and determine a match result between the spot-checked structured query statement and the corresponding first natural language question, wherein the match result is used to indicate whether the spot-checked structured query statement can be used to query and obtain data required for the corresponding first natural language question; The combining submodule is used to combine a compliant structured query statement and a corresponding first natural language question into a training corpus for training the first large language model when the matching result meets a preset condition.
[0067] Reference below Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing an embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0068] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.
[0069] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0070] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0071] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.
[0072] In some embodiments, communications may be conducted using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0073] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0074] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains a first error sample, an execution environment, and a natural language question sample for data query, wherein the first error sample includes a sample natural language question for data query, an erroneous structured query statement generated by a first large language model based on the sample natural language question, and a correct structured query statement corresponding to the erroneous structured query statement, and the execution environment is used to represent a query data table obtained based on the first error sample; inputs the first error sample, the execution environment, and the natural language question sample into a second large language model to obtain a first natural language question for data query and a first structured query statement corresponding to the first natural language question, wherein the second large language model is used to understand the first error sample, obtain a first query keyword that causes the first large language model to generate the erroneous structured query statement, rewrite the natural language question sample according to the first query keyword and the execution environment, output the first natural language question, and output a first structured query statement based at least on the first natural language question and the execution environment; and obtains training corpus for training the first large language model based on the first natural language question and the first structured query statement.
[0075] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0077] The modules described in the embodiments of the present disclosure may be implemented in software or hardware, wherein the name of a module does not necessarily limit the module itself.
[0078] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0079] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0080] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the scope of the above disclosure. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0081] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0082] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.
Claims
1. A method for generating training corpus based on a large language model, characterized in that: The training corpus generation method based on the large language model includes: Obtaining a first error sample, an execution environment, and a natural language question sample for data query, wherein the first error sample includes a sample natural language question for data query, an erroneous structured query statement generated by a first large language model based on the sample natural language question, and a correct structured query statement corresponding to the erroneous structured query statement, and the execution environment is used to represent a query data table obtained based on the first error sample; Inputting the first error sample, the execution environment, and the natural language question example into a second large language model to obtain a first natural language question for data query and a first structured query statement corresponding to the first natural language question, wherein the second large language model is used to understand the first error sample, obtain a first query keyword that causes the first large language model to generate the erroneous structured query statement, rewrite the natural language question example according to the first query keyword and the execution environment, output the first natural language question, and output the first structured query statement based on at least the first natural language question and the execution environment; Based on the first natural language question and the first structured query statement, a training corpus for training the first language model is obtained.
2. The method for generating training corpus based on a large language model according to claim 1, characterized in that: The first error sample further includes an error label, wherein the error label represents the error cause of the incorrectly structured query statement. The second language model is configured to output the first natural language question in the following manner: Determining a target error type of the erroneous structured query statement based on the error label; Determining a second execution environment from a plurality of preset first execution environments according to the target error type, wherein each of the first execution environments is pre-associated with an error type of the erroneous structured query statement generated by the first large language model, and the error type associated with the second execution environment is the same as the target error type; The natural language question example is rewritten according to the first query keyword and the second execution environment, and the first natural language question is output.
3. The method for generating training corpus based on a large language model according to claim 2, characterized in that: The plurality of preset first execution environments are obtained in the following manner: Acquire a plurality of different first error samples, and for each of the first error samples, determine an error type of the incorrectly structured query statement in the first error sample based on an error label included in the first error sample; Performing field recognition on first error samples having the same error type to obtain an initial data field corresponding to the error type, and performing deduplication processing on the initial data field to obtain a deduplication data field corresponding to the error type; A data table is constructed based on the deduplication data field corresponding to the error type and the field type corresponding to the deduplication data field to obtain a first execution environment corresponding to the error type.
4. The method for generating training corpus based on a large language model according to any one of claims 1 to 3, characterized in that: The second language model is used to output the first structured query statement in the following manner: The sample natural language question and the erroneous structured query statement in the first error sample are understood to obtain first knowledge for clarifying the erroneous structured query statement, and the sample natural language question and the correct structured query statement in the first error sample are understood to obtain second knowledge for generating a correct structured query statement; Generate a first structured query statement corresponding to the first natural language question based on the first knowledge, the second knowledge, the first natural language question, and the execution environment; The first structured query statement is output.
5. The method for generating training corpus based on a large language model according to claim 4, characterized in that: The obtaining, based on the first natural language question and the first structured query statement, a training corpus for training the first language model includes: Combining the first natural language question and the first structured query statement into a first training corpus; Training the first language model based on the first training corpus to obtain a third structured query statement generated by the first language model based on the first natural language question during the training process; In a case where the third structured query statement is an erroneous structured query statement, obtaining a second error sample based on the first natural language question, the first structured query statement, and the third structured query statement; Inputting the second error sample, the execution environment, and the natural language question example into the second large language model, obtaining a second natural language question for data query output by the second large language model and a second structured query statement corresponding to the second natural language question; Based on the second natural language question and the second structured query statement, a second training corpus for training the first language model is obtained.
6. The method for generating training corpus based on a large language model according to any one of claims 1 to 3, characterized in that: The first error sample is obtained in the following manner: Obtaining an error definition for the first language model, a sample natural language question for data query, and a third structured query statement generated by the first language model based on the sample natural language question, wherein the error definition includes processing steps required for the first language model to generate the structured query statement, an error type for each processing step, and error details corresponding to each error type; The sample natural language question, the third structured query statement and the error definition are input into the second largest language model to obtain a first error sample, wherein the second largest language model is also used to understand the sample natural language question and the third structured query statement based on the error definition to obtain an erroneous structured query statement, generate a correct structured query statement based on the erroneous structured query statement and the sample natural language question corresponding to the erroneous structured query statement, and output the first error sample based on the erroneous structured query statement, the sample question corresponding to the erroneous structured query statement and the correct structured query statement.
7. The method for generating training corpus based on a large language model according to any one of claims 1 to 3, characterized in that: The obtaining, based on the first natural language question and the first structured query statement, a training corpus for training the first language model includes: Performing a syntax check on the first structured query statement to obtain a syntax check result, and screening a grammatically compliant structured query statement from the first structured query statement based on the syntax check result; Performing a spot check on the compliant structured query statement to obtain a spot-checked structured query statement, and determining a match result between the spot-checked structured query statement and the corresponding first natural language question, wherein the match result is used to indicate whether the spot-checked structured query statement can be used to query and obtain data required for the corresponding first natural language question; When the matching result satisfies a preset condition, the compliant structured query statement and the corresponding first natural language question are combined into a training corpus for training the first large language model.
8. A training corpus generation device based on a large language model, characterized in that: The training corpus generation device based on the large language model includes: an acquisition module, configured to acquire a first error sample, an execution environment, and a sample natural language question for data query, wherein the first error sample includes a sample natural language question for data query, an erroneous structured query statement generated by a first large language model based on the sample natural language question, and a correct structured query statement corresponding to the erroneous structured query statement; and the execution environment is used to represent a query data table obtained based on the first error sample; a first processing module configured to input the first error sample, the execution environment, and the natural language question example into a second large language model to obtain a first natural language question for data query and a first structured query statement corresponding to the first natural language question, wherein the second large language model is configured to understand the first error sample, obtain a first query keyword that causes the first large language model to generate the erroneous structured query statement, rewrite the natural language question example based on the first query keyword and the execution environment, output the first natural language question, and output the first structured query statement based at least on the first natural language question and the execution environment; The second processing module is used to obtain training corpus for training the first language model based on the first natural language question and the first structured query statement.
9. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Model training and corpus generation method and device, equipment and storage medium
CN114281968A
Method and device for querying data based on large model, electronic equipment and storage medium
CN119149582A
Natural language query sample generation method and device, equipment, medium and product
CN119226437A
Generation method and device of language model for automatically generating query statement
CN119396855A
Method and apparatus for outputting structured query sentence
US20210200763A1