Big model-based training data generation method, model training method and device
By reflecting on and rewriting questions and retrieval fragments through a large model, automatically correcting errors, and combining multi-model review, the problem of retrieval errors in RAG training data generation was solved, data quality and generation efficiency were improved, and the effectiveness of the knowledge question-answering model was enhanced.
Patent Information
- Application Number
- CN202411899226.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-12-20
AI Technical Summary
The existing technology lacks effective means to generate RAG training data, resulting in the inability to automatically correct retrieval errors, affecting the quality of training data and generation efficiency.
Through large-scale model reflection and rewriting of questions, search fragment errors are automatically corrected, and an iterative reflection mechanism is used to generate corrected answers. Multiple large models are combined for review to ensure the quality of the answers.
It improves the quality and generation efficiency of RAG training data, reduces the failure of training data generation due to retrieval errors, and enhances the effectiveness of the knowledge question-answering model in business scenarios.
Smart Images

Figure CN119862272B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computers, in particular to the technical field of artificial intelligence such as natural language processing, large models, intelligent search, knowledge graphs, and the like, and specifically relates to a large model-based training data generation method, a knowledge question and answer model training method, an apparatus, and an electronic device, which can be applied to knowledge question and answer scenarios. BACKGROUND
[0002] A RAG (Retrieval-Augmented Generation) system is a model combining retrieval and generation capabilities, which answers questions by retrieving relevant information from a large amount of text and generating answers. However, there is currently a lack of effective means for generating RAG training data. SUMMARY
[0003] The present disclosure provides a large model-based training data generation method, a knowledge question and answer model training method, an apparatus, an electronic device, and a storage medium.
[0004] According to a first aspect of the present disclosure, a large model-based training data generation method is provided, comprising:
[0005] Based on historical operation data, N triple data are obtained, wherein N is a positive integer; each triple data includes a question, an answer corresponding to the question, and a retrieval segment;
[0006] From the N triple data, triple data with incorrect answers are selected as to-be-corrected triple data;
[0007] Based on the iteration reflection of the large model, the question in the to-be-corrected triple data is rewritten, and a corrected retrieval segment is generated based on the rewritten question and the large model;
[0008] Based on the rewritten question and the corrected retrieval segment, the large model is used to generate a corrected answer corresponding to the rewritten question;
[0009] Based on the rewritten question, the corrected retrieval segment, and the corrected answer, the to-be-corrected triple data is updated to obtain RAG (Retrieval-Augmented Generation) training data.
[0010] According to a second aspect of the present disclosure, a knowledge question and answer model training method is provided, comprising:
[0011] Obtaining training data, the training data being generated based on the large model-based training data generation method of the first aspect;
[0012] Training a knowledge question and answer model based on the training data.
[0013] According to a third aspect of the present disclosure, a large model-based training data generation apparatus is provided, comprising:
[0014] an acquisition module configured to acquire N triple data based on historical operation data, N being a positive integer; wherein each of the triple data comprises a question, an answer corresponding to the question, and a retrieval fragment;
[0015] a screening module configured to screen out triple data with incorrect answers from the N triple data as to-be-corrected triple data;
[0016] a rewriting module configured to rewrite the question in the to-be-corrected triple data based on iterative reflection of the large model, and generate a corrected retrieval fragment based on the rewritten question and the large model;
[0017] a generation module configured to generate a corrected answer corresponding to the rewritten question by using the large model based on the rewritten question and the corrected retrieval fragment;
[0018] an update module configured to update the to-be-corrected triple data based on the rewritten question, the corrected retrieval fragment, and the corrected answer, to obtain retrieval enhancement generation (RAG) training data.
[0019] According to a fourth aspect of the present disclosure, a knowledge question and answer model training apparatus is provided, comprising:
[0020] an acquisition module configured to acquire training data, the training data being generated based on the large model-based training data generation method of the first aspect;
[0021] a training module configured to train a knowledge question and answer model based on the training data.
[0022] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising:
[0023] at least one processor; and
[0024] a memory communicatively connected to the at least one processor; wherein
[0025] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of the first aspect and the second aspect.
[0026] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the method of the first aspect and the second aspect.
[0027] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the steps of the method of the first aspect and the second aspect when executed by a processor.
[0028] According to the technology of the present disclosure, by letting the large model reflect the rewriting problem, the problem of automatically correcting the retrieval fragment error can be solved, the problem that the RAG data generation tool in the prior art cannot automatically correct the retrieval error can be solved, the quality of the RAG training data can be improved, and the generation efficiency of the training data is improved.
[0029] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0030] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0031] Figure 1 is a flowchart of the training data generation method based on a large model provided by the embodiments of the present disclosure;
[0032] Figure 2 is a flowchart of the training data generation method based on a large model provided by the embodiments of the present disclosure;
[0033] Figure 3 is a flowchart of the training data generation method based on a large model provided by the embodiments of the present disclosure;
[0034] Figure 4 is a flowchart of the training data generation method based on a large model provided by the embodiments of the present disclosure;
[0035] Figure 5 is a flowchart of the knowledge question and answer model training method provided by the embodiments of the present disclosure;
[0036] Figure 6 is a block diagram of the training data generation device based on a large model provided by the embodiments of the present disclosure;
[0037] Figure 7 is a block diagram of the knowledge question and answer model training device provided by the embodiments of the present disclosure;
[0038] Figure 8 is a block diagram of an electronic device for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION
[0039] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0040] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.
[0041] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0042] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0043] It is worth noting that in the embodiments of the present disclosure, certain software, components, models, etc. that already exist in the industry may be mentioned. They should be considered as exemplary and their purpose is only to illustrate the feasibility of implementing the technical solution of the present disclosure, but it does not mean that the applicant has or will necessarily use the solution.
[0044] The embodiments of the present disclosure relate to artificial intelligence technology fields such as natural language processing, large models, intelligent search, and knowledge graphs.
[0045] Artificial Intelligence, abbreviated as AI. It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence.
[0046] Natural Language Processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between people and computers using natural language. It is a discipline that uses computer technology to analyze, understand and process natural language, that is, it uses computers as a powerful tool for language research, and under the support of computers, it conducts quantitative research on language information and provides language descriptions that can be used by people and computers.
[0047] Large model refers to a "large parameter" model trained using large-scale data and powerful computing power. These models usually have high versatility and generalization ability and can be applied to natural language processing, image recognition, speech recognition and other fields. For example, the large model in this paper can be a large language model (LLM), which refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. Large models can handle a variety of natural language tasks such as text classification, question answering, and dialogue, and are an important way to artificial intelligence. Alternatively, the large model can also be other large models, such as multi-modal large models or basic large models, etc.
[0048] Intelligent search (Intelligent Retrieval) is based on the relevance of documents and search terms, and considers the importance of documents and other indicators to sort the search results to provide higher search efficiency. The result sorting of intelligent retrieval considers relevance and importance at the same time, relevance uses weighted mixed index of each field, relevance analysis is more accurate, importance refers to the evaluation of document quality through analysis of document source authority and citation relationship, such result sorting is more accurate, and can put the documents most relevant to user's wishes in the front, improving search efficiency.
[0049] A knowledge graph is a semantic network used to describe the relationships between entities. It is a semi-structured data representation method used to describe entities, attributes, and relationships between entities. Knowledge graphs are commonly used to build intelligent search engines, recommendation systems, question and answer systems, and other artificial intelligence applications. The core idea of a knowledge graph is to transform information in the real world into a graph, where nodes represent entities and edges represent relationships between entities. A knowledge graph is not just a graphical knowledge base, it also contains semantic descriptions of entities and relationships that can be understood and processed by computers. The construction of a knowledge graph requires information extraction, fusion, and reasoning, which usually requires the use of natural language processing, image recognition, machine learning, and other technologies.
[0050] A RAG (Retrieval-Augmented Generation) system is a model that combines retrieval and generation capabilities, answering questions by retrieving relevant information from a large amount of text and generating answers. However, there is currently a lack of effective means for generating RAG training data.
[0051] Therefore, the present disclosure provides a large model-based training data generation method, device and electronic equipment. By having the large model reflect and rewrite the problem, the present disclosure can automatically correct the retrieval fragment error, improve the quality and efficiency of RAG training data. The large model-based training data generation method provided by the present disclosure can be applied to enterprise office scenarios, including but not limited to internal knowledge question and answer. For example, the training data generation method provided by the present disclosure can be applied to knowledge question and answer model training, which can quickly generate a large number of high-quality RAG training data based on historical operation data, and further improve the effect of knowledge question and answer model in various business scenarios.
[0052] The large model-based training data generation method, device and electronic equipment of the present disclosure are described below with reference to the accompanying drawings.
[0053] It should be noted that the execution subject of the large model-based training data generation method of the present disclosure can be a training data generation device, which can be implemented by software and / or hardware, and can be configured in an electronic device. For example, the electronic device can include but is not limited to a terminal, a server, etc.
[0054] Figure 1 is a flowchart of the large model-based training data generation method provided by the present disclosure. As shown in Figure 1 The large model-based training data generation method can include but is not limited to the following steps.
[0055] In step 101, based on historical operation data, N triple data are obtained, N being a positive integer; wherein each triple data includes a question, an answer corresponding to the question, and a retrieval fragment.
[0056] In some embodiments, the collected historical operation data can be parsed and processed, and the question (which can refer to the input Query), the answer, and the retrieval fragment are parsed from the historical operation data to construct triple data. For example, the historical operation data can be the historical operation data of the RAG system, that is, the historical operation data of the RAG system is used, which includes the operation data recorded each time the RAG system is used, including the input user Query (question), the retrieval fragment obtained based on the user Query, and the answer generated based on the user Query and the retrieval fragment.
[0057] In some embodiments, the RAG system can implement knowledge question and answer based on a large model. The retrieval fragment obtained based on the user Query can be obtained by using a large model, that is, the large model retrieves based on the user Query to obtain the retrieval fragment. For example, the answer generated based on the user Query and the retrieval fragment can be obtained by using a large model, that is, the large model generates an answer based on the user Query and the retrieval fragment.
[0058] In some embodiments, the above-mentioned retrieval fragment can refer to the retrieval result obtained based on the user Query, for example, which can include but is not limited to the retrieved document and / or document fragment, etc.
[0059] In step 102, the triple data with incorrect answers are filtered out from the N triple data as the to-be-corrected triple data.
[0060] In some embodiments, the triple data with incorrect answers can be filtered out from the N triple data as the to-be-corrected triple data by using a large model.
[0061] In one possible implementation, based on the question in the N triple data and the retrieval fragment corresponding to the question, a first prompt information is constructed by using a first prompt template; based on the first prompt information and a large model, a reference answer corresponding to the question is obtained; based on the answer corresponding to the question and the reference answer corresponding to the question, the triple data with incorrect answers are filtered out from the N triple data as the to-be-corrected triple data.
[0062] For each of the N triple data, the question in the triple data, the search snippet corresponding to the question, can be filled into a first prompt template to construct a first prompt information. The first prompt information is input into the large model, and a reference answer to the question is generated by the large model. Based on the answer corresponding to the question in the triple data and the reference answer corresponding to the question, the triple data with incorrect answers is screened from the N triple data as the triple data to be corrected.
[0063] For example, the first prompt information can be expressed as follows:
[0064] "Current user question (Query): Question A
[0065] Search snippet corresponding to the current question A: Search snippet a
[0066] Please generate an answer based on the user question and the search snippet".
[0067] After the first prompt information is constructed, the first prompt information can be input into the large model. The large model can generate a corresponding answer based on the question A and the search snippet a, and the answer generated by the large model is taken as the reference answer Aa of the question A. The answer corresponding to the question A in the triple data is obtained, and the answer corresponding to the question A in the triple data is compared with the reference answer Aa of the question A obtained by the large model, for example, similarity calculation can be performed. If the answer corresponding to the question A in the triple data is consistent with the reference answer Aa of the question A, or the similarity between the answer corresponding to the question A in the triple data and the reference answer Aa of the question A is greater than or equal to a threshold value, it can be considered that the answer of the large model to the question A is correct, that is, the triple data is the triple data with correct answer. If the answer corresponding to the question A in the triple data is inconsistent with the reference answer Aa of the question A, or the similarity between the answer corresponding to the question A in the triple data and the reference answer Aa of the question A is less than the threshold value, it can be considered that the answer of the large model to the question A is incorrect, that is, the triple data is the triple data with incorrect answer.
[0068] In step 103, the question in the triple data to be corrected is rewritten based on the iterative reflection of the large model, and the corrected search snippet is generated based on the rewritten question and the large model.
[0069] In some embodiments, the search snippet in the to-be-corrected triple data can be evaluated for accuracy and / or relevance to obtain an evaluation result (e.g., including accuracy and / or relevance) of the search snippet. Based on the evaluation result of the search snippet, the large model is iteratively reflected to rewrite the question to generate a new search snippet, and the new search snippet is continuously evaluated and the question is rewritten until a satisfactory search result is obtained, for example, until the evaluation result of the generated search snippet meets the preset condition. After completing the rewriting of the question, the large model can search based on the rewritten question to obtain a search snippet of the question, and the search snippet is taken as the corrected search snippet.
[0070] In some embodiments, the large model has reflection capability, which refers to the ability of the model to evaluate and improve its own output to improve the quality and accuracy of the output. The importance of the reflection capability of the large model is reflected in the following aspects: (1) error correction, the model can identify and correct its own errors, which is crucial to improve the quality of the output; (2) learning efficiency, through self-reflection, the model can learn from errors more quickly; (3) adaptability, self-reflection enables the model to adapt to new or unseen tasks, and optimizes performance through self-adjustment; (4) robustness, enhances the robustness of the model, so that it can maintain stable performance in the face of uncertainty and noise. In the embodiments of the present disclosure, the iterative reflection of the large model can refer to that after the large model obtains a search snippet based on a question and generates an answer based on the question and the search snippet, the large model reflects once based on the reflection result to continue rewriting the question. Through continuous reflection and rewriting of the question, the large model can output a satisfactory search result.
[0071] It is worth noting that in the embodiments of the present disclosure, the question in the to-be-corrected triple data is rewritten based on the iterative reflection of the large model, and the purpose is to automatically correct the search result through the iterative reflection of the large model to reduce the situation that training data cannot be generated due to search errors.
[0072] In step 104, based on the rewritten question and the corrected search snippet, a large model is used to generate a corrected answer corresponding to the rewritten question.
[0073] In some embodiments, a prompt template can be used to construct a prompt information based on the rewritten question and the corrected search snippet, and the prompt information can be used to instruct the large model to generate an answer based on the rewritten question and the corrected search snippet. The prompt information is input into the large model, and the large model can generate a corrected answer corresponding to the rewritten question based on the prompt information.
[0074] For example, the prompt information can be represented as follows:
[0075] “Rewritten question: question A1
[0076] The modified search snippet (i.e., the search snippet obtained based on the question A1): search snippet a1
[0077] Please generate an answer based on the rewritten question and the modified search snippet.
[0078] After constructing the prompt information, the prompt information can be input into the large model, and the large model can generate a corresponding answer based on the question A1 and the search snippet a1, and the answer generated by the large model is taken as the updated answer corresponding to the question A1.
[0079] In step 105, based on the rewritten question, the modified search snippet, and the corrected answer, the to-be-corrected triple data is updated to obtain RAG training data.
[0080] In some embodiments, the question in the to-be-corrected triple data can be replaced by the rewritten question, the search snippet in the to-be-corrected triple data can be replaced by the modified search snippet, and the answer in the to-be-corrected triple data can be replaced by the corrected answer, thereby obtaining the RAG training data.
[0081] In the above embodiments, the present disclosure automatically generates a large number of RAG training data with high quality based on online real historical operation data, can fully exert the advantages of the field scene, can continuously fit the real user data distribution after optimization, and finally achieves better question and answer effect in the real business scene. By introducing the reflection-based search correction mechanism, the present disclosure can automatically correct the part of the search error, reduce the situation that training data cannot be generated due to search error, thereby improving the training data quality and the generation efficiency of the training data, and solving the problem that the RAG data generation tool in the prior art cannot automatically correct the search error.
[0082] Figure 2 is a flowchart of the training data generation method based on a large model provided by the embodiments of the present disclosure. As shown in Figure 2 , on the basis of the embodiments shown in Figure 1 , the optional implementation manner of the above-mentioned large model-based iterative reflection for rewriting the question in the to-be-corrected triple data includes but is not limited to the following steps.
[0083] In step 201, the search snippet in the to-be-corrected triple data is evaluated to obtain an evaluation result of the search snippet.
[0084] In some embodiments, the accuracy and / or correlation of the search snippet in the to-be-corrected triple data can be evaluated based on the to-be-corrected triple data to obtain an evaluation result of the search snippet, and the evaluation result includes the accuracy and / or correlation.
[0085] Exemplarily, based on the to-be-corrected triple data, a large model is used to evaluate the accuracy and / or relevance of the retrieval fragment in the to-be-corrected triple data. For example, the large model can be used to evaluate the accuracy and relevance of the retrieval fragment in the to-be-corrected triple data, and the evaluation result of the retrieval fragment is obtained. In some embodiments, the relevance can be used to represent the relevance of the retrieval fragment and the question.
[0086] In some embodiments, the accuracy can be used to represent whether the retrieval fragment has no answer and / or whether the retrieval fragment is complete. Exemplarily, the accuracy can be represented as: whether the answer generated based on the retrieval fragment is correct based on the question.
[0087] In step 202, based on the to-be-corrected triple data and the evaluation result of the retrieval fragment, a second prompt information is constructed by using a second prompt template.
[0088] In some embodiments, the retrieval fragment in the to-be-corrected triple data, the user question (Query) and the evaluation result of the retrieval fragment can be filled into the second prompt template to construct the second prompt information, wherein the second prompt information can be used to instruct the large model to rewrite the question. For example, the second prompt information can be represented as follows:
[0089] “Current retrieval fragment: XXXX
[0090] Evaluation result of the current retrieval fragment: accuracy is XX, relevance is XX
[0091] Please rewrite the Query based on the user question (Query) and the evaluation result of the current retrieval fragment, so that the answer can be correct.”
[0092] In step 203, based on the second prompt information and the large model, the question in the to-be-corrected triple data is rewritten to obtain a candidate question, and the retrieval fragment is generated based on the candidate question.
[0093] In some embodiments, the second prompt information can be input into the large model, and the large model can rewrite the question in the to-be-corrected triple data based on the second prompt information. The question rewritten at this time is used as a candidate question, and the relevant fragment is retrieved based on the candidate question. The retrieved fragment is used as the retrieval fragment corresponding to the candidate question.
[0094] In step 204, based on the candidate question and the generated retrieval fragment, a first answer is generated by using the large model.
[0095] In some embodiments, the candidate question and the generated retrieval snippet can be filled into a preset prompt template to construct a prompt information, and the prompt information can be input into the large model, which can be used to instruct the large model to generate an answer. Through the large model, a corresponding answer, i.e., the first answer, is generated based on the prompt information, in combination with the candidate question and the generated retrieval snippet.
[0096] Optionally, in some embodiments, based on the candidate question and the generated retrieval snippet, a plurality of large models can be used to generate a plurality of candidate answers, and the first answer can be obtained from the plurality of candidate answers. For example, a prompt template can be used to construct a prompt information, e.g., the candidate question and the generated retrieval snippet can be filled into a preset prompt template to construct the prompt information. The prompt information can be input into a plurality of large models, each of which generates a corresponding answer based on the prompt information, in combination with the candidate question and the generated retrieval snippet, i.e., a candidate answer generated by each large model. Then, the first answer can be obtained from the plurality of candidate answers. For example, the best answer can be selected from the plurality of candidate answers as the first answer.
[0097] In some embodiments, a third prompt template can be used to construct a third prompt information, which is used to instruct the large model to score a plurality of candidate answers in at least one scoring indicator dimension. The at least one scoring indicator dimension includes one or more of the following: answer accuracy, relevance of the answer to the retrieval snippet, completeness of the answer, and readability of the answer format. Based on the third prompt information, a plurality of large models can be used to score the plurality of candidate answers in at least one scoring indicator dimension, to obtain scoring information of the plurality of candidate answers. The first answer can be obtained from the plurality of candidate answers based on the scoring information of the plurality of candidate answers.
[0098] Exemplarily, the plurality of large models can be employed to score the plurality of candidate answers in at least one scoring indicator dimension. In a possible implementation, each of the plurality of large models can score the plurality of candidate answers in at least one scoring indicator dimension. For example, assuming that the plurality of large models includes large model A and large model B, and the candidate answers include candidate answer 1 and candidate answer 2, large model A can score candidate answer 1 in answer accuracy, answer relevance to the retrieval snippet, answer completeness, answer format readability, and the like, obtain scores of the candidate answer 1 in each dimension, and average the scores in the dimensions to obtain an average value as the score of the candidate answer 1 by large model A. Large model A can score candidate answer 2 in answer accuracy, answer relevance to the retrieval snippet, answer completeness, answer format readability, and the like, obtain scores of the candidate answer 2 in each dimension, and average the scores in the dimensions to obtain an average value as the score of the candidate answer 2 by large model A. Large model B scores candidate answer 1 and candidate answer 2 in the above manner of large model A to obtain the score of candidate answer 1 by large model B and the score of candidate answer 2 by large model B. Based on the scores of candidate answer 1 by large model A and large model B, final score information of the candidate answer 1 is calculated, based on the scores of candidate answer 2 by large model A and large model B, final score information of the candidate answer 2 is calculated, and based on the final score information of candidate answer 1 and the final score information of candidate answer 2, the candidate answer with the highest score is selected from candidate answer 1 and candidate answer 2 as the first answer.
[0099] In another possible implementation, each of the plurality of large models can score one candidate answer, exemplarily, the large model can score the candidate answer generated by the large model. For example, assuming that the plurality of large models includes large model A and large model B, and the candidate answers include candidate answer 1 and candidate answer 2, candidate answer 1 is generated by large model A, and candidate answer 2 is generated by large model B, large model A can score candidate answer 1 in answer accuracy, answer relevance to the retrieval snippet, answer completeness, answer format readability, and the like, obtain scores of the candidate answer 1 in each dimension, and average the scores in the dimensions to obtain an average value as the score of the candidate answer 1 by large model A. Large model B scores candidate answer 2 in the above manner of large model A to obtain the score of candidate answer 2 by large model B, and based on the score of candidate answer 1 and the score of candidate answer 2, the candidate answer with the highest score is selected from candidate answer 1 and candidate answer 2 as the first answer, which can be greater than or equal to a threshold value (such as 0.8) optionally.
[0100] That is, multiple large models can be used to review multiple candidate answers, for example, by having multiple large models act as reviewers, voting on each large model, such as having each large model vote on all other large models, and determining the candidate answer generated by the large model with the most votes as the first answer. As can be seen, by combining multiple large models to generate answers and using multiple large models for review, the best answer is selected, which can further ensure the quality of training data.
[0101] In step 205, based on the first answer generated by the large model and the candidate question, the generated retrieval segment is evaluated to obtain an evaluation result of the generated retrieval segment.
[0102] In some embodiments, the accuracy, relevance and relevance of the generated retrieval segment can be evaluated based on the first answer generated by the large model and the candidate question. For example, the accuracy can be represented as: based on the rewritten question, a new retrieval segment is obtained, and whether the answer generated based on the new retrieval segment is correct. The relevance can be represented as: the relevance of the rewritten question and the original question. The relevance can be represented as: whether the new answer and the rewritten question are related.
[0103] In step 206, in the case where the evaluation result of the generated retrieval segment does not meet the preset condition, reflection iteration is performed to continue to evaluate and rewrite the candidate question until the evaluation result of the generated retrieval segment meets the preset condition.
[0104] For example, the evaluation result of the retrieval segment meeting the preset condition can include: the accuracy of the retrieval segment is greater than or equal to the accuracy threshold; the relevance of the retrieval segment is greater than or equal to the relevance threshold. Correspondingly, the evaluation result of the retrieval segment not meeting the preset condition can include: the accuracy of the retrieval segment is less than the accuracy threshold; the relevance of the retrieval segment is less than the relevance threshold.
[0105] In the above embodiments, through the iterative reflection of the large model, the retrieval error can be automatically corrected, and the situation that training data cannot be generated due to retrieval error can be reduced.
[0106] Figure 3 is a flowchart of the large model-based training data generation method provided by the embodiments of the present disclosure. As shown in Figure 3 The large model-based training data generation method can include but is not limited to the following steps.
[0107] In step 301, based on historical operation data, N triple data are obtained, N is a positive integer; wherein each triple data includes a question, an answer corresponding to the question and a retrieval segment.
[0108] Optionally, step 301 can be implemented by any of the implementation manners of the embodiments of the present disclosure respectively, and the embodiments of the present disclosure do not make any limitation thereto, and will not be repeated here.
[0109] In step 302, the triple data with answer errors are filtered out from the N triple data as the triple data to be corrected.
[0110] Optionally, step 302 can be implemented by any of the implementation manners of the embodiments of the present disclosure respectively, and the embodiments of the present disclosure do not make any limitation thereto, and will not be repeated here.
[0111] In step 303, the problem in the triple data to be corrected is rewritten based on the iterative reflection of the large model, and the corrected retrieval segment is generated based on the rewritten problem and the large model.
[0112] Optionally, step 303 can be implemented by any of the implementation manners of the embodiments of the present disclosure respectively, and the embodiments of the present disclosure do not make any limitation thereto, and will not be repeated here.
[0113] In step 304, a plurality of candidate answers are generated based on the rewritten problem and the corrected retrieval segment by using a plurality of large models.
[0114] For example, after completing the problem rewriting, i.e., completing the automatic correction of the retrieval error part, the answer can be generated based on the rewritten problem and the corrected retrieval segment by using a plurality of large models. The prompt information can be constructed by using the prompt template, for example, the above-mentioned rewritten problem and the corrected retrieval segment are filled into the preset prompt template to construct the prompt information. The prompt information is input into a plurality of large models, and each large model generates a corresponding answer based on the prompt information combined with the rewritten problem and the corrected retrieval segment, i.e., the candidate answer generated by each large model is obtained.
[0115] In step 305, the correct answer is obtained from the plurality of candidate answers.
[0116] In some embodiments, the plurality of candidate answers can be scored in at least one scoring indicator dimension by using a plurality of large models to obtain scoring information of the plurality of candidate answers; and the correct answer is obtained from the plurality of candidate answers based on the scoring information of the plurality of candidate answers. The optional implementation manner of obtaining the correct answer from the plurality of candidate answers can refer to the optional implementation manner of obtaining the first answer from the plurality of candidate answers in the examples after step 204, which will not be repeated here.
[0117] That is, multiple large models can be used to review multiple candidate answers, for example, by having multiple large models act as reviewers, voting on each large model, for example, having each large model vote on all other large models, and determining the candidate answer generated by the large model with the most votes as the corrected answer.
[0118] In step 306, the to-be-corrected triple data is updated based on the rewritten question, the corrected retrieval fragment, and the corrected answer, to obtain RAG training data.
[0119] Optionally, step 306 can be implemented by using any one of the implementation manners in the embodiments of the present disclosure, and the present disclosure does not limit this, and will not be repeated here.
[0120] In the above embodiments, the present disclosure can automatically generate a large number of RAG training data with high quality based on real online historical operation data, can fully exert the advantages of the field scene, can continuously fit the real user data distribution after optimization, and can finally achieve better question and answer effects in real business scenarios. By introducing the reflection-based retrieval correction mechanism, the present disclosure can automatically correct the part of the retrieval error, reduce the situation that training data cannot be generated due to retrieval error, thereby improving the training data quality and generation efficiency, and solving the problem that the RAG data generation tool in the prior art cannot automatically correct the retrieval error. The present disclosure can further generate answers by combining multiple large models, and review by using multiple large models, and select the best answer, which can further ensure the quality of the training data.
[0121] The large model-based training data generation method provided by the embodiments of the present disclosure can mainly include the following three parts: an online historical operation data acquisition stage, a retrieval data automatic generation stage, and an answer automatic generation stage. As shown in Figure 4 the online historical operation data acquisition stage can collect online historical operation data, parse <user Query, large model answer, retrieval fragment> triple data from the historical operation data, and select triple data with incorrect answers from the historical operation data by using a large model.
[0122] As shown in Figure 4 after the triple data with incorrect answers is selected by using the large model, the accuracy, relevance, and correlation of the retrieval fragment in the triple data can be evaluated, if the evaluation result of the retrieval fragment does not meet the preset condition, for example, the retrieval fragment has no answer or is incomplete, or the correlation between the retrieval fragment and the user question is poor, reflection iteration is performed to generate a new retrieval fragment. The new retrieval fragment generation method can be as follows:
[0123] The query is iteratively rethought and rewritten by using a large model, so as to generate a new retrieval fragment. The specific generation method is as follows: the rewritten query is generated by using a prompt template, the relevant fragment is retrieved by using a retriever, and a new answer is generated by using a large model. Subsequently, the retrieved fragment is evaluated in terms of accuracy, relevance and correlation. If the retrieved fragment has no answer or is incomplete, or the correlation with the new query is not high, the query is iteratively rethought and rewritten, and the query is continuously evaluated and rewritten until a satisfactory retrieval result is obtained. The relevance refers to the relevance between the rewritten question and the original question. The accuracy refers to whether the answer generated based on the rewritten question and the new retrieval fragment is correct. The correlation refers to whether the new answer is related to the rewritten query.
[0124] In the answer automatic generation stage: multiple large models can be used to generate answers, and multiple large models can be used for review to select the best answer.
[0125] In summary, the model, method and idea adopted by the present disclosure are not dependent on products and are applicable to any knowledge Q&A scene. For example, the present disclosure can be applied to a knowledge Q&A model construction scene, the retrieval result can be automatically corrected by reflection, and the answer effect can be gradually improved by introducing multiple model answering and review, so that the answer is more in line with user habits.
[0126] Figure 5 is a flowchart of a knowledge Q&A model training method provided by an embodiment of the present disclosure. As shown in Figure 5 The knowledge Q&A model training method can include but is not limited to the following steps.
[0127] In step 501, training data is obtained.
[0128] In an embodiment of the present disclosure, the training data can be generated based on the large model-based training data generation method described in any of the preceding embodiments, which will not be described here.
[0129] In step 502, the knowledge Q&A model is trained based on the training data.
[0130] In the above embodiment, a large number of RAG training data with high quality can be automatically generated based on historical operation data. Training the knowledge Q&A model based on the RAG training data can improve the effect of the knowledge Q&A model in various business scenarios.
[0131] Figure 6 is a block diagram of a large model-based training data generation device provided by an embodiment of the present disclosure. As shown in Figure 6As shown, the large model-based training data generation apparatus can include an acquisition module 601, a screening module 602, a rewriting module 603, a generation module 604, and an updating module 605.
[0132] The acquisition module 601 is configured to acquire N triple data based on historical operation data, where N is a positive integer. Each triple data includes a question, an answer corresponding to the question, and a retrieval segment.
[0133] The screening module 602 is configured to screen out triple data with incorrect answers from the N triple data as to-be-corrected triple data. In some embodiments, the screening module 602 is configured to: construct first prompt information by using a first prompt template based on the question in the N triple data and the retrieval segment corresponding to the question; obtain a reference answer corresponding to the question based on the first prompt information and a large model; and screen out triple data with incorrect answers from the N triple data as to-be-corrected triple data based on the answer corresponding to the question and the reference answer corresponding to the question.
[0134] The rewriting module 603 is configured to rewrite the question in the to-be-corrected triple data based on iterative reflection of the large model, and generate a corrected retrieval segment based on the rewritten question and the large model.
[0135] In some embodiments, the rewriting module 603 is configured to: evaluate the retrieval segment in the to-be-corrected triple data to obtain an evaluation result of the retrieval segment; construct second prompt information by using a second prompt template based on the to-be-corrected triple data and the evaluation result of the retrieval segment, where the second prompt information is used to instruct the large model to rewrite the question; rewrite the question in the to-be-corrected triple data based on the second prompt information and the large model to obtain a candidate question, and generate a retrieval segment based on the candidate question; generate a first answer by using the large model based on the candidate question and the generated retrieval segment; evaluate the generated retrieval segment based on the first answer generated by the large model and the candidate question to obtain an evaluation result of the generated retrieval segment; and in a case where the evaluation result of the generated retrieval segment does not satisfy a preset condition, perform reflection iteration to continue to evaluate and rewrite the candidate question until the evaluation result of the generated retrieval segment satisfies the preset condition.
[0136] In some embodiments, the rewriting module 603 is configured to: evaluate the retrieval segment in the to-be-corrected triple data in terms of accuracy and / or correlation to obtain an evaluation result of the retrieval segment, where the evaluation result includes the accuracy and / or the correlation; and the accuracy is used to represent whether the retrieval segment has no answer and / or the retrieval segment is complete, and the correlation is used to represent the correlation between the retrieval segment and the question.
[0137] In some embodiments, the rewriting module 603 is configured to generate a plurality of candidate answers based on the candidate question and the generated retrieval segment; and obtain the first answer from the plurality of candidate answers.
[0138] In some embodiments, the rewriting module 603 is configured to construct third prompt information using a third prompt template, the third prompt information being used to instruct the large model to score the plurality of candidate answers in at least one scoring indicator dimension; the at least one scoring indicator dimension comprises one or more of: answer accuracy, relevance of the answer to the retrieval segment, completeness of the answer, answer format readability; based on the third prompt information, the plurality of large models are used to score the plurality of candidate answers in the at least one scoring indicator dimension to obtain scoring information of the plurality of candidate answers; and the first answer is obtained from the plurality of candidate answers based on the scoring information of the plurality of candidate answers.
[0139] The generating module 604 is configured to generate a corrected answer corresponding to the rewritten question based on the rewritten question and the corrected retrieval segment using a large model.
[0140] In some embodiments, the generating module 604 is configured to generate a plurality of candidate answers based on the rewritten question and the corrected retrieval segment using a plurality of large models; and obtain the corrected answer from the plurality of candidate answers.
[0141] In some embodiments, the generating module 604 is configured to score the plurality of candidate answers in at least one scoring indicator dimension using a plurality of large models to obtain scoring information of the plurality of candidate answers; the at least one scoring indicator dimension comprises one or more of: answer accuracy, relevance of the answer to the retrieval segment, completeness of the answer, answer format readability; and the corrected answer is obtained from the plurality of candidate answers based on the scoring information of the plurality of candidate answers.
[0142] The updating module 605 is configured to update the to-be-corrected triple data based on the rewritten question, the corrected retrieval segment, and the corrected answer to obtain retrieval enhancement generation RAG training data.
[0143] As to the apparatus in the above-mentioned embodiments, the specific manners in which various modules perform operations have been described in details in the embodiments of the method, and thus will not be described in details here.
[0144] Figure 7 is a block diagram of a knowledge question and answer model training apparatus provided by the embodiments of the present disclosure. As shown in Figure 7 The knowledge question and answer model training apparatus can include an obtaining module 701 and a training module 702.
[0145] The obtaining module 701 is configured to obtain training data. In an embodiment of the present disclosure, the training data can be generated based on the large model-based training data generation method described in any of the preceding embodiments.
[0146] The training module 702 is configured to train the knowledge question and answer model based on the training data.
[0147] As for the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described here in detail.
[0148] According to embodiments of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0149] As shown in Figure 8 is a block diagram of an electronic device according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown in the electronic device, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0150] As shown in Figure 8 The electronic device includes one or more processors 801, a memory 802, and an interface that connects the components, including a high-speed interface and a low-speed interface. The components are interconnected using different buses, and can be mounted on a common main board or otherwise installed as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information on a GUI on an external input / output device, such as a display device coupled to the interface. In other embodiments, multiple processors and / or buses can be used with multiple memories and multiple storage devices, if desired. Similarly, multiple electronic devices can be connected, each providing part of the necessary operations (e.g., as a server array, a set of blade servers, or a multi-processor system). Figure 8 The processor 801 is taken as an example in the electronic device.
[0151] The memory 802 is a non-transitory computer-readable storage medium provided by the present disclosure. The memory stores instructions executable by at least one processor, so that the at least one processor executes the large model-based training data generation method or the knowledge question and answer model training method provided by the present disclosure. The non-transitory computer-readable storage medium of the present disclosure stores computer instructions for causing a computer to execute the large model-based training data generation method or the knowledge question and answer model training method provided by the present disclosure.
[0152] The memory 802 is a non-transitory computer-readable storage medium, which can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the large model-based training data generation method or the knowledge question and answer model training method in the embodiments of the present disclosure (for example, the acquisition module 601, the screening module 602, the rewriting module 603, the generation module 604 and the updating module 605 shown in the embodiments of the present disclosure). Figure 6 The processor 801 executes various function applications and data processing of the server by running the non-transitory software programs, instructions and modules stored in the memory 802, that is, implements the large model-based training data generation method or the knowledge question and answer model training method in the above method embodiments.
[0153] The memory 802 can include a program area and a data area, wherein the program area can store an operating system and application programs required by at least one function; the data area can store data created according to the use of the electronic device, etc. In addition, the memory 802 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory 802 can optionally include a memory disposed remotely with respect to the processor 801, which can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0154] The electronic device can further include an input device 803 and an output device 804. The processor 801, the memory 802, the input device 803 and the output device 804 can be connected by a bus or other means, Figure 8 For example, by a bus connection.
[0155] The input device 803 can receive input of numeric or character information, and generate key signal input relating to user settings and function controls of the electronic device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a jog wheel, one or more mouse buttons, a trackball, a joystick, etc. The output device 804 can include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The display device can include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display, and a plasma display. In some embodiments, the display device can be a touch screen.
[0156] Various embodiments of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0157] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0158] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0159] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0160] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0161] It should be understood that various forms of flow shown above can be used with orders of steps reordered, added to, or deleted from. For example, each step recited in the present disclosure can be executed in parallel, in sequence, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.
[0162] The above detailed description does not limit the scope of the disclosure. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed description. Any modification, equivalent replacement and improvement etc. made within the spirit and principle of the disclosure shall be included in the scope of the disclosure.
Claims
1. A method for generating training data based on a large model, comprising: Based on the historical operation data, N triple data are obtained, where N is a positive integer; wherein each triple data includes a question, an answer corresponding to the question, and a search fragment; Filtering out triple data with incorrect answers from the N triple data as triple data to be corrected; The problem in the triple data to be corrected is rewritten based on the iterative reflection of the large model, including: evaluating the retrieval segment in the triple data to be corrected to obtain an evaluation result of the retrieval segment; constructing a second prompt information using a second prompt template based on the evaluation result of the triple data to be corrected and the retrieval segment; the second prompt information is used to instruct the large model to rewrite the problem; rewriting the problem in the triple data to be corrected based on the second prompt information and the large model to obtain a candidate question, and generating a retrieval segment based on the candidate question; generating a first answer based on the candidate question and the generated retrieval segment using the large model; evaluating the generated retrieval segment based on the first answer generated by the large model and the candidate question to obtain an evaluation result of the generated retrieval segment; if the evaluation result of the generated retrieval segment does not meet the preset condition, performing iterative reflection, continuing to evaluate and rewrite the candidate question until the evaluation result of the generated retrieval segment meets the preset condition; generating a revised search snippet based on the rewritten question and the large model; Based on the rewritten question and the revised search fragment, using the large model to generate a corrected answer corresponding to the rewritten question; Based on the rewritten question, the revised search fragment and the corrected answer, the triple data to be revised is updated to obtain retrieval enhancement generation RAG training data.
2. The method according to claim 1, wherein The step of selecting the triple data with incorrect answers from the N triple data as the triple data to be corrected includes: Based on the questions in the N triple data and the search segments corresponding to the questions, construct first prompt information using a first prompt template; Obtaining a reference answer corresponding to the question based on the first prompt information and the large model; Based on the answer corresponding to the question and the reference answer corresponding to the question, the triple data with incorrect answers are screened out from the N triple data as the triple data to be corrected.
3. The method according to claim 1, wherein The evaluation results include accuracy and / or relevance; The accuracy is used to characterize whether the search segment has no answer and / or whether the search segment is complete, and the relevance is used to characterize the relevance between the search segment and the question.
4. The method according to claim 1 or 3, wherein: Generating a first answer using the large model based on the candidate question and the generated search fragment includes: Based on the candidate question and the generated search fragment, using multiple large models to generate multiple candidate answers; The first answer is obtained from the multiple candidate answers.
5. The method according to claim 4, wherein: The obtaining the first answer from the multiple candidate answers includes: Constructing third prompt information using a third prompt template, wherein the third prompt information is used to instruct the macro model to score the multiple candidate answers based on at least one scoring indicator dimension; the at least one scoring indicator dimension includes one or more of the following: answer accuracy, relevance of the answer to the search fragment, completeness of the answer, and readability of the answer format; Based on the third prompt information, using the multiple large models to score the multiple candidate answers on the at least one scoring indicator dimension to obtain scoring information for the multiple candidate answers; The first answer is obtained from the plurality of candidate answers based on the score information of the plurality of candidate answers.
6. The method of claim 1, wherein: Generating a corrected answer corresponding to the rewritten question using the large model based on the rewritten question and the revised search fragment includes: Based on the rewritten question and the revised search fragment, multiple large models are used to generate multiple candidate answers; The corrected answer is obtained from the plurality of candidate answers.
7. The method according to claim 6, wherein: The obtaining the corrected answer from the plurality of candidate answers comprises: Scoring the multiple candidate answers using the multiple large models based on at least one scoring indicator dimension to obtain scoring information for the multiple candidate answers; the at least one scoring indicator dimension includes one or more of the following: answer accuracy, relevance of the answer to the search fragment, completeness of the answer, and readability of the answer format; The corrected answer is obtained from the plurality of candidate answers based on the score information of the plurality of candidate answers.
8. A knowledge question answering model training method, comprising: Acquire training data, where the training data is generated based on the large model-based training data generation method according to any one of claims 1 to 7; The knowledge question answering model is trained based on the training data.
9. A training data generation device based on a large model, comprising: An acquisition module is configured to acquire N triple data based on historical operation data, where N is a positive integer; wherein each triple data includes a question, an answer corresponding to the question, and a search fragment; a screening module, configured to screen out triple data with incorrect answers from the N triple data as triple data to be corrected; a rewriting module for rewriting the question in the triple data to be corrected based on the iterative reflection of the large model, and generating a corrected retrieval segment based on the rewritten question and the large model; wherein the rewriting module is used to: evaluate the retrieval segment in the triple data to be corrected to obtain an evaluation result of the retrieval segment; construct a second prompt information using a second prompt template based on the evaluation result of the triple data to be corrected and the retrieval segment; the second prompt information is used to instruct the large model to rewrite the question; rewrite the question in the triple data to be corrected based on the second prompt information and the large model to obtain a candidate question, and generate a retrieval segment based on the candidate question; generate a first answer based on the candidate question and the generated retrieval segment using the large model; evaluate the generated retrieval segment based on the first answer generated by the large model and the candidate question to obtain an evaluation result of the generated retrieval segment; if the evaluation result of the generated retrieval segment does not meet the preset condition, perform iterative reflection, continue to evaluate and rewrite the candidate question until the evaluation result of the generated retrieval segment meets the preset condition; a generation module, configured to generate a corrected answer corresponding to the rewritten question using the large model based on the rewritten question and the corrected search fragment; An updating module is used to update the triple data to be corrected based on the rewritten question, the corrected search fragment and the corrected answer to obtain retrieval enhancement and generate RAG training data.
10. The device according to claim 9, wherein The screening module is used to: Based on the questions in the N triple data and the search segments corresponding to the questions, construct first prompt information using a first prompt template; Obtaining a reference answer corresponding to the question based on the first prompt information and the large model; Based on the answer corresponding to the question and the reference answer corresponding to the question, the triple data with incorrect answers are screened out from the N triple data as the triple data to be corrected.
11. The device according to claim 9, wherein The evaluation results include accuracy and / or relevance; The accuracy is used to characterize whether the search segment has no answer and / or whether the search segment is complete, and the relevance is used to characterize the relevance between the search segment and the question.
12. The device according to claim 9 or 11, wherein The rewriting module is used to: Based on the candidate question and the generated search fragment, using multiple large models to generate multiple candidate answers; The first answer is obtained from the multiple candidate answers.
13. The device of claim 12, wherein: The rewriting module is used to: Constructing third prompt information using a third prompt template, wherein the third prompt information is used to instruct the macro model to score the multiple candidate answers based on at least one scoring indicator dimension; the at least one scoring indicator dimension includes one or more of the following: answer accuracy, relevance of the answer to the search fragment, completeness of the answer, and readability of the answer format; Based on the third prompt information, using the multiple large models to score the multiple candidate answers on the at least one scoring indicator dimension to obtain scoring information for the multiple candidate answers; The first answer is obtained from the plurality of candidate answers based on the score information of the plurality of candidate answers.
14. The apparatus of claim 9, wherein: The generation module is used to: Based on the rewritten question and the revised search fragment, multiple large models are used to generate multiple candidate answers; The corrected answer is obtained from the plurality of candidate answers.
15. The apparatus of claim 14, wherein: The generation module is used to: Scoring the multiple candidate answers using the multiple large models based on at least one scoring indicator dimension to obtain scoring information for the multiple candidate answers; the at least one scoring indicator dimension includes one or more of the following: answer accuracy, relevance of the answer to the search fragment, completeness of the answer, and readability of the answer format; The corrected answer is obtained from the plurality of candidate answers based on the score information of the plurality of candidate answers.
16. A knowledge question answering model training device, comprising: An acquisition module, configured to acquire training data, wherein the training data is generated based on the large model-based training data generation method according to any one of claims 1 to 7; A training module is used to train the knowledge question answering model based on the training data.
17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7 and 8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7 and 8.
19. A computer program product comprising a computer program, wherein When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 7 and 8.
Citation Information
Patent Citations
Model training method, data processing method, electronic equipment and storage medium
CN118446338A
Process optimization design method and system for improving RAG accuracy of large model
CN118503350A
Question and answer method and device based on large model, electronic equipment, storage medium, intelligent agent and program product
CN118981527A