Large model-based training sample generation method and apparatus, and electronic device
By extracting high-quality reference question-and-answer pairs from manual customer service data and generating training samples with semantic matching algorithms, the problem of poor Q&A in the vertical field of RAG system is solved, and more efficient and accurate training sample generation is achieved, improving the model's Q&A ability.
Patent Information
- Application Number
- CN202510138415.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
When generating training samples, it is difficult for existing RAG systems to effectively utilize high-quality Q&A pairs in vertical fields, resulting in insufficient accuracy and targeting of training samples, which in turn affects the model's Q&A effect in vertical fields.
By extracting high-quality reference question-and-answer pairs from the manual customer service data, and combining semantic matching algorithms, matching knowledge fragments are retrieved from the knowledge base, and training samples for retrieving the target training task of generating large models are generated.
It improves the accuracy and pertinence of training samples, enhances the Q&A ability of retrieving and generating large models in vertical fields, and improves the generalization ability and response accuracy of the model.
Smart Images

Figure CN120069033A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, especially in the field of artificial intelligence such as deep learning and large models. Specifically, it relates to a method, device, and electronic device for generating training samples based on a large model. Background Art
[0002] The RAG (Retrieval-Augmented Generation) system is an artificial intelligence model architecture that combines retrieval and generation technologies, mainly used in natural language processing tasks. The RAG system retrieves relevant information from a large number of documents and then uses this information to generate more accurate and context-appropriate responses, thereby improving the quality of predictions. Summary of the Invention
[0003] This application provides a method, device, and electronic device for generating training samples based on a large model.
[0004] According to one aspect of this application, there is provided a method for generating training samples based on a large model, including:
[0005] Obtain the input query information;
[0006] In response to the query information matching the reference question in the reference Q&A pairs of the target domain, retrieve the first knowledge fragment matching the query information from the knowledge base; wherein, the reference Q&A pairs are extracted from artificial customer service data;
[0007] Generate a training sample for the target training task of the retrieval and generation large model according to the query information, the reference question, and the first knowledge fragment.
[0008] According to another aspect of this application, there is provided a training method for a retrieval and generation large model, including:
[0009] Obtain the training sample for the target training task of the retrieval and generation large model; wherein, the training sample is generated according to the training sample generation method described in the above-mentioned first aspect embodiment;
[0010] Use the training sample to train the retrieval and generation large model for the target training task to obtain the trained retrieval and generation large model.
[0011] According to another aspect of this application, there is provided a method for generating answers based on a large model, including:
[0012] Obtain the input query information;
[0013] Retrieve in the knowledge base according to the query information to obtain a knowledge fragment matching the query information;
[0014] Based on the query information and the knowledge fragments, use a retrieval and generation large model to obtain an answer to the query information; wherein, the retrieval and generation large model is trained according to the training method described in the above-mentioned embodiment of the other aspect.
[0015] According to another aspect of the present application, there is provided a training sample generation device based on a large model, including:
[0016] An acquisition module, configured to acquire input query information;
[0017] A retrieval module, configured to retrieve, from a knowledge base, a first knowledge fragment that matches the query information in response to a match between the query information and a reference question in a reference Q&A pair in a target domain; wherein, the reference Q&A pair is extracted from artificial customer service data;
[0018] A generation module, configured to generate a training sample for a target training task of the retrieval and generation large model according to the query information, the reference question, and the first knowledge fragment.
[0019] According to another aspect of the present application, there is provided a training device for a retrieval and generation large model, including:
[0020] An acquisition module, configured to acquire a training sample for a target training task of the retrieval and generation large model; wherein, the training sample is generated according to the training sample generation method described in the above-mentioned embodiment of one aspect;
[0021] A training module, configured to train the retrieval and generation large model for the target training task by using the training sample, to obtain a trained retrieval and generation large model.
[0022] According to another aspect of the present application, there is provided an answer generation device based on a large model, including:
[0023] An acquisition module, configured to acquire input query information;
[0024] A retrieval module, configured to retrieve in a knowledge base according to the query information to obtain a knowledge fragment that matches the query information;
[0025] A generation module, configured to use the retrieval and generation large model to obtain an answer to the query information according to the query information and the knowledge fragment; wherein, the retrieval and generation large model is trained according to the training method described in the above-mentioned embodiment of the other aspect.
[0026] According to another aspect of the present application, there is provided an electronic device, including:
[0027] At least one processor; and
[0028] A memory communicatively connected to the at least one processor; wherein,
[0029] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the above embodiments.
[0030] According to another aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the above embodiments.
[0031] According to another aspect of the present application, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in the above embodiments are implemented.
[0032] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings are used to better understand the solution and do not constitute a limitation to the present application. Among them:
[0034] Figure 1 It is a schematic flowchart of a method for generating training samples based on a large model provided by an embodiment of the present application;
[0035] Figure 2 It is a schematic flowchart of a method for generating training samples based on a large model provided by another embodiment of the present application;
[0036] Figure 3 It is a schematic flowchart of a method for generating training samples based on a large model provided by another embodiment of the present application;
[0037] Figure 4 It is a schematic diagram of a process for generating training samples based on a large model provided by an embodiment of the present application;
[0038] Figure 5 It is a schematic flowchart of a method for training a retrieval and generation large model provided by another embodiment of the present application;
[0039] Figure 6 It is a schematic flowchart of a method for generating answers based on a large model provided by another embodiment of the present application;
[0040] Figure 7 It is a schematic structural diagram of a device for generating training samples based on a large model provided by an embodiment of the present application;
[0041] Figure 8 Structural schematic diagram of a training device for a retrieval generation large model provided by an embodiment of the present application;
[0042] Figure 9 Structural schematic diagram of an answer generation device based on a large model provided by an embodiment of the present application;
[0043] Figure 10 It is a block diagram of an electronic device for implementing a training sample generation method based on a large model in an embodiment of the present application. Detailed implementation manners
[0044] The following describes exemplary embodiments of the present application with reference to the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0045] It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solution of the present application all comply with the relevant regulations of national laws and regulations and do not violate public order and good customs.
[0046] Next, a training sample generation method, device, electronic device, and storage medium based on a large model in an embodiment of the present application will be described with reference to the accompanying drawings.
[0047] Figure 1 Flow schematic diagram of a training sample generation method based on a large model provided by an embodiment of the present application.
[0048] The training sample generation method based on a large model in an embodiment of the present application can be executed by the training sample generation device based on a large model in an embodiment of the present application, and this device can be configured in an electronic device.
[0049] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc., which are hardware devices with various operating systems, touch screens, and / or display screens.
[0050] As Figure 1 shown, the training sample generation method based on a large model includes:
[0051] Step 101, obtain the input query information.
[0052] In the present application, the user can input query information on the client interface, and the client can send the query information to the RAG system, so that the RAG system can obtain the query information input by the user.
[0053] Step 102: In response to the query information matching the reference question in the reference Q&A pairs of the target domain, retrieve the first knowledge fragment matching the query information from the knowledge base.
[0054] Among them, the reference Q&A pairs of the target domain can be extracted from the artificial customer service data. Among them, the reference Q&A pairs include reference questions and reference answers. The reference question can refer to the user's question, and the reference answer can refer to the answer of the artificial customer service to the reference question. The target domain can refer to any domain, and there is no limitation on this.
[0055] It should be noted that the reference Q&A pairs can be one or more, and there is no limitation on this.
[0056] Exemplarily, the artificial customer service data can include artificial customer service work orders. The artificial customer service work orders can include, but are not limited to, work order numbers, basic user information, description information of user questions, customer service information, solutions to user questions, processing status, etc. Among them, the description information of the user question can include the user's specific question, the urgency of the question, etc.
[0057] Exemplarily, the artificial customer service data can include user questions and customer service answers in the target domain, and can also include user questions and customer service answers in other domains, etc.
[0058] For example, the artificial customer service data can include user questions and customer service answers in multiple domains such as the home appliance domain, the pharmaceutical domain, the clothing domain, etc.
[0059] In this application, the query information can be compared with the reference questions in the reference Q&A pairs. If the query information matches the reference questions in the reference Q&A pairs, the query information can be used to retrieve in the knowledge base to obtain the first knowledge fragment matching the query information.
[0060] Exemplarily, a semantic matching algorithm can be used to match the query information with the knowledge fragments in the knowledge base, and the knowledge fragment with the highest semantic matching degree with the query information can be used as the first knowledge fragment.
[0061] Step 103: Generate a training sample for the target training task of the retrieval and generation large model according to the query information, reference questions, and the first knowledge fragment.
[0062] Among them, the retrieval and generation large model can refer to the generation model in the RAG system.
[0063] Exemplarily, the target training task can include, but is not limited to, model generation training tasks, model reflection generation training tasks, query rewriting training tasks, etc.
[0064] Among them, the model generation training task is used to train the ability of the retrieval and generation large model to generate answers, the model reflection generation training task is used to train the ability of the retrieval and generation large model to generate answers and the ability to reflect on the generated answers, and the query rewriting training task is used to train the query information rewriting ability of the retrieval and generation large model.
[0065] In this application, training samples for the target training task of the retrieval and generation large model can be generated according to the reference answer, the first knowledge fragment, and the reference answer to the reference question in the reference Q&A pair.
[0066] Exemplarily, according to whether the first knowledge fragment is relevant to the reference answer, combined with the query information, the target large model can be used to generate an answer or rewrite the query information to obtain the rewritten query information, so that the training samples for the target training task of the retrieval and generation large model can be determined based on the input and output of the target large model.
[0067] In this application, the training samples for the target training task of the retrieval and generation large model can be used to train the retrieval and generation large model for the target training task to obtain the trained retrieval and generation large model.
[0068] In the embodiments of this application, by obtaining the input query information, if the query information matches the reference question in the reference Q&A pair of the target domain, the first knowledge fragment matching the query information is retrieved in the knowledge base, where the reference Q&A pair is extracted from the artificial customer service data, and training samples for the target training task of the retrieval and generation large model are generated according to the query information, the reference question, and the first knowledge fragment. Thus, by matching the input query information with the reference Q&A pair of the vertical domain extracted from the artificial customer service data and retrieving the relevant knowledge fragment from the knowledge base accordingly to generate the training samples of the retrieval and generation large model, the accuracy and pertinence of the training samples can be improved, so that the accuracy of the answers generated by the retrieval and generation large model can be improved, and further the Q&A effect of the retrieval and generation large model in the vertical domain can be enhanced.
[0069] In addition, training samples for different training tasks can be obtained, and the retrieval and generation large model is trained based on the training samples for different training tasks, so that the abilities of multiple aspects of the retrieval and generation large model can be enhanced, and the generalization ability of the retrieval and generation large model can be improved.
[0070] Figure 2 It is a schematic flowchart of a method for generating training samples based on a large model provided by another embodiment of this application.
[0071] As Figure 2 shown, the method for generating training samples based on a large model includes:
[0072] Step 201, obtain the input query information.
[0073] In this application, step 201 can adopt any implementation manner in the various embodiments of this application, so it will not be elaborated here.
[0074] Step 202: In response to the query information matching the reference question in the reference Q&A pairs of the target field, retrieve the first knowledge fragment matching the query information from the knowledge base.
[0075] Among them, the reference Q&A pairs of the target field can be extracted from the artificial customer service data.
[0076] To improve the quality of the Q&A pairs, exemplarily, the quality of the artificial customer service data can be detected. If it is determined that the artificial customer service data passes the quality detection, it can be considered that the user's question has been completely solved, and then the user's question and the answer of the artificial customer service are extracted from the artificial customer service data to obtain candidate Q&A pairs, and then the reference Q&A pairs of the target field are determined from the candidate Q&A pairs.
[0077] Thus, when it is determined that the artificial customer service data passes the quality detection, and then the Q&A pairs are extracted from the artificial customer service data, the quality of the mined Q&A pairs can be improved.
[0078] Exemplarily, the question and the corresponding answer can be extracted from the artificial customer service data, and the extracted question and answer are matched. If the extracted question and answer match, it can be considered that the question has been solved, and it can be determined that the artificial customer service data passes the quality detection. Thus, based on the matching situation between the question and the answer, it is determined whether the artificial customer service data passes the quality detection, and the accuracy is high.
[0079] Exemplarily, the question processing status can also be obtained from the artificial customer service data. If the question processing status is the target status, it is determined that the artificial customer service data passes the quality detection. Among them, the question processing status can include to be processed, processing, resolved, closed, etc., and the target status can be resolved.
[0080] Thus, based on the question processing status in the artificial customer service data, it is determined whether the artificial customer service data passes the quality detection, and the method is simple and the accuracy is high.
[0081] For determining the reference Q&A pairs of the target field from the candidate Q&A pairs, exemplarily, the candidate Q&A pairs can be quality scored to obtain the quality scores of the candidate Q&A pairs, and the reference Q&A pairs with quality scores greater than the second threshold are screened out from the candidate Q&A pairs. Since the artificial customer service data may contain questions and answers in multiple fields, the reference Q&A pairs with quality scores greater than the second threshold can be classified by field to obtain the reference Q&A pairs of each field, and then the reference Q&A pairs of the target field are obtained from the reference Q&A pairs of each field.
[0082] Thus, by performing quality scoring on candidate question-and-answer pairs and filtering reference question-and-answer pairs based on the quality score, the quality of the reference question-and-answer pairs can be improved. Additionally, classifying the reference question-and-answer pairs with quality scores greater than the second threshold into domains can facilitate subsequent targeted optimization in vertical domains.
[0083] Step 203, in response to the first knowledge fragment being related to the reference answer of the reference question, according to the query information and the first knowledge fragment, use the target large model to generate the target answer to the query information.
[0084] Among them, the target large model and the retrieval and generation large model can be the same model or different models, and there is no limitation on this.
[0085] In this application, the first knowledge fragment can be compared with the reference answer of the reference question in the reference question-and-answer pair. If the first knowledge fragment is related to the reference answer, the target large model can be used to process the query information and the first knowledge fragment to generate the target answer to the query information.
[0086] Step 204, according to the reference answer, obtain the first attribute information of the target answer.
[0087] Among them, the first attribute information can be used to characterize the quality of the target answer. Exemplarily, the first attribute information can be an evaluation score, or other parameters used to characterize the quality of the target answer, and there is no limitation on this.
[0088] Exemplarily, the target answer can be evaluated according to the reference answer to obtain the evaluation score of the target answer. For example, a preset scoring mechanism can be used to evaluate the target answer from multiple aspects such as accuracy, fluency, and coverage according to the reference answer.
[0089] Step 205, according to the query information, the first knowledge fragment, the target answer, and the first attribute information, generate training samples for the model generation training task.
[0090] In this application, the target training task can include a model generation training task. Training samples for the model generation training task can be generated according to the query information, the first knowledge fragment, the target answer, and the first attribute information of the target answer.
[0091] Exemplarily, if the model generation training task uses the first training method to train the retrieval and generation large model, then the training samples for the model generation training task can include the training samples of the first training method. If the value of the first attribute information is greater than the first threshold, it can be considered that the quality of the target answer is qualified, and training samples corresponding to the first training method of the model generation training task can be generated according to the query information, the first knowledge fragment, and the target answer. Among them, the training samples corresponding to the first training method can include the query information, the target answer, etc.
[0092] For example, the first training method can be fine-tuning supervision. If the value of the first attribute information is greater than the first threshold, the query information, the first knowledge fragment, and the target answer can be used as the training samples for fine-tuning supervision.
[0093] Exemplarily, if the value of the first attribute information is less than or equal to the first threshold, it can be considered that the quality of the target answer is unqualified. The answer can be regenerated using the target large model based on the query information, the first knowledge fragment, and the target answer. Then, according to the reference answer, the second attribute information of the regenerated answer is obtained. If the value of the second attribute information is greater than the first threshold, it indicates that the quality of the regenerated answer is qualified. The training samples corresponding to the first training method can be generated based on the query information, the first knowledge fragment, and the regenerated answer. Among them, the training samples corresponding to the first training method can include the query information, the regenerated answer, etc.
[0094] Thus, according to the magnitude relationship between the value of the first attribute information and the first threshold, it can be determined whether it is necessary to regenerate the answer, and based on this, the training samples of the first training method for the model generation training task can be obtained, thereby improving the quality of the training samples of the first training method, and further improving the answer accuracy of the retrieval and generation large model.
[0095] Optionally, if the first knowledge fragment is not relevant to the reference answer, the query information can be rewritten using the target large model. If the second knowledge fragment retrieved based on the rewritten query information is relevant to the reference answer, then according to the query information and the second knowledge fragment, the answer to the query information can be generated using the target large model. The query information, the second knowledge fragment, and the answer generated based on the query information and the second knowledge fragment can be used as the training samples of the first training method for the model generation training task.
[0096] Exemplarily, the model generation training task can also use the second training method to train the retrieval and generation large model. Then, the training samples of the model generation training task can also include the training samples of the second training method. If the value of the second attribute information of the above regenerated answer is greater than the first threshold, it can be considered that the quality of the regenerated answer is qualified. The positive samples corresponding to the second training method of the model generation training task can be determined based on the query information, the first knowledge fragment, and the regenerated answer, and the negative samples corresponding to the second training method can be determined based on the query information, the first knowledge fragment, and the target answer. Then, according to the positive samples and negative samples corresponding to the second training method, the training samples corresponding to the second training method are determined.
[0097] Among them, the training samples corresponding to the second training method can include positive samples and negative samples.
[0098] For example, the second training method is reward optimization. The query information, the first knowledge fragment, and the regenerated answer can be used as positive samples, and the query information, the first knowledge fragment, and the target answer as negative samples. Based on the positive and negative samples, training samples for reward optimization, i.e., the preference dataset, can be obtained.
[0099] It can be understood that a positive sample and a corresponding negative sample can be regarded as a sample pair of the second training method. The sample pair can be used to train the retrieval generation large model using the second training method.
[0100] Thus, if the value of the second attribute information of the regenerated answer is greater than the first threshold, training samples for the second training method of the training task generated by the model can also be generated based on the query information, the first knowledge fragment, the target answer, and the regenerated answer, thereby improving the quality of the training samples of the second training method, and further enhancing the performance and generalization ability of the retrieval generation large model.
[0101] Regarding the regeneration of the answer based on the query information, the first knowledge fragment, and the target answer, exemplarily, based on the query information and the first knowledge fragment, the target large model can be used to reflect on the target answer, analyze the deficiencies in the target answer, obtain answer improvement suggestions, and then, based on the query information, the first knowledge fragment, and the target answer, use the answer improvement suggestions to guide the target large model to regenerate the answer.
[0102] Thus, by using the target large model to reflect on the answer generated by itself and guiding the target large model to regenerate the answer based on the reflection result, the quality of the regenerated answer can be improved.
[0103] Optionally, the target training task can also include a model reflection training task. If the value of the second attribute information of the regenerated answer is greater than the first threshold, i.e., the quality of the regenerated answer is qualified, training samples for the model reflection training task can be generated based on the query information, the first knowledge fragment, the reply improvement suggestions, the target answer, and the regenerated answer.
[0104] Among them, the training samples for the model reflection training task can include query information, reply improvement suggestions, target answers, regenerated answers, etc.
[0105] Thus, if the quality of the regenerated answer is qualified, training samples for the model reflection training task can also be obtained based on the query information, the first knowledge fragment, the reply improvement suggestions, the target answer, and the regenerated answer. Therefore, by using the training samples for the model reflection training task to train the retrieval generation large model, the quality of the answers generated by the model can be improved.
[0106] In the embodiments of the present application, when the first knowledge fragment is relevant to the reference answer of the reference question, according to the query information and the first knowledge fragment, the target large model is used to generate the target answer of the query information, which can improve the accuracy of the generated answer. And according to the query information, the first knowledge fragment and the target answer, combined with the first attribute information of the target answer obtained based on the reference answer, the training sample of the model generation training task is obtained, so that the accuracy of the training sample of the model generation training task can be improved, and further the accuracy of the answer generated by the retrieval generation large model can be improved.
[0107] Figure 3 It is a schematic flowchart of a method for generating a training sample based on a large model provided by another embodiment of the present application.
[0108] As Figure 3 shown, the method for generating a training sample based on a large model includes:
[0109] Step 301, obtain the input query information.
[0110] Step 302, in response to the query information matching the reference question in the reference Q&A pair of the target domain, retrieve the first knowledge fragment matching the query information from the knowledge base.
[0111] In the present application, steps 301-302 can adopt any implementation manner in the embodiments of the present application, so details are not described herein again.
[0112] Step 303, in response to the first knowledge fragment being irrelevant to the reference answer of the reference question, use the target large model to rewrite the query information to obtain the rewritten query information.
[0113] In the present application, if the first knowledge fragment is irrelevant to the reference answer, the query information can be input into the target large model, and the target large model is used to rewrite the query information to obtain the rewritten query information.
[0114] Optionally, according to the first knowledge fragment, the reference answer, etc., the target large model can be used to reflect on the query information to obtain a query rewriting suggestion, and then the query rewriting suggestion is used to guide the target large model to rewrite the query information to obtain the rewritten query information.
[0115] Optionally, the original question and the essential question of the original question can be obtained, and the target large model uses the original question and its essential question as a reference to rewrite the query information to obtain the rewritten query information. Among them, the essential question can be understood as the essential problem of the original question. Thus, rewriting the query information based on the original question and its essential question as a reference can improve the accuracy of the rewritten query information.
[0116] For example, the original question is "The mobile phone suddenly cannot connect to the WiFi. I followed the prompts to connect, but it has been unable to log in", and the essential problem is "The mobile phone WiFi connection method".
[0117] Another example, the query information is "What other movies has the director of Movie A directed?". The knowledge fragments retrieved from the knowledge base based on this query information are not relevant to the reference answer or no knowledge fragments are retrieved. Use the target large model to rewrite the query information. The rewritten query information can be "Who is the director of Movie A? What other movies has this director directed?".
[0118] Step 304, according to the rewritten query information, retrieve in the knowledge base to obtain a second knowledge fragment that matches the rewritten query information.
[0119] In this application, using the semantic matching algorithm, match the rewritten query information with the knowledge fragments in the knowledge base, and the knowledge fragment with the highest semantic matching degree with the rewritten query information can be used as the second knowledge fragment.
[0120] Step 305, in response to the second knowledge fragment being relevant to the reference answer, generate a training sample for the query rewriting training task according to the query information and the rewritten query information.
[0121] In this application, the second knowledge fragment can be compared with the reference answer. If the second knowledge fragment is relevant to the reference answer, the query information and the rewritten query information can be used as the training sample for the query rewriting training task.
[0122] Optionally, the target training task can also include a query reflection generation training task. The query information, the rewritten query information, and the query rewriting suggestion can be used as the training sample for the query reflection generation training task. Among them, the query reflection generation training task can be used to train the query information rewriting ability and the query reflection ability of the retrieval generation large model.
[0123] In the embodiment of this application, if the first knowledge fragment is not relevant to the reference answer, use the target large model to rewrite the query information to obtain the rewritten query information. If the second knowledge fragment that matches the rewritten query information is relevant to the reference answer, generate a training sample for the query rewriting training task according to the query information and the rewritten query information, which can improve the quality of the training sample for the query rewriting training task, thereby improving the query information rewriting ability of the retrieval generation large model and enhancing the generalization ability of the retrieval generation large model.
[0124] To facilitate understanding of the training sample generation method based on a large model in this application, the following is combined with Figure 4 for illustration. Figure 4Schematic diagram of a process for generating training samples based on a large model provided by an embodiment of the present application.
[0125] As Figure 4 shown, this process includes:
[0126] Step 401, input query information.
[0127] Step 402, the RAG system obtains the query information.
[0128] Step 403, whether to transfer to manual processing. If the user determines to transfer to manual processing, then proceed to Step 404, where the customer service processes the query information.
[0129] Step 404, customer service processing.
[0130] Step 405, analyze the work order.
[0131] In the present application, user problems and customer service answers can be mined from the work order, and the quality of the work order can be detected.
[0132] Step 406, whether to pass the quality inspection.
[0133] In the present application, the work order can be quality-tested based on the mined questions and answers. If the work order passes the quality inspection, then proceed to Step 407, otherwise discard the work order.
[0134] Step 407, retrieve knowledge fragments.
[0135] In the present application, the query information can be matched with the questions in the mined question-and-answer pairs. If the query information matches the questions in the question-and-answer pairs, then retrieve in the knowledge base according to the query information to obtain knowledge fragments that match the query information. If the query information does not match the questions in the question-and-answer pairs, then discard the query information.
[0136] Step 408, whether the knowledge fragments meet the requirements. If so, execute Step 411, otherwise execute Step 409.
[0137] In the present application, if the knowledge fragments are relevant to the answers in the question-and-answer pairs, then the retrieved knowledge fragments meet the requirements, otherwise they do not meet the requirements. Here, the answer refers to the answer to the question that matches the query information.
[0138] Step 409, reflect on the query information.
[0139] Step 410, rewrite the query information.
[0140] In this application, the target large model can be used to reflect on the query information to obtain query score rewriting suggestions. Then, according to the query score rewriting suggestions, the target large model is used to rewrite the query information to obtain the rewritten query information. The rewritten query information is used to continue the retrieval. If the retrieved knowledge fragments meet the requirements, step 411 is executed; otherwise, the reflection and rewriting continue. If the number of rewriting times reaches the upper limit and the knowledge fragments retrieved based on the rewritten query information do not meet the requirements, the rewritten query information is discarded.
[0141] Step 411, model answer.
[0142] In this application, according to the query information or the rewritten query information, and the retrieved knowledge fragments, the target large model can be used to answer to obtain the answer to the query information.
[0143] Step 412, whether the answer meets the requirements. If not, step 413 is executed.
[0144] In this application, the generated answer can be evaluated. If the evaluation score is greater than the preset threshold, the answer can be considered to meet the requirements; otherwise, the target large model can be used to reflect on the answer.
[0145] Exemplarily, the answer in the question-answer pair matching the query information can be used to verify the accuracy of the answer. For example, the answer can be evaluated through a scoring mechanism.
[0146] Step 413, reflect on the answer.
[0147] In this application, the target large model can be used to reflect on the answer to obtain answer rewriting suggestions. Then, according to the answer rewriting suggestions, the target large model is used to regenerate the answer, and then it is judged whether the answer meets the requirements. If not, continue to reflect and regenerate the answer. If the upper limit of regenerating the answer is reached and the regenerated answer does not meet the requirements, the answer is discarded.
[0148] In this embodiment, according to the input and output of the target large model involved in steps 407 - 413, the training samples for different training tasks of the retrieval and generation large model can be determined. For details, refer to the above embodiments and will not be elaborated here.
[0149] For example, according to the evaluation score of the answer, the answers with an evaluation score greater than the preset score can be screened out to be used for constructing the training samples for supervised fine-tuning to perform supervised fine-tuning on the generation model in the RAG system, and the answers with an evaluation score less than the preset score and the answers with an evaluation score greater than the preset score are used to construct the training samples for reward optimization to perform reward optimization on the generation model in the RAG system.
[0150] In addition, the generative model in the RAG system can be used to analyze the deficiencies in its own answers, improve the generation logic, provide optimization suggestions for unqualified answers, and update the training data according to the latest evaluation results to improve the generation quality of the model. After optimizing the generative model in the RAG system each time, the model performance is re-evaluated to ensure the effectiveness of the improvement. That is to say, a feedback mechanism is introduced in the model training closed-loop to continuously optimize the model performance.
[0151] The training sample generation method based on the large model in the embodiments of the present application has the following beneficial effects:
[0152] (1) Efficient data utilization: By extracting high-quality question-and-answer pairs from the artificial customer service data, fully exploiting the value of existing data, and combining quality assessment to ensure data quality.
[0153] (2) Vertical domain optimization: Using the question-and-answer pairs and query information rewriting mechanism screened by quality inspection, deeply customize for vertical domains such as medical, legal, financial, etc., to improve the accuracy and coverage of the RAG system.
[0154] (3) Dynamic feedback mechanism: Introduce reflection and query information rewriting to optimize the retrieval and generation capabilities in a closed loop, and gradually improve the response accuracy and robustness of the model.
[0155] (4) Model training enhancement: High-quality training data can be generated through supervised fine-tuning and reward optimization, and the model output can be adjusted in combination with user preferences to improve the answer quality.
[0156] (5) Cyclic optimization: Continuously improve the system performance through data update, model evaluation and improvement to adapt to the requirements of dynamic scenarios.
[0157] To implement the above embodiments, the embodiments of the present application also propose a training method for a retrieval and generation large model. Figure 5 It is a schematic flowchart of the training method for the retrieval and generation large model provided by another embodiment of the present application.
[0158] As Figure 5 shown, the training method for the retrieval and generation large model includes:
[0159] Step 501, obtain training samples for the target training task of the retrieval and generation large model.
[0160] In the present application, the training sample generation method described in the above embodiments can be used to obtain training samples for the target training task of the retrieval and generation large model.
[0161] Step 502, use the training samples to train the retrieval and generation large model for the target training task to obtain the trained retrieval and generation large model.
[0162] In this application, after obtaining the training samples of the target training task, the retrieval and generation large model can be trained for the target training task by using one or more training methods according to the training samples of the target training task, and the trained retrieval and generation large model can be obtained.
[0163] Exemplarily, the target training task may include, but is not limited to, a model generation training task, a model reflection generation training task, a query rewriting training task, etc. The specific explanations of these training tasks can be referred to the above embodiments, so they will not be elaborated here.
[0164] Exemplarily, the training method may include, but is not limited to, supervised fine-tuning, reward optimization, etc.
[0165] In the embodiment of this application, by using the training samples of the target training task obtained by the above method for generating training samples, the retrieval and generation large model is trained for the target training task, so as to improve the accuracy of the retrieval and generation large model in generating answers, and further improve the question-answering effect of the retrieval and generation large model in the vertical field.
[0166] In an embodiment of this application, the target training task may include a model generation training task, and the training samples of this training task may include the training samples corresponding to the first training method.
[0167] Exemplarily, the training samples corresponding to the first training method include query information, the first knowledge fragment, and the target answer. The query information and the first knowledge fragment can be input into the retrieval and generation large model to obtain the predicted answer output by the retrieval and generation large model. According to the predicted answer and the target answer, the model loss is determined. According to the model loss, the parameters of the retrieval and generation large model are adjusted, and then the retrieval and generation large model with adjusted parameters is continuously trained until the training end condition is met, and the trained retrieval and generation large model is obtained.
[0168] Exemplarily, the training samples corresponding to the first training method include query information, the first knowledge fragment, and the regenerated answer. For example, the first training method may be supervised fine-tuning. The query information and the first knowledge fragment can be input into the retrieval and generation large model to obtain the predicted answer output by the retrieval and generation large model. According to the predicted answer and the regenerated answer, the model loss is determined. According to the model loss, the parameters of the retrieval and generation large model are adjusted, and then the retrieval and generation large model with adjusted parameters is continuously trained until the training end condition is met, and the trained retrieval and generation large model is obtained.
[0169] Thus, by training the retrieval and generation large model through the first training method, the accuracy of the retrieval and generation large model in generating answers can be improved.
[0170] Optionally, the training samples for the model generation training task may also include the training samples corresponding to the second training method. The training samples corresponding to the second training method include positive samples and negative samples. The positive samples include query information, the first knowledge fragment, and the regenerated answer. The negative samples include query information, the first knowledge fragment, and the target answer.
[0171] For example, the second training method can be reward optimization. The positive samples and negative samples can be input into the retrieval and generation large model, and it is stated that the regenerated answer is adopted and the target answer is rejected to obtain the predicted answer output by the retrieval and generation large model. According to the predicted answer, the regenerated answer, and the target answer, the model loss is calculated. According to the model loss, the parameters of the retrieval and generation large model are adjusted, and the retrieval and generation large model after parameter adjustment is continuously trained until the training end condition is met.
[0172] Thus, training the retrieval and generation large model through the second training method can improve the accuracy of the answer generated by the retrieval and generation large model.
[0173] It should be noted that the retrieval and generation large model can be trained first using the first training method, and then using the second training method. It can also be trained using only the first training method or the second training method alone, and there is no limitation in this regard.
[0174] In an embodiment of the present application, the target training task may include a model reflection training task.
[0175] Exemplarily, the training samples corresponding to the model reflection training task may include query information, the first knowledge fragment, reply improvement suggestions, the target answer, and the regenerated answer. The query information and the first knowledge fragment can be input into the retrieval and generation large model, and it is stated that the generated answer is evaluated. If the quality is unqualified, reflection is performed, and the answer before reflection, the suggestions obtained from reflection, and the answer after reflection are output. Then, according to the answer before reflection and the target answer, the suggestions obtained from reflection and the answer improvement suggestions, and the answer after reflection and the regenerated answer, the model loss is obtained. Based on the model loss, the parameters of the retrieval and generation large model are adjusted, and the retrieval and generation large model after parameter adjustment is continuously trained until the training end condition is met, until the training result condition is met.
[0176] It can be understood that if the training samples corresponding to the model reflection training task include the rewritten query information, the knowledge fragment related to the rewritten query information, the answer generated based on the rewritten query information and the related knowledge fragment, and the regenerated answer, a similar method can be used to train the retrieval and generation large model for the model reflection training task, and this step will not be elaborated here.
[0177] Thus, while training the generation ability of the retrieval generation large model, the reflection ability of the model can be trained, thereby improving the accuracy of the answers generated by the retrieval generation large model.
[0178] In one embodiment of the present application, the target training task may include a query rewriting training task. The training samples corresponding to the query rewriting training task may include query information and rewritten query information. The query information can be input into the retrieval generation large model for rewriting, and the rewritten query information output by the retrieval generation large model can be obtained. Based on the rewritten query information and the rewritten query information in the training samples, the model loss is determined. The parameters of the retrieval generation large model are adjusted based on the model loss, and the retrieval generation large model with adjusted parameters is continuously trained until the training end condition is met, obtaining the trained retrieval generation large model.
[0179] Thus, the training samples corresponding to the query rewriting training task can be used to train the query information rewriting ability of the retrieval generation large model, thereby improving the generalization ability of the retrieval generation large model.
[0180] The solution of the present application can be applied to the customer service system of large enterprises to improve customer service efficiency and reduce manual operations. It is applicable to various online customer service tools and products, can help Internet platforms or e-commerce websites optimize user support processes, can be applied to public affairs and brand public opinion management systems, and can improve crisis handling efficiency by analyzing problem emotions and quickly retrieving corresponding solutions. It can also be used in vertical fields such as financial information, patent retrieval, and academic literature to achieve efficient and accurate search in combination with RAG, etc.
[0181] To implement the above embodiments, the embodiments of the present application also propose a large model-based answer generation method. Figure 6 It is a schematic flowchart of the large model-based answer generation method provided by another embodiment of the present application.
[0182] As Figure 6 shown, the large model-based answer generation method includes:
[0183] Step 601, obtain the input query information.
[0184] In the present application, step 601 can adopt the method of obtaining query information described in the above embodiments, which will not be elaborated here.
[0185] Step 602, retrieve in the knowledge base according to the query information to obtain knowledge fragments matching the query information.
[0186] In the present application, step 602 can adopt the method of obtaining knowledge fragments matching the query information described in the above embodiments, which will not be elaborated here.
[0187] Step 603: According to the query information and knowledge fragments, use the retrieval and generation large model to obtain the answer to the query information.
[0188] Among them, the retrieval and generation large model can be trained by using the training method described in the above embodiments.
[0189] In this application, the query information and knowledge fragments can be input into the retrieval and generation large model, and the retrieval and generation large model is used to generate the answer to the query information. Thus, the generated answer can be displayed on the client to provide the answer to the user.
[0190] In the embodiments of this application, by retrieving relevant knowledge fragments in the knowledge base according to the query information, and using the retrieval and generation large model trained based on the above training method to generate the answer to the query information according to the query information and the retrieved knowledge fragments, the accuracy of the answer can be improved.
[0191] To implement the above embodiments, the embodiments of this application also propose a training sample generation device based on a large model. Figure 7 It is a schematic structural diagram of a training sample generation device based on a large model provided in an embodiment of this application.
[0192] As Figure 7 shown, the training sample generation device 700 based on a large model includes:
[0193] An acquisition module 710, configured to acquire the input query information;
[0194] A retrieval module 720, configured to retrieve the first knowledge fragment matching the query information from the knowledge base in response to the query information matching the reference question in the reference Q&A pairs of the target domain; wherein, the reference Q&A pairs are extracted from the artificial customer service data;
[0195] A generation module 730, configured to generate a training sample for the target training task of the retrieval and generation large model according to the query information, the reference question, and the first knowledge fragment.
[0196] Optionally, the target training task includes a model generation training task, and the generation module 730 is configured to:
[0197] In response to the first knowledge fragment being related to the reference answer of the reference question, generate a target answer to the query information according to the query information and the first knowledge fragment by using the target large model;
[0198] Obtain the first attribute information of the target answer according to the reference answer;
[0199] Generate training samples for the training task of the model generation according to the query information, the first knowledge fragment, the target answer, and the first attribute information.
[0200] Optionally, the generating module 730 is configured to:
[0201] In response to the value of the first attribute information being greater than a first threshold, generate training samples corresponding to the first training method of the training task of the model generation according to the query information, the first knowledge fragment, and the target answer;
[0202] In response to the value of the first attribute information being less than or equal to the first threshold, use the target large model to regenerate an answer according to the query information, the first knowledge fragment, and the target answer;
[0203] Obtain second attribute information of the regenerated answer according to the reference answer;
[0204] In response to the value of the second attribute information being greater than the first threshold, generate training samples corresponding to the first training method according to the query information, the first knowledge fragment, and the regenerated answer.
[0205] Optionally, the generating module 730 is further configured to:
[0206] In response to the value of the second attribute information being greater than the first threshold, determine positive samples corresponding to the second training method of the training task of the model generation according to the query information, the first knowledge fragment, and the regenerated answer;
[0207] Determine negative samples corresponding to the second training method according to the query information, the first knowledge fragment, and the target answer;
[0208] Determine training samples corresponding to the second training method according to the positive samples and negative samples corresponding to the second training method.
[0209] Optionally, the generating module 730 is configured to:
[0210] Reflect on the target answer using the target large model according to the query information and the first knowledge fragment to obtain answer improvement suggestions;
[0211] Regenerate an answer using the target large model according to the query information, the first knowledge fragment, the target answer, and the answer improvement suggestions.
[0212] Optionally, the target training task further includes a model reflection training task, and the generating module 730 is further configured to:
[0213] In response to the value of the second attribute information being greater than the first threshold, a training sample for the model reflection training task is generated according to the query information, the first knowledge fragment, the reply improvement suggestion, the target answer, and the regenerated answer.
[0214] Optionally, the target training task includes a query rewriting training task, and the generating module 730 is configured to:
[0215] In response to the first knowledge fragment being irrelevant to the reference answer of the reference question, use the target large model to rewrite the query information to obtain the rewritten query information;
[0216] Retrieve in the knowledge base according to the rewritten query information to obtain a second knowledge fragment that matches the rewritten query information;
[0217] In response to the second knowledge fragment being relevant to the reference answer, generate a training sample for the query rewriting training task according to the query information and the rewritten query information.
[0218] Optionally, the generating module 730 is configured to:
[0219] Obtain the original question and the essential question corresponding to the original question;
[0220] Rewrite the query information using the target large model according to the original question and the essential question to obtain the rewritten query information.
[0221] Optionally, the apparatus may further include:
[0222] An extraction module, configured to extract candidate question-and-answer pairs from the artificial customer service data in response to determining that the artificial customer service data passes the quality inspection;
[0223] A determination module, configured to determine a reference question-and-answer pair in the target domain from the candidate question-and-answer pairs.
[0224] Optionally, the apparatus further includes:
[0225] An extraction module, configured to extract candidate question-and-answer pairs from the artificial customer service data in response to determining that the artificial customer service data passes the quality inspection;
[0226] A determination module, configured to determine a reference question-and-answer pair in the target domain from the candidate question-and-answer pairs.
[0227] Optionally, the extraction module is configured to:
[0228] Extract questions and the corresponding answers to the questions from the artificial customer service data;
[0229] In response to the question matching the answer, it is determined that the artificial customer service data passes the quality inspection.
[0230] Optionally, the extraction module is configured to:
[0231] Obtain the problem handling status from the artificial customer service data;
[0232] In response to the problem handling status being the target status, it is determined that the artificial customer service data passes the quality inspection.
[0233] It should be noted that the explanations of the foregoing embodiments of the training sample generation method based on the large model also apply to the training sample generation device based on the large model in this embodiment, so they will not be elaborated here.
[0234] In the embodiments of the present application, by matching the input query information with the reference Q&A pairs in the vertical domain extracted from the artificial customer service data, and retrieving relevant knowledge fragments from the knowledge base accordingly to generate training samples for the retrieval and generation large model, the accuracy and pertinence of the training samples can be improved, thereby the accuracy of the answers generated by the retrieval and generation large model can be improved, and further the Q&A effect of the retrieval and generation large model in the vertical domain can be enhanced.
[0235] To implement the above embodiments, the embodiments of the present application also propose a training device for the retrieval and generation large model. Figure 8 It is a schematic structural diagram of a training device for the retrieval and generation large model provided in an embodiment of the present application.
[0236] As Figure 8 shown, the training device 800 for the retrieval and generation large model includes:
[0237] An acquisition module 810, configured to acquire training samples for the target training task of the retrieval and generation large model; wherein, the training samples are generated according to the training sample generation method described in the above embodiments;
[0238] A training module 820, configured to train the retrieval and generation large model for the target training task by using the training samples, to obtain a trained retrieval and generation large model.
[0239] It should be noted that the explanations of the foregoing embodiments of the training method for the retrieval and generation large model also apply to the training device for the retrieval and generation large model in this embodiment, so they will not be elaborated here.
[0240] In the embodiments of the present application, by training the retrieval and generation large model for the target training task by using the training samples obtained based on the above method for generating training samples, the accuracy of the answers generated by the retrieval and generation large model is improved, and further the Q&A effect of the retrieval and generation large model in the vertical domain is enhanced.
[0241] To implement the above embodiments, an embodiment of the present application also proposes an answer generation device based on a large model. Figure 9 FIG. is a schematic structural diagram of an answer generation device based on a large model provided by an embodiment of the present application.
[0242] As Figure 9 shown, the answer generation device 900 based on a large model includes:
[0243] An acquisition module 910, configured to acquire input query information;
[0244] A retrieval module 920, configured to retrieve in a knowledge base according to the query information to obtain a knowledge fragment matching the query information;
[0245] A generation module 930, configured to use a retrieval and generation large model to obtain an answer to the query information according to the query information and the knowledge fragment; wherein, the retrieval and generation large model is trained according to the training method described in the above embodiments.
[0246] It should be noted that the explanations of the embodiments of the above-mentioned answer generation method based on a large model also apply to the answer generation device based on a large model in this embodiment, so details are not described herein again.
[0247] In the embodiments of the present application, by retrieving relevant knowledge fragments in a knowledge base according to query information, and using a retrieval and generation large model trained based on the above training method to generate an answer to the query information according to the query information and the retrieved knowledge fragment, the accuracy of the answer can be improved.
[0248] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium, and a computer program product.
[0249] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present application described herein and / or claimed.
[0250] As Figure 10As shown, device 1000 includes a computing unit 1001, which can execute various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 1002 or a computer program loaded from a storage unit 1008 into a RAM (Random Access Memory) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An I / O (Input / Output) interface 1005 is also connected to the bus 1004.
[0251] Multiple components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disc, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0252] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the training sample generation method based on a large model. For example, in some embodiments, the training sample generation method based on a large model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the training sample generation method based on a large model described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the training sample generation method based on a large model in any other suitable manner (e.g., by means of firmware).
[0253] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0254] The program code for implementing the methods of this application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0255] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0256] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0257] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.
[0258] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (Virtual Private Server). The server may also be a server of a distributed system or a server combined with a blockchain.
[0259] It should be noted that the electronic devices for implementing the training method of the retrieval generation large model in the embodiments of the present application and the electronic devices for implementing the answer generation method based on the large model in the embodiments of the present application are similar in structure to the above-mentioned electronics, so they will not be elaborated here.
[0260] According to an embodiment of the present application, the present application also provides a computer program product. When the instruction processor in the computer program product executes, it executes the training sample generation method based on the large model, or the training method of the retrieval generation large model, or the answer generation method based on the large model proposed in the above embodiments of the present application.
[0261] It should be understood that various forms of the processes shown above can be used, reordering, adding or deleting steps. For example, the steps described in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, and no limitation is made herein.
[0262] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A method for generating training samples based on a large model, comprising: Get the input query information; In response to the query information matching a reference question in a reference question-answer pair in a target domain, retrieving a first knowledge segment matching the query information from a knowledge base; wherein the reference question-answer pair is extracted from manual customer service data; A training sample for a target training task of retrieving and generating a large model is generated according to the query information, the reference question and the first knowledge fragment.
2. The method of claim 1, wherein: The target training task includes a model generation training task, and the generating of training samples for the target training task of retrieving and generating a large model according to the query information, the reference question and the first knowledge fragment includes: In response to the first knowledge fragment being related to a reference answer to the reference question, generating a target answer to the query information using a target macro model according to the query information and the first knowledge fragment; According to the reference answer, obtaining first attribute information of the target answer; A training sample for the model generation training task is generated according to the query information, the first knowledge fragment, the target answer and the first attribute information.
3. The method of claim 2, wherein: The step of generating a training sample for the model generation training task according to the query information, the first knowledge fragment, the target answer and the first attribute information includes: In response to the value of the first attribute information being greater than a first threshold, generating a training sample corresponding to a first training method of the model generation training task according to the query information, the first knowledge fragment and the target answer; In response to the value of the first attribute information being less than or equal to the first threshold, regenerating an answer using the target macro model according to the query information, the first knowledge fragment, and the target answer; According to the reference answer, obtaining second attribute information of the regenerated answer; In response to the value of the second attribute information being greater than the first threshold, a training sample corresponding to the first training method is generated according to the query information, the first knowledge fragment and the regenerated answer.
4. The method of claim 3, further comprising: In response to the value of the second attribute information being greater than the first threshold, determining, according to the query information, the first knowledge fragment, and the regenerated answer, a positive sample corresponding to the second training method of the model generation training task; Determining negative samples corresponding to the second training method according to the query information, the first knowledge fragment and the target answer; According to the positive samples and negative samples corresponding to the second training method, training samples corresponding to the second training method are determined.
5. The method of claim 3, wherein: The step of regenerating the answer using the target macro model according to the query information, the first knowledge fragment and the target answer includes: Based on the query information and the first knowledge fragment, using the target big model to reflect on the target answer to obtain answer improvement suggestions; The answer is regenerated using the target big model according to the query information, the first knowledge fragment, the target answer and the answer improvement suggestion.
6. The method of claim 5, wherein: The target training task also includes a model reflection training task, and the method further includes: In response to the value of the second attribute information being greater than the first threshold, a training sample for the model reflection training task is generated based on the query information, the first knowledge fragment, the answer improvement suggestion, the target answer and the regenerated answer.
7. The method of claim 1, wherein: The target training task includes a query rewriting training task, and the generating of training samples for the target training task of retrieving and generating a large model according to the query information, the reference question and the first knowledge fragment includes: In response to the first knowledge fragment being irrelevant to the reference answer to the reference question, rewriting the query information using the target macro model to obtain rewritten query information; According to the rewritten query information, searching the knowledge base to obtain a second knowledge fragment matching the rewritten query information; In response to the second knowledge fragment being related to the reference answer, a training sample for the query rewriting training task is generated according to the query information and the rewritten query information.
8. The method of claim 7, wherein: The step of rewriting the query information by using the target macro model to obtain the rewritten query information includes: Obtaining the original problem and the essential problem corresponding to the original problem; According to the original question and the essential question, the query information is rewritten using the target big model to obtain the rewritten query information.
9. The method according to any one of claims 1 to 8, further comprising: In response to determining that the manual customer service data passes the quality inspection, extracting candidate question-answer pairs from the manual customer service data; A reference question-answer pair in the target domain is determined from the candidate question-answer pairs.
10. The method of claim 9, wherein: The step of determining a reference question-answer pair in the target domain from the candidate question-answer pairs includes: Performing a quality score on the candidate question-answer pair to obtain a quality score of the candidate question-answer pair; Filtering out reference question-answer pairs having a quality score greater than a second threshold from the candidate question-answer pairs; Performing field classification on the reference question-answer pairs whose quality scores are greater than a second threshold to obtain reference question-answer pairs in various fields; A reference question-answer pair in the target field is obtained from the reference question-answer pairs in each field.
11. The method of claim 9, wherein: Determining that the manual customer service data passes the quality inspection includes: Extracting questions and answers corresponding to the questions from the manual customer service data; In response to the question matching the answer, it is determined that the manual customer service data passes the quality inspection.
12. The method of claim 9, wherein: Determining that the manual customer service data passes the quality inspection includes: Obtaining a problem handling status from the manual customer service data; In response to the problem handling status being the target status, it is determined that the manual customer service data passes the quality inspection.
13. A training method for retrieving and generating a large model, comprising: Acquire a training sample for a target training task of retrieving and generating a large model; wherein the training sample is generated by a training sample generation method according to any one of claims 1 to 12; The training samples are used to train the retrieval generation model for the target training task to obtain a trained retrieval generation model.
14. A method for generating an answer based on a large model, comprising: Get the input query information; According to the query information, searching in the knowledge base to obtain knowledge fragments matching the query information; According to the query information and the knowledge fragments, a large model is generated by retrieval to obtain an answer to the query information; wherein the large model is trained according to the method of claim 13.
15. A training sample generation device based on a large model, comprising: An acquisition module is used to obtain input query information; A retrieval module, configured to retrieve a first knowledge segment matching the query information from a knowledge base in response to matching the query information with a reference question in a reference question-answer pair in a target domain; wherein the reference question-answer pair is extracted from manual customer service data; A generation module is used to generate training samples for a target training task of retrieving and generating a large model based on the query information, the reference question and the first knowledge fragment.
16. The device of claim 15, wherein: The target training task includes a model generation training task, and the generation module is used to: In response to the first knowledge fragment being related to a reference answer to the reference question, generating a target answer to the query information using a target macro model according to the query information and the first knowledge fragment; According to the reference answer, obtaining first attribute information of the target answer; A training sample for the model generation training task is generated according to the query information, the first knowledge fragment, the target answer and the first attribute information.
17. The device of claim 16, wherein: The generating module is used for: In response to the value of the first attribute information being greater than a first threshold, generating a training sample corresponding to a first training method of the model generation training task according to the query information, the first knowledge fragment and the target answer; In response to the value of the first attribute information being less than or equal to the first threshold, regenerating an answer using the target macro model according to the query information, the first knowledge fragment, and the target answer; According to the reference answer, obtaining second attribute information of the regenerated answer; In response to the value of the second attribute information being greater than the first threshold, a training sample corresponding to the first training method is generated according to the query information, the first knowledge fragment and the regenerated answer.
18. The device of claim 17, wherein: The generating module is further used for: In response to the value of the second attribute information being greater than the first threshold, determining, according to the query information, the first knowledge fragment, and the regenerated answer, a positive sample corresponding to the second training method of the model generation training task; Determining negative samples corresponding to the second training method according to the query information, the first knowledge fragment and the target answer; According to the positive samples and negative samples corresponding to the second training method, training samples corresponding to the second training method are determined.
19. The device of claim 17, wherein: The generating module is used for: Based on the query information and the first knowledge fragment, using the target big model to reflect on the target answer to obtain answer improvement suggestions; The answer is regenerated using the target big model according to the query information, the first knowledge fragment, the target answer and the answer improvement suggestion.
20. The device of claim 19, wherein: The target training task also includes a model reflection training task, and the generation module is further used to: In response to the value of the second attribute information being greater than the first threshold, a training sample for the model reflection training task is generated based on the query information, the first knowledge fragment, the answer improvement suggestion, the target answer and the regenerated answer.
21. The apparatus of claim 15, wherein: The target training task includes a query rewriting training task, and the generating module is used to: In response to the first knowledge fragment being irrelevant to the reference answer to the reference question, rewriting the query information using the target macro model to obtain rewritten query information; According to the rewritten query information, searching the knowledge base to obtain a second knowledge fragment matching the rewritten query information; In response to the second knowledge fragment being related to the reference answer, a training sample for the query rewriting training task is generated according to the query information and the rewritten query information.
22. The device of claim 21, wherein: The generating module is used for: Obtaining the original problem and the essential problem corresponding to the original problem; According to the original question and the essential question, the query information is rewritten using the target big model to obtain the rewritten query information.
23. The apparatus of any one of claims 15 to 22, further comprising: an extraction module, configured to extract candidate question-answer pairs from the manual customer service data in response to determining that the manual customer service data passes the quality inspection; A determination module is used to determine a reference question-answer pair in the target domain from the candidate question-answer pairs.
24. The device of claim 23, wherein: The determining module is used to: Performing a quality score on the candidate question-answer pair to obtain a quality score of the candidate question-answer pair; Filtering out reference question-answer pairs having a quality score greater than a second threshold from the candidate question-answer pairs; Performing field classification on the reference question-answer pairs whose quality scores are greater than a second threshold to obtain reference question-answer pairs in various fields; A reference question-answer pair in the target field is obtained from the reference question-answer pairs in each field.
25. The apparatus of claim 23, wherein: The extraction module is used to: Extracting questions and answers corresponding to the questions from the manual customer service data; In response to the question matching the answer, it is determined that the manual customer service data passes the quality inspection.
26. The apparatus of claim 23, wherein: The extraction module is used to: Obtaining a problem handling status from the manual customer service data; In response to the problem handling status being the target status, it is determined that the manual customer service data passes the quality inspection.
27. A training device for retrieving and generating a large model, comprising: An acquisition module, used to acquire training samples for a target training task of retrieving and generating a large model; wherein the training samples are generated according to a training sample generation method according to any one of claims 1 to 12; The training module is used to train the retrieval generation model for the target training task using the training samples to obtain the trained retrieval generation model.
28. A device for generating an answer based on a large model, comprising: An acquisition module is used to obtain input query information; A retrieval module, used to search the knowledge base according to the query information to obtain knowledge fragments matching the query information; A generation module is used to generate a large model by retrieval based on the query information and the knowledge fragments to obtain an answer to the query information; wherein the large model generated by retrieval is trained according to the method of claim 13.
29. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.
30. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-14.
31. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 14.
Citation Information
Cited By
LLM-RAG sample construction method and device, and storage medium
CN120705581A