Question and answer model training and question and answer method, device and equipment based on large model
Through the training method of big model-based question-and-answer model, bad example query text is processed and corrected, and the question-and-answer model is trained, which solves the problem of a large number of bad examples in the business and improves the generation efficiency and accuracy of answer text.
Patent Information
- Application Number
- CN202510125107.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-10
AI Technical Summary
During the business operation, there are a large number of bad cases, and it is necessary to effectively reduce the number of bad cases to ensure the normal operation of the business.
The big model-based question-answer model training method is adopted to process bad example query text through the initial question-and-answer model, correct the answer text based on the big model, generate preferred sample pairs, and train the initial question-and-answer model through these sample pairs to obtain the target question-and-answer model.
It effectively reduces the number of bad examples and improves the generation efficiency and accuracy of answer text.
Smart Images

Figure CN120124706A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the fields of large models, intelligent systems, natural language processing, etc. In particular, it relates to a question-answering model training and question-answering method, device, equipment, medium, and product based on a large model. Background Art
[0002] During the operation of a business, some bad cases will be generated. To ensure the normal operation of the business, it is necessary to reduce the bad cases.
[0003] How to reduce the number of bad cases is a problem that needs to be solved. Summary of the Invention
[0004] The present disclosure provides a question-answering model training and question-answering method, device, equipment, medium, and product based on a large model.
[0005] According to one aspect of the present disclosure, there is provided a question-answering model training method based on a large model, including: processing a bad case query text by using an initial question-answering model to obtain an initial answer text; correcting the initial answer text based on the large model to obtain a corrected answer text; the preference degree of the corrected answer text is higher than that of the initial answer text; generating a first preference sample pair based on the corrected answer text and the initial answer text; obtaining a target preference sample pair at least based on the first preference sample pair; training the initial question-answering model based on the target preference sample pair to obtain a target question-answering model.
[0006] According to another aspect of the present disclosure, there is provided a question-answering method, including: obtaining a query text; processing the query text by using a target question-answering model to obtain an answer text; the target question-answering model is trained by using the method described in any one of the above aspects.
[0007] According to another aspect of the present disclosure, there is provided a question-answering model training device based on a large model, including: a question-answering module for processing a bad case query text by using an initial question-answering model to obtain an initial answer text; a correction module for correcting the initial answer text based on the large model to obtain a corrected answer text; the preference degree of the corrected answer text is higher than that of the initial answer text; a generation module for generating a first preference sample pair based on the corrected answer text and the initial answer text; an obtaining module for obtaining a target preference sample pair at least based on the first preference sample pair; a training module for training the initial question-answering model based on the target preference sample pair to obtain a target question-answering model.
[0008] According to another aspect of the present disclosure, there is provided a question-answering device, including: an acquisition module configured to acquire a query text; a question-answering module configured to process the query text by using a target question-answering model to obtain an answer text; the target question-answering model is trained by using the method described in any one of the above aspects.
[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in any one of the above aspects.
[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in any one of the above aspects.
[0011] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method described in any one of the above aspects.
[0012] According to the embodiments of the present disclosure, the number of bad examples can be efficiently reduced, thereby improving the generation efficiency and accuracy of the answer text.
[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. Description of the Drawings
[0014] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0015] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;
[0016] Figure 2 is a schematic diagram of an implementation system for implementing the question-answering method of the embodiments of the present disclosure;
[0017] Figure 3 is a schematic diagram of the overall architecture for implementing the question-answering model training method based on a large model of the embodiments of the present disclosure;
[0018] Figure 4 is a schematic diagram according to the second embodiment of the present disclosure;
[0019] Figure 5 is a schematic diagram according to the third embodiment of the present disclosure;
[0020] Figure 6 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0021] Figure 7 is a schematic diagram according to the fifth embodiment of the present disclosure;
[0022] Figure 8 is a schematic diagram of an electronic device for implementing the large model-based question answering model training method or question answering method of the embodiments of the present disclosure. Detailed implementation manners
[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0024] In the related art, the method of manually correcting bad examples is usually adopted to reduce the number of bad examples, but there are problems in terms of efficiency.
[0025] In order to efficiently reduce the number of bad examples, the present disclosure provides the following embodiments.
[0026] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. This embodiment provides a large model-based question answering model training method, as Figure 1 shown, the method includes:
[0027] 101. Process the bad example query text using an initial question answering model to obtain an initial answer text.
[0028] 102. Correct the initial answer text based on a large model to obtain a corrected answer text; the preference degree of the corrected answer text is higher than that of the initial answer text.
[0029] 103. Generate a first preference sample pair based on the corrected answer text and the initial answer text.
[0030] 104. Obtain a target preference sample pair based at least on the first preference sample pair.
[0031] 105. Adjust the model parameters of the initial question answering model based on the target preference sample pair to obtain a target question answering model.
[0032] In a question answering scenario, a question answering model can process an input query text (query) and the output is an answer text.
[0033] A bad case query text refers to a query text corresponding to an answer text that does not conform to the preferences of the target object.
[0034] For example, taking the target object as a user, feedback data (such as likes and dislikes) of the user for the answer text can be collected, and the query text corresponding to the answer text that does not conform to the user's preferences (such as dislikes) is used as the bad case query text.
[0035] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0036] An initial question and answer model is a pre-set deep learning model for question and answer. For example, the pre-trained large model can be fine-tuned using sample data of the question and answer scenario to obtain the initial question and answer model.
[0037] After obtaining the bad case query text, the bad case query text is input into the initial question and answer model, and the initial text large model can output at least one initial answer text based on the bad case query text.
[0038] The number of initial answer texts can be configured. For example, it can be preset to generate N (N is a positive integer) initial answer texts.
[0039] The initial answer text is generated based on the bad case query text and usually does not conform to the preferences of the target object. Therefore, the initial answer text can be corrected so that the corrected text is more in line with the preferences of the target object.
[0040] Specifically, the initial answer text can be corrected based on the large model, and the corrected text obtained based on the large model is called the corrected answer text.
[0041] The large model refers to the Large Language Model (LLM). LLM is a hot technology in the field of AI in recent years. LLM is a natural language processing model based on deep learning, which has a huge number of parameters and a complex structure, enabling it to process and understand a large amount of natural language data and perform well in various natural language processing tasks, including but not limited to text generation, language understanding, machine translation, etc.
[0042] After obtaining the corrected answer text, a preference sample pair is formed based on the corrected answer text and the initial answer text. For the sake of distinction, this preference sample pair is called the first preference sample pair.
[0043] The preference sample pair includes a positive sample and a negative sample. The positive sample is a sample that conforms to the target object's preference, and the negative sample is a sample that does not conform to the target object's preference. Based on the above example, the corrected answer text can be used as the positive sample, and the initial answer text can be used as the negative sample.
[0044] After obtaining the first preference sample pair, obtain the target preference sample pair based on at least the first preference sample pair.
[0045] Specifically, the first preference sample pair can be used as the target preference sample pair; or, other preference sample pairs can be obtained in other ways, and then, based on the first preference sample pair and other preference sample pairs, the target preference sample pair is obtained.
[0046] After obtaining the target preference sample pair, use the target preference sample pair to train the initial question-and-answer model to obtain the target question-and-answer model.
[0047] The training objective is to maximize the generation probability of the answer text that conforms to the target object's preference, so that the target question-and-answer model can generate an answer text that is more in line with the target object's preference.
[0048] Specifically, preference-based reinforcement training methods such as KTO and DPO can be used. KTO refers to Kahneman-Tversky Optimization, and DPO refers to Direct Preference Optimization.
[0049] In this embodiment, training the initial question-and-answer model based on the target preference sample pair to obtain the target question-and-answer model can enable the target question-and-answer model to generate an answer text that conforms to the target object's preference, reduce the generation probability of bad cases that do not conform to the preference, thereby efficiently reducing the number of bad cases, and improving the generation efficiency and accuracy of the answer text.
[0050] To better understand the present disclosure, the application scenarios involved in the present disclosure are described as follows:
[0051] Figure 2 It is a schematic diagram of an implementation system for implementing the question-and-answer method of the embodiments of the present disclosure.
[0052] Taking the question-and-answer process as an example, as Figure 2 shown, the implementation system includes: a user terminal 201 and a server 202.
[0053] The user terminal 201 may include: mobile devices such as a personal computer (PC), a laptop computer, a mobile phone, etc. The server 202 may be a local server or a cloud server, and may be a single server or a cluster server.
[0054] A client is deployed on the user terminal 201, and a server is deployed on the server 202.
[0055] The server performs a question-and-answer task based on a question-and-answer model.
[0056] The user can send a query text (query) to the server through the client. The server processes the input query text, and the output is the answer text. After that, the server can feedback the answer text to the client and display it to the user through the client.
[0057] The answer text may not meet the user's preferences. These answer texts that do not meet the user's preferences can be called bad cases. In the related art, in order to correct bad cases, manual analysis and correction can be performed, but this will lead to problems of poor efficiency and poor accuracy.
[0058] To solve the problems caused by manual correction of bad cases, in the embodiments of the present disclosure, an automatic correction method can be adopted to correct bad cases. Specifically, the model parameters of the question-and-answer model can be adjusted so that the question-and-answer model can generate answer texts that meet the user's preferences as much as possible, gradually reduce the number of bad cases, and achieve automatic correction of bad cases.
[0059] Figure 3 It is a schematic diagram of the overall architecture for implementing the question-and-answer model training method based on a large model in the embodiments of the present disclosure.
[0060] As Figure 3 shown, the initially configured question-and-answer model is called the initial question-and-answer model 301. The initial question-and-answer model 301 is trained based on the large model 302 to obtain the target question-and-answer model 303. Initially, the initial question-and-answer model is used to perform the question-and-answer task. After obtaining the target question-and-answer model, the target question-and-answer model is used to replace the initial question-and-answer model, and the target question-and-answer model is used to perform the question-and-answer task.
[0061] The question-and-answer model can also be a large model. That is, relative to the large model that performs the answer correction task, the question-and-answer model can be a relatively small-scale large model. Based on this, the question-and-answer model can be called a student model, and the large model that corrects the initial answer text can be called a teacher model. The student model and the teacher model can both be large models, and the scale of the student model is smaller than that of the teacher model.
[0062] For a large model, the large model performs related operations according to the prompt information (prompt).
[0063] Assume that the prompt information corresponding to the student model (question-and-answer model) is called the first prompt information, and the prompt information corresponding to the teacher model (large model for correcting answers) is called the second prompt information.
[0064] Based on this, the user can send a first prompt to the server through the client. The first prompt contains the query text input by the user. The Q&A model processes the query text in the input first prompt and outputs the answer text.
[0065] During the operation of the Q&A model, answer texts that do not conform to the user's preferences (which can be called bad cases) may be generated. The query text corresponding to the answer text that does not conform to the user's preferences can be called the bad case query text. To ensure the normal execution of the Q&A task, these bad cases need to be corrected to generate answer texts that conform to the user's preferences as much as possible. For this purpose, the model parameter adjustment method can be adopted, that is, training the initial Q&A model to obtain the target Q&A model, so that the target Q&A model can generate answer texts that conform to the user's preferences as much as possible.
[0066] During the training process, after obtaining the bad case query text, input the bad case query text into the initial Q&A model, and the output is the initial answer text. The number of initial answer texts is a preset value.
[0067] Among them, the initial answer text includes at least one of the following: the first answer text with the highest preference degree and the second answer text with the lowest preference degree;
[0068] In the case where the initial answer text includes the first answer text, Figure 1 Step 102 described in, correcting the initial answer text based on the large model to obtain the corrected answer text, includes:
[0069] Correcting the first answer text based on the large model to obtain the corrected answer text;
[0070] In the case where the initial answer text includes the second answer text, Figure 1 Step 103 described in, generating the first preference sample pair based on the corrected answer text and the initial answer text, includes:
[0071] Generating the first preference sample pair based on the corrected answer text and the second answer text.
[0072] This embodiment takes four as an example, and uses Answer A to Answer D to represent them respectively.
[0073] After obtaining the initial answer text, the preference degree of each initial answer text can be obtained, and then the initial answer text with the highest preference degree (which can be called the first answer text) and the initial answer text with the lowest preference degree (which can be called the second answer text) can be obtained. A first preference sample pair is formed based on the first answer text and the second answer text.
[0074] In the fields of natural language processing, information retrieval, or question answering, text preference is used to reflect the preference degree of a target object (such as a user) for text. For example, if a certain text better meets the user's preferences in terms of meeting user needs, relevance, readability, etc. compared to other texts, it can be said that the preference degree of this text is higher than that of other texts.
[0075] Based on this, the first answer text is the text that best meets the user's preferences among multiple initial answer texts, and the second answer text is the text that least meets the user's preferences among multiple initial answer texts.
[0076] Specifically, a preset reward model can be used to calculate the preference score of each initial answer text, and the initial answer text with the highest score is used as the first answer text, and the initial answer text with the lowest score is used as the second answer text. Based on Figure 3 the example, Answer C is the first answer text, and Answer D is the second answer text.
[0077] After obtaining the first answer text, the first answer text is corrected based on a large model. Specifically, the large model (teacher model) can give correction opinions, and then the initial question-answering model (student model) corrects the initial answer text (specifically the first answer text) based on the correction opinions to obtain a corrected answer text.
[0078] For example, the second prompt information corresponding to the teacher model can be generated according to a preset prompt information template, such as "Please act as an expert who understands user preferences in the field of..., and give correction opinions on the following answer:...", fill the first answer text and its corresponding query text into the second prompt information, and the teacher model outputs correction opinions based on the second prompt information.
[0079] After obtaining the correction opinions, the correction opinions can be input into the initial question-answering model to instruct the initial question-answering model to regenerate an answer, such as "Please correct the previous answer text based on the following correction opinions:...", so that the initial question-answering model can output a corrected answer text based on the above prompt information.
[0080] After obtaining the corrected answer text, the large model (teacher model) can also verify the corrected answer text, such as "Please verify whether the following answer text meets the user's preferences and give the verification result:...", if it passes the verification, the corrected answer text and the second answer text (the initial answer text with the lowest preference degree, such as Answer D) are formed into the first preference sample pair, where the corrected answer text is used as the positive sample and the second answer text is used as the negative sample.
[0081] After obtaining the first preference sample pair, it can be used as the target preference sample pair to train the initial question-answering model. To expand the number of samples, other methods can also be used to obtain the second preference sample pair.
[0082] For example, a correction operation for the above initial answer text can be obtained, and based on this correction operation, an annotated answer text can be obtained. Then, based on this annotated answer text and the initial answer text, a second preferred sample pair can be generated.
[0083] Specifically, as Figure 3 shown, after processing the bad example query text using the initial question and answer model, an initial answer text is obtained. After obtaining the initial answer text, the initial answer text can be corrected through an annotation platform, and the corrected answer text is called the annotated answer text. Then, the annotated answer text and the initial answer text are combined into a second preferred sample pair, where the annotated answer text is used as the positive sample and the initial answer text is used as the negative sample.
[0084] After obtaining the first preferred sample pair and the second preferred sample pair, the first preferred sample pair and the second preferred sample pair are combined into a target preferred sample pair, and the initial question and answer model is trained using the target preferred sample pair to obtain a target question and answer model. The training process can be carried out using preference-based reinforcement training methods such as KTO or DPO.
[0085] After obtaining the target question and answer model, the target question and answer model is used to replace the initial question and answer model to perform the question and answer task. In this way, an answer text that meets the user's preferences can be generated, effectively reducing the number of bad examples (answer texts that do not meet the user's preferences), and improving the generation efficiency and accuracy of the answer text.
[0086] Combined with the above application scenarios, the present disclosure also provides the following embodiments.
[0087] Figure 4 is a schematic diagram according to the second embodiment of the present disclosure. This embodiment provides a method for training a question and answer model based on a large model, and the method includes:
[0088] 401. Process the bad example query text using an initial question and answer model to obtain an initial answer text.
[0089] Among them, the bad example query text can be obtained based on user feedback data. For example, the query text corresponding to the answer text that the user has downvoted is used as the bad example query text. Then, the bad example query text is input into the initial question and answer model, and a preset number of initial answer texts are output.
[0090] 402. Use a preset reward model to score the initial answer text to obtain the preference degree score of the initial answer text.
[0091] 403. Based on the preference degree score, obtain the first answer text with the highest preference degree and the second answer text with the lowest preference degree.
[0092] Among them, the reward model is a pre-trained deep learning model. Its input is text, and its output is the preference score of the text. The higher the preference score, the higher the preference for the corresponding text, that is, the more it conforms to the user's preference.
[0093] Therefore, after obtaining the initial answer text, the initial answer text can be input into the reward model, and the output is the preference score of the initial answer text.
[0094] After obtaining the preference scores of the initial answer texts, the initial answer text with the highest preference score is used as the first answer text, and the initial answer text with the lowest preference score is used as the second answer text.
[0095] In this embodiment, the reward model is used to calculate the preference scores of the initial answer texts, and the first answer text and the second answer text are obtained based on the preference scores, which can improve the accuracy of the first answer text and the second answer text, and further improve the accuracy of the target question-answering model.
[0096] 404. Correct the first answer text based on the large model to obtain a corrected answer text.
[0097] Since the large model has good analysis ability, correcting the first answer text through the large model can improve the accuracy of the corrected answer text, and further improve the accuracy of the target question-answering model.
[0098] Specifically, the large model can output a corrected answer text based on the input first answer text; or, the large model can be used to process the input initial answer text to output the correction opinion of the initial answer text; the initial question-answering model is used to correct the initial answer text based on the input correction opinion to output the corrected answer text.
[0099] That is, specifically, the large model gives the correction opinion, and then the initial answer model performs the specific correction operation.
[0100] Generally speaking, the large model is universal in various fields, and the initial question-answering model is dedicated to a specific field. The large model has strong analysis ability and can analyze the content, grammar, conciseness, etc. of the first answer text to analyze whether it conforms to the user's preference and give a correction opinion. The initial answer model can give a corrected answer text that better meets the requirements of the specific field based on this correction opinion.
[0101] Therefore, by the large model giving the correction opinion and the initial question-answering model obtaining the corrected answer text based on the correction opinion, the advantages of the large model and the initial question-answering model can be combined to improve the matching degree of the corrected answer text with the specific field, and further improve the matching performance of the target question-answering model.
[0102] In some embodiments, after obtaining the corrected answer text output by the initial Q&A model, the large model may also be used to verify the corrected answer text, so as to perform subsequent processing based on the verified corrected answer text.
[0103] In this embodiment, verifying the corrected answer text through the large model can further improve the accuracy of the corrected answer text, and thus improve the accuracy of the target Q&A model.
[0104] 405. Generate the first preference sample pair based on the corrected answer text and the second answer text.
[0105] Specifically, the corrected answer text can be used as the positive sample, and the second answer text can be used as the negative sample to form the first preference sample pair.
[0106] In this embodiment, since the corrected answer text is obtained based on the first answer text with the highest preference degree, and the second answer text has the lowest preference degree, forming the first preference sample pair based on the corrected answer text and the second answer text can improve the contrast between the positive and negative samples in the first preference sample pair, and thus improve the accuracy of the target Q&A model.
[0107] 406. Obtain the correction operation for the initial answer text; obtain the labeled answer text according to the correction operation.
[0108] 407. Generate the second preference sample pair based on the labeled answer text and the initial answer text.
[0109] In this embodiment, obtaining the second preference sample pair through the correction operation can effectively utilize the resources of the correction platform and improve the resource utilization rate.
[0110] 408. Obtain the target preference sample pair based on the first preference sample pair and the second preference sample pair.
[0111] After obtaining the initial answer text, the correction method can also be used to obtain the labeled answer text. Then, the labeled answer text is used as the positive sample, and the initial answer text is used as the negative sample to form the second preference sample pair.
[0112] Then, the first preference sample pair and the second preference sample pair are combined to form the target preference sample pair.
[0113] In this embodiment, obtaining the target preference sample pair based on the first preference sample pair and the second preference sample pair can obtain preference sample pairs from multiple sources, expand the number of training samples, and improve the accuracy of the target Q&A model.
[0114] 409. Train the initial Q&A model based on the target preference sample pair to obtain the target Q&A model.
[0115] After obtaining the target preference sample pairs, training methods such as KTO and DPO can be used to train the initial question-and-answer model to obtain the target question-and-answer model.
[0116] After obtaining the target question-and-answer model, use the target question-and-answer model to replace the initial question-and-answer model to perform the question-and-answer task. In this way, answer texts that conform to the user's preferences can be generated, efficiently reducing the number of bad cases (answer texts that do not conform to the user's preferences), and improving the generation efficiency and accuracy of the answer texts.
[0117] Figure 5 It is a schematic diagram according to the third embodiment of the present disclosure. This embodiment provides a question-and-answer method, and the method includes:
[0118] 501. Obtain a query text.
[0119] 502. Use the target question-and-answer model to process the query text to obtain an answer text.
[0120] Wherein, the target question-and-answer model is trained by the method described in any of the above embodiments.
[0121] During the question-and-answer process, the query text input by the user can be input into the target question-and-answer model, and the output is the answer text.
[0122] In this embodiment, since the target question-and-answer model can generate answer texts that conform to the preferences of the target object and reduce the generation probability of bad cases that do not conform to the preferences, the number of bad cases can be efficiently reduced, and the generation efficiency and accuracy of the answer texts are improved.
[0123] Figure 6 It is a schematic diagram according to the fourth embodiment of the present disclosure. This embodiment provides a question-and-answer model training device based on a large model. The device 600 includes: a question-and-answer module 601, a correction module 602, a generation module 603, an acquisition module 604, and a training module 605.
[0124] The question-and-answer module 601 is used to process the bad case query text with the initial question-and-answer model to obtain an initial answer text; the correction module 602 is used to correct the initial answer text based on the large model to obtain a corrected answer text; the preference degree of the corrected answer text is higher than that of the initial answer text; the generation module 603 is used to generate a first preference sample pair based on the corrected answer text and the initial answer text; the acquisition module 604 is used to obtain the target preference sample pair at least based on the first preference sample pair; the training module 605 is used to train the initial question-and-answer model based on the target preference sample pair to obtain the target question-and-answer model.
[0125] In this embodiment, by training an initial question-and-answer model based on target preference samples, a target question-and-answer model can be obtained, enabling the target question-and-answer model to generate answer texts that conform to the preferences of the target object, reducing the generation probability of bad examples that do not conform to the preferences, thereby efficiently reducing the number of bad examples and improving the generation efficiency and accuracy of the answer texts.
[0126] In some embodiments, the initial answer text includes at least one of the following: a first answer text with the highest preference degree and a second answer text with the lowest preference degree;
[0127] When the initial answer text includes the first answer text, the correction module 602 is further configured to: correct the first answer text based on the large model to obtain the corrected answer text;
[0128] When the initial answer text includes the second answer text, the generation module 603 is further configured to: generate the first preference sample pair based on the corrected answer text and the second answer text.
[0129] In this embodiment, by correcting the first answer text through the large model, the accuracy of the corrected answer text can be improved, and further the accuracy of the target question-and-answer model can be improved.
[0130] In this embodiment, since the corrected answer text is obtained based on the first answer text with the highest preference degree and the second answer text has the lowest preference degree, forming the first preference sample pair based on the corrected answer text and the second answer text can improve the contrast between the positive and negative samples in the first preference sample pair, and further improve the accuracy of the target question-and-answer model.
[0131] In some embodiments, the initial answer text includes: a first answer text with the highest preference degree and a second answer text with the lowest preference degree; the apparatus 600 further includes: a scoring module, configured to score the initial answer text using a preset reward model to obtain the preference score of the initial answer text; and obtain the first answer text and the second answer text based on the preference score.
[0132] In this embodiment, calculating the preference score of the initial answer text using the reward model and obtaining the first answer text and the second answer text based on the preference score can improve the accuracy of the first answer text and the second answer text, and further improve the accuracy of the target question-and-answer model.
[0133] In some embodiments, the correction module 602 is further configured to: process the input initial answer text using the large model to output a correction opinion on the initial answer text; and correct the initial answer text based on the input correction opinion using the initial question-answering model to output the corrected answer text.
[0134] In this embodiment, by using the large model to give a correction opinion and the initial question-answering model to obtain the corrected answer text based on the correction opinion, the advantages of the large model and the initial question-answering model can be combined to improve the matching degree of the corrected answer text with a specific field, and further improve the matching performance of the target question-answering model.
[0135] In some embodiments, the generation module 603 is further configured to: verify the corrected answer text using the large model; and in response to determining that the corrected answer text passes the verification, generate the first preference sample pair based on the corrected answer text and the initial answer text.
[0136] In this embodiment, by using the large model to verify the corrected answer text, the accuracy of the corrected answer text can be further improved, and thus the accuracy of the target question-answering model can be improved.
[0137] In some embodiments, the apparatus 600 further includes: an annotation module, configured to obtain a correction operation for the initial answer text; obtain an annotated answer text according to the correction operation; and generate a second preference sample pair based on the annotated answer text and the initial answer text.
[0138] In this embodiment, by obtaining the second preference sample pair through the correction method, the correction platform resources can be effectively utilized, and the resource utilization rate can be improved.
[0139] In some embodiments, the obtaining module 604 is further configured to: obtain the target preference sample pair based on the first preference sample pair and the second preference sample pair.
[0140] In this embodiment, by obtaining the target preference sample pair based on the first preference sample pair and the second preference sample pair, preference sample pairs from multiple sources can be obtained, the number of training samples can be expanded, and the accuracy of the target question-answering model can be improved.
[0141] Figure 7 FIG. is a schematic diagram according to the fifth embodiment of the present disclosure. In this embodiment, a question-answering apparatus is provided. The apparatus 700 includes: an obtaining module 701 and a question-answering module 702.
[0142] The obtaining module 701 is configured to obtain a query text; the question-answering module 702 is configured to process the query text using a target question-answering model to obtain an answer text; the target question-answering model is trained by the method described in any of the above embodiments.
[0143] In this embodiment, since the target Q&A model can generate answer texts that conform to the preferences of the target object, reducing the generation probability of bad examples that do not conform to the preferences, the number of bad examples can be efficiently reduced, improving the generation efficiency and accuracy of the answer texts.
[0144] It can be understood that in the embodiments of the present disclosure, the same or similar content in different embodiments can be referred to each other.
[0145] It can be understood that the "first", "second", etc. in the embodiments of the present disclosure are only used for distinction and do not represent the level of importance, the sequence of time, etc.
[0146] It can be understood that if there is no special limitation on the sequence of steps involved in the process, it means that the temporal relationship between these steps is not limited.
[0147] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0148] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0149] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device 800 is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0150] As Figure 8 shown, the electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0151] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0152] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as a large model-based question-answering model training method or a question-answering method. For example, in some embodiments, the large model-based question-answering model training method or the question-answering method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the large model-based question-answering model training method or the question-answering method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the large model-based question-answering model training method or the question-answering method in any other suitable manner (e.g., by means of firmware).
[0153] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0154] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable task processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0155] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0156] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0157] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0158] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0159] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0160] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A question-answering model training method based on a large model, comprising: Use the initial question-answering model to process the bad query text to obtain the initial answer text; Modifying the initial answer text based on the large model to obtain a modified answer text; The preference of the revised answer text is higher than the preference of the initial answer text; Based on the revised answer text and the initial answer text, generating a first preference sample pair; Based at least on the first preference sample pair, obtaining a target preference sample pair; Based on the target preference sample pairs, the initial question-answering model is trained to obtain a target question-answering model.
2. The method according to claim 1, wherein: The initial answer text includes at least one of the following: a first answer text with the highest preference and a second answer text with the lowest preference; In the case where the initial answer text includes the first answer text, the step of correcting the initial answer text based on the large model to obtain a corrected answer text includes: Modifying the first answer text based on the large model to obtain the modified answer text; In a case where the initial answer text includes the second answer text, generating a first preference sample pair based on the revised answer text and the initial answer text comprises: Based on the revised answer text and the second answer text, the first preference sample pair is generated.
3. The method according to claim 1, wherein: The initial answer text includes: a first answer text with the highest preference and a second answer text with the lowest preference; The method further comprises: Using a preset reward model, scoring the initial answer text to obtain a preference score for the initial answer text; The first answer text and the second answer text are obtained based on the preference score.
4. The method according to claim 1, wherein: The step of correcting the initial answer text based on the large model to obtain a corrected answer text includes: Using the large model, the input initial answer text is processed to output revised opinions on the initial answer text; The initial question-answer model is used to revise the initial answer text based on the input revision opinion to output the revised answer text.
5. The method according to claim 1, wherein: The step of generating a first preference sample pair based on the revised answer text and the initial answer text comprises: Using the large model to verify the revised answer text; In response to determining that the revised answer text passes verification, the first preference sample pair is generated based on the revised answer text and the initial answer text.
6. The method according to claim 1, further comprising: Obtaining an input correction operation for the initial answer text; Obtaining the marked answer text according to the correction operation; Based on the marked answer text and the initial answer text, a second preference sample pair is generated.
7. The method according to claim 6, wherein: The acquiring a target preference sample pair based at least on the first preference sample pair comprises: Based on the first preference sample pair and the second preference sample pair, the target preference sample pair is obtained.
8. A question-answering method, comprising: Get the query text; Using a target question-answering model, the query text is processed to obtain an answer text; The target question answering model is trained using the method described in any one of claims 1-7.
9. A question-answering model training device based on a large model, comprising: A question-answering module is used to process the bad query text using the initial question-answering model to obtain the initial answer text; A correction module, used for correcting the initial answer text based on the large model to obtain a corrected answer text; The preference of the revised answer text is higher than the preference of the initial answer text; A generating module, configured to generate a first preference sample pair based on the revised answer text and the initial answer text; An acquisition module, configured to acquire a target preference sample pair based at least on the first preference sample pair; A training module is used to train the initial question-answering model based on the target preference sample pairs to obtain a target question-answering model.
10. The device according to claim 9, wherein: The initial answer text includes at least one of the following: a first answer text with the highest preference and a second answer text with the lowest preference; In the case where the initial answer text includes the first answer text, the correction module is further used to: correct the first answer text based on the large model to obtain the corrected answer text; In the case where the initial answer text includes the second answer text, the generating module is further configured to generate the first preference sample pair based on the revised answer text and the second answer text.
11. The device according to claim 9, wherein: The initial answer text includes: a first answer text with the highest preference and a second answer text with the lowest preference; The device also includes: A scoring module is used to score the initial answer text using a preset reward model to obtain a preference score for the initial answer text; and to obtain the first answer text and the second answer text based on the preference score.
12. The device according to claim 9, wherein: The correction module is further used for: Using the large model, the input initial answer text is processed to output revised opinions on the initial answer text; The initial question-answer model is used to revise the initial answer text based on the input revision opinion to output the revised answer text.
13. The device according to claim 9, wherein: The generating module is further used for: Using the large model to verify the revised answer text; In response to determining that the revised answer text passes verification, the first preference sample pair is generated based on the revised answer text and the initial answer text.
14. The apparatus according to claim 9, further comprising: A marking module, used for obtaining a correction operation inputted for the initial answer text; Obtaining the marked answer text according to the correction operation; And, based on the marked answer text and the initial answer text, a second preference sample pair is generated.
15. The device according to claim 14, wherein: The acquisition module is further used for: Based on the first preference sample pair and the second preference sample pair, the target preference sample pair is obtained.
16. A question-answering device, comprising: The acquisition module is used to obtain the query text; A question-answering module, used to process the query text using a target question-answering model to obtain an answer text; The target question answering model is trained using the method described in any one of claims 1-7.
17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
19. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.