Large model-based question and answer method, device, medium, equipment and program product
By introducing a dual question-answering model mechanism into the large language model, the first question-answering model is used to generate and evaluate the answer. If the answer is not satisfactory, it is regenerated by the second question-answering model. This solves the problem of low quality of LLM-generated content and achieves efficient answer evaluation and quality assurance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-27
AI Technical Summary
Existing Large Language Models (LLMs) suffer from illusions when generating content, resulting in low-quality generated content. Effective evaluation mechanisms are needed to improve the quality of generated content.
The first question-answering model generates and evaluates the answer. If the evaluation result is unsatisfactory, the second question-answering model, which has a higher performance, is used to regenerate the answer. The evaluation result is combined with the identifier to ensure the quality of the answer.
This allows for simultaneous evaluation during answer generation, improving evaluation efficiency, reducing evaluation costs, and ensuring the quality of the final answers provided to users.
Smart Images

Figure CN121303368B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of large models, in particular, to a large model-based question answering method, device, medium, equipment and program product. BACKGROUND
[0002] With the development of large models, such as LLM (Large Language Model), in order to bring convenience to people's life, LLM has been widely used in various scenarios, such as knowledge question answering and content creation, etc.
[0003] In the related art, the generation ability of LLM has reached a certain level, but due to various factors, the generation of LLM still has illusion. Therefore, in order to provide high-quality generated content for people, the evaluation of generated content becomes an important research direction. SUMMARY
[0004] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0005] In a first aspect, the present disclosure provides a large model-based question answering method, comprising:
[0006] obtaining a first question;
[0007] generating, by a first question answering large model, a first reply to the first question based on the first question, the first reply comprising a first answer to the first question and an evaluation result of the first answer;
[0008] determining whether an answer to the first question needs to be regenerated based on the evaluation result of the first answer in the first reply;
[0009] in a case where it is determined that an answer to the first question needs to be regenerated, generating, by a second question answering large model, a second answer to the first question based on the first question, the inference performance of the second question answering large model being higher than that of the first question answering large model.
[0010] In a second aspect, the present disclosure provides a large model-based question answering device, comprising:
[0011] a first obtaining module configured to obtain a first question;
[0012] The first generation module is configured to generate, by the first question and answer large model, a first reply to the first question based on the first question, the first reply including a first answer to the first question and an evaluation result of the first answer.
[0013] The first determination module is configured to determine, based on the evaluation result of the first answer in the first reply, whether the answer to the first question needs to be regenerated.
[0014] The second generation module is configured to generate, by a second question and answer large model, a second answer to the first question based on the first question in a case where it is determined that the answer to the first question needs to be regenerated, the inference performance of the second question and answer large model being higher than the inference performance of the first question and answer large model.
[0015] In a third aspect, the present disclosure provides a computer readable medium having a computer program stored thereon, the computer program being executed by a processing device to implement the steps of the question and answer method of the first aspect.
[0016] In a fourth aspect, the present disclosure provides an electronic device, comprising:
[0017] A storage device having a computer program stored thereon;
[0018] A processing device configured to execute the computer program in the storage device to implement the steps of the question and answer method of the first aspect.
[0019] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program, the computer program being executed by a processor to implement the steps of the question and answer method of the first aspect.
[0020] Through the above technical solution, while generating the first answer to the first question by the first question and answer large model, the evaluation result of the first answer is generated by the first question and answer large model, so that the generation of the answer and the evaluation of the answer are realized by the first question and answer large model without introducing other evaluation models, the evaluation efficiency is improved and the evaluation cost is reduced, and in addition, in a case where it is determined based on the evaluation result that the answer to the first question needs to be regenerated, the second answer to the first question is generated by the second question and answer large model based on the first question, the performance of which is higher than that of the first question and answer large model, thereby ensuring the quality of the answer finally provided to the user.
[0021] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0022] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings. The same or similar components have the same or similar reference labels. It should be understood that the drawings are not necessarily to scale, with emphasis instead being placed upon illustrating the principles of the embodiments of the present disclosure. In the drawings:
[0023] Figure 1 is a flowchart of a large model-based question answering method according to an embodiment of the present disclosure;
[0024] Figure 2 is a process diagram of a large model-based question answering method according to an embodiment of the present disclosure;
[0025] Figure 3 is a process diagram of determining a noise sample according to an embodiment of the present disclosure;
[0026] Figure 4 is a process diagram of a third question answering large model in a forward propagation process according to an embodiment of the present disclosure;
[0027] Figure 5 is a process diagram of updating a third question answering large model according to an embodiment of the present disclosure;
[0028] Figure 6 is another process diagram of a third question answering large model in a forward propagation process according to an embodiment of the present disclosure;
[0029] Figure 7 is a block diagram of a large model-based question answering apparatus according to an embodiment of the present disclosure;
[0030] Figure 8 is a structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several embodiments of the present disclosure have been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the present disclosure. It is to be understood that the drawings and descriptions are illustrative and explanatory only, and are not intended to be limiting to the present disclosure.
[0032] It should be understood that the various steps of the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this regard.
[0033] As used herein, the term "includes" and its variants are to be read to be analogous to "comprises," or "comprising." The term "based on" is to be read as "based, at least in part, on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments." Related terms have corresponding meanings.
[0034] It should be noted that the terms "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0035] It should be noted that the terms "one", "multiple" mentioned in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.
[0036] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0037] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the permission of the user should be obtained through appropriate means according to relevant laws and regulations.
[0038] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0039] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may, for example, carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0040] It can be understood that the above notification and user permission obtaining process is only illustrative, and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0041] Meanwhile, it can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the technical solution should comply with the requirements of the corresponding laws, regulations and relevant provisions.
[0042] Figure 1 is a flowchart of a large model-based question answering method according to an embodiment of the present disclosure, referring to Figure 1 The large model-based question answering method can include steps 110, 120, 130, and 140.
[0043] In step 110, a first question is obtained.
[0044] Here, the first question can be provided by a user to a first question answering large model.
[0045] In step 120, a first reply to the first question is generated by the first question answering large model based on the first question, the first reply including a first answer to the first question and an evaluation result of the first answer.
[0046] Here, the evaluation result is determined based at least on a target identifier generated by the first question answering large model for the first question, as an example, the evaluation result can be the target identifier itself, as another example, the evaluation result can also be a quality score determined based on the target identifier and a confidence of other identifiers, it should be noted that the target identifier is a character output synchronously when the first question answering large model outputs the first answer, used to represent a positive evaluation result or a negative evaluation result. The target identifier can be a first identifier or a second identifier, the first identifier being used to represent a positive result of the first question answering large model for the first answer, the second identifier being used to represent a negative result of the first question answering large model for the first answer, and the other identifier being another one of the first identifier and the second identifier other than the target identifier.
[0047] In the case where the evaluation result is a quality score determined based on the target identifier and a confidence of other identifiers, the above step 120 can be implemented in the following manner: determining the evaluation result of the first answer based on a confidence vector corresponding to the last character output by the first question answering large model, the last character being the target identifier, the confidence vector including a confidence of the target identifier and a confidence of the other identifier, the other identifier being another one of the first identifier and the second identifier other than the target identifier.
[0048] It should be noted that in the process of outputting each character, the first question and answer large model generates a corresponding confidence vector, which includes the confidence of each character in the vocabulary table, and the first question and answer large model selects the character with the highest confidence in the confidence vector as the current output. All output characters are combined according to the output time, that is, the first reply can be obtained, wherein the confidence can be understood as the intermediate output result of the large model in the process of outputting the first reply.
[0049] In determining the evaluation result of the first answer, only the confidence of the target identifier in the confidence vector corresponding to the last character (target identifier) output by the first question and answer large model and the confidence of other identifiers in the confidence vector are extracted. The quality score is determined based on the confidence of the target identifier and the confidence of the other identifiers. The quality score is the evaluation result of the first answer.
[0050] In the case where the target identifier is the first identifier, the evaluation result of the first answer is determined using the following formula (1):
[0051] (1);
[0052] In the above formula (1), is the quality score, is the confidence of the first identifier in the confidence vector corresponding to the last character output by the first question and answer large model, is the confidence of the second identifier in the confidence vector corresponding to the last character output by the first question and answer large model, i is the index of the first identifier, and j is the index of the second identifier. For explanations and descriptions of the index, refer to the related embodiments described below, which will not be repeated here.
[0053] In the case where the target identifier is the second identifier, the evaluation result of the first answer is determined using the following formula (2):
[0054] (2);
[0055] In the above formula (2), the explanations and descriptions of the various parameters can refer to the related embodiments described above, which will not be repeated here.
[0056] In a possible manner, the confidence vector includes the index corresponding to each character in the vocabulary table and the confidence corresponding to the character. The first identifier, the index of the first identifier, the second identifier, and the index of the second identifier can be pre-stored in the vocabulary table. Therefore, when extracting the confidence of the target identifier and the confidence of the other identifiers from the confidence vector, the index of the target identifier and the index of the other identifiers can be used to query the index in the confidence vector, and then the confidence corresponding to the corresponding index can be extracted.
[0057] The following is the first identifier <ga>, the second identifier is <ba>The embodiments of the present disclosure are explained and described.
[0058] As an example, the target identifier is a character that the first question and answer large model continues to output after outputting the answer. For the first question "Who is the teacher of Sun Wukong", the first reply generated by the first question and answer large model can be "Sun Wukong has two teachers, corresponding to his two key life stages of becoming a god and becoming a Buddha, and the two are significantly different in identity, teaching content and influence on Sun Wukong, respectively, the teacher of the Dharma-Pho Ti Zu Shi and the guide of the journey-Tang Seng. <ga>". Among them, the first reply "Sun Wukong's two teachers correspond to his two key life stages of becoming a god and becoming a Buddha. The two are significantly different in identity, content of transmission, and influence on Sun Wukong. They are the teacher of the Dharma - Bodhi Master and the guide of the pilgrimage - Tang Monk." is the first answer, and the first reply " <ga>" as a target identifier.
[0059] In step 130, it is determined whether the answer to the first question needs to be regenerated based on the evaluation result of the first answer in the first reply.
[0060] Taking the above evaluation result as an example, the target identifier itself is <ga>In the case where the target identifier in the first reply is <ba>In a case where it is determined that the answer to the first question needs to be regenerated, the second answer to the first question is regenerated based on the first question by using the second Q&A large model.
[0061] In a case where the quality score is less than or equal to the preset first score threshold, it is determined that the answer to the first question needs to be regenerated; and in a case where the quality score is greater than the preset first score threshold, it is determined that the answer to the first question does not need to be regenerated. As an example, the preset first score threshold can be 0.6.
[0062] It is worth noting that for a case where the confidence of the first identifier and the confidence of the second identifier are relatively close, such an answer should essentially be more inclined to be regarded as a low-quality answer, and therefore, by using the confidence as the fine-grained information to determine the evaluation result, the rationality of the answer quality judgment can be improved.
[0063] In step 140, in a case where it is determined that the answer to the first question needs to be regenerated, a second answer to the first question is regenerated based on the first question by using the second Q&A large model, and the inference performance of the second Q&A large model is higher than the inference performance of the first Q&A large model.
[0064] It can be understood that in a case where it is determined that the answer to the first question does not need to be regenerated, the first answer is provided to the user; and in a case where it is determined that the answer to the first question needs to be regenerated, the second answer is provided to the user.
[0065] Figure 2 is a process schematic diagram of a large model-based Q&A method according to an embodiment of the present disclosure. Referring to Figure 2 The router determines the subsequent data flow based on the size relationship between the quality score and the preset first score threshold, and in a case where the router determines that the quality score is greater than the preset first score threshold, the first answer output by the first Q&A large model is provided to the user as the answer to the first question. In a case where the router determines that the quality score is less than or equal to the preset first score threshold, the router provides the first question to the second Q&A large model, so that a second answer to the first question is regenerated based on the first question by using the second Q&A large model, and the second answer output by the second Q&A large model is provided to the user as the answer to the first question.
[0066] By the technical solution, the first answer generation model is used to generate the evaluation result of the first answer generated by the first question and answer large model, so that the other evaluation model does not need to be introduced, the generation of the answer and the evaluation of the answer are realized by the first question and answer large model at the same time, the evaluation efficiency is improved, and the evaluation cost is reduced. In addition, in the case where it is determined based on the evaluation result that the answer of the first question needs to be regenerated, a second question and answer large model with higher performance than the first question and answer large model is used to generate a second answer of the first question based on the first question, thereby ensuring the quality of the answer finally provided to the user.
[0067] The training of the first question and answer large model is exemplarily described below.
[0068] In a possible manner, the above question and answer method can further include the following steps: obtaining a first positive sample, the first positive sample including a question sample and an answer sample corresponding to the question sample; based on the first positive sample, constructing a first negative sample; adding a first identifier at the end of the answer sample in the first positive sample to construct a second positive sample; adding a second identifier at the end of the answer sample in the first negative sample to construct a second negative sample, the second positive sample and the second negative sample being used to train a third question and answer large model to obtain the first question and answer large model.
[0069] Here, the question sample corresponds to the first question, and the answer sample corresponds to the first answer. The explanation and description of the question sample and the answer sample can refer to the above related embodiments, which will not be repeated here.
[0070] After obtaining the first positive sample and the first negative sample, the first identifier or the second identifier can be added after the corresponding answer sample, so that the model can continue to output the corresponding first identifier or second identifier after outputting the answer sample.
[0071] In the above manner, the negative sample is constructed directly by using the positive sample, which reduces the construction cost of the negative sample. After obtaining the positive sample and the negative sample, the first identifier or the second identifier is added to obtain the training data for training the third question and answer large model.
[0072] In a possible manner, the third question and answer large model is a pre-trained large model, and the third question and answer large model is fine-tuned based on the training data, so that the first question and answer large model can be obtained.
[0073] In a possible manner, the answer samples in two first positive samples can be exchanged to obtain two first negative samples.
[0074] Since the way of swapping the answer sample in the two first positive samples will destroy the correlation between the question sample and the answer sample, the model can easily identify this sample as a negative sample in the training process, and thus the negative sample constructed in this way has a low contribution value to the evaluation ability of the model in the model training process. Therefore, when constructing the first negative sample based on the first positive sample, a first constraint is set to prohibit destroying the correlation between the question sample and the answer sample in the first negative sample, that is, each first negative sample is constructed from the corresponding first positive sample. The first negative sample includes a question sample and an answer sample corresponding to the question sample. The question sample in the first negative sample is the same as the question sample in the first positive sample, and the answer sample in the first negative sample is obtained by sampling the answer sample in the first positive sample, that is, the answer sample in the first negative sample can be part of the answer sample in the first positive sample.
[0075] In the way of constructing the first negative sample from the corresponding first positive sample, the answer sample in the first positive sample can be randomly sampled. Here, random sampling refers to randomly selecting content at any position as the answer sample in the negative sample. For example, for the answer sample "Isaac Newton" in the first positive sample, the content "Isaac Newton" can be randomly sampled as the answer sample in the first negative sample. Isaac Newton is one of the most influential scientists in the 17th and 18th centuries. His achievements span multiple fields such as physics, mathematics, and astronomy, laying the foundation for classical physics and even shaping the research paradigm of modern science. In this case, the content "Isaac Newton" is sampled from the answer sample as the answer sample in the negative sample. Isaac Newton, whose achievements span physics.
[0076] However, since random sampling results in incoherent sentences in the sampled answer sample, the model can easily identify the sample constructed by random sampling as a negative sample in the training process, that is, the negative sample constructed by random sampling has a low contribution value to the evaluation ability of the model in the model training process. Therefore, a second constraint can be set to prohibit destroying the coherence of the sentences in the answer sample in the first negative sample.
[0077] When constructing the first negative sample based on the second constraint, random truncation can be performed from the beginning or end of the answer sample in the first positive sample. Here, random truncation refers to randomizing the number of characters. The remaining content after truncation is used as the answer sample of the first negative sample.
[0078] From the above content, it can be seen that the above step of constructing the first negative sample based on the first positive sample can be implemented in the following way: based on the first positive sample, the first negative sample is constructed according to the constraint. The sample quality of the first negative sample constructed according to the constraint is higher than that of the first negative sample not constructed according to the constraint.
[0079] As an example, the above constraints include at least one of a first constraint and a second constraint, and the explanation and description of the first constraint and the second constraint can refer to the above related embodiments. The following is an example of constructing a first negative sample based on a first positive sample with constraints.
[0080] The following is an example of the first positive sample:
[0081] Question sample: What are Newton's achievements?
[0082] Answer sample: Isaac Newton (Isaac Newton) is one of the most influential scientists in the 17th and 18th centuries, whose achievements span physics, mathematics, astronomy and other fields, laying the foundation of classical physics, and even deeply shaping the research paradigm of modern science.
[0083] Starting from the first character of the above answer sample, truncate it backward to get the truncated answer sample "18th century's most influential scientist, whose achievements span physics, mathematics, astronomy and other fields, laying the foundation of classical physics, and even deeply shaping the research paradigm of modern science." Newton (Isaac Newton) is one of the most influential scientists in the 17th and 18th centuries, whose achievements span physics, mathematics, astronomy and other fields, laying the foundation of classical physics, and even deeply shaping the research paradigm of modern science.
[0084] Question sample: What are Newton's achievements?
[0085] Answer sample: 18th century's most influential scientist, whose achievements span physics, mathematics, astronomy and other fields, laying the foundation of classical physics, and even deeply shaping the research paradigm of modern science.
[0086] As another example, truncate the above answer sample from the last character forward to get the truncated answer sample "Isaac Newton (Isaac Newton) is one of the most influential scientists in the 17th and 18th centuries, whose achievements span physics, mathematics, and
[0087] Question sample: What are Newton's achievements?
[0088] Answer sample: Isaac Newton (Isaac Newton) is one of the most influential scientists in the 17th and 18th centuries, whose achievements span physics, mathematics,
[0089] By the above scheme, the coherence of the answers in the negative samples and the relevance of the questions and answers in the negative samples can be ensured by constructing the negative samples according to the constraints, the contribution value of the constructed negative samples to the evaluation ability of the model in the model training process is improved, and data basis is provided for the model to learn how to identify and evaluate the negative samples.
[0090] In a possible manner, the first positive sample can include a first sub-positive sample and a second sub-positive sample, the first sub-positive sample can include a first question sample and a first answer sample corresponding to the first question sample, and the second sub-positive sample includes a second question sample and a second answer sample corresponding to the second question sample.
[0091] Here, the difficulty of the large model generating the first answer sample based on the first question sample is higher than the difficulty of the large model generating the second answer sample based on the second question sample, that is, the first sub-positive sample is a sample with higher complexity than the second sub-positive sample. As an example, question and answer data about subjects such as mathematics and physics can be obtained from open source corpus, and the question and answer data is sorted to obtain the first sub-positive sample; as an example, question and answer data of subjects such as history and humanities can be obtained from open source corpus, and the question and answer data is sorted to obtain the second sub-positive sample. By collecting simple and complex question and answer data as training samples at the same time, the ability of the first question and answer large model in complex reasoning scenarios can be considered.
[0092] Among them, the first sub-positive sample and the second sub-positive sample can be respectively constructed based on the above technical scheme for constructing negative samples. For construction methods, please refer to the above related embodiments, which will not be repeated here.
[0093] In a possible manner, before adding the identifier (first identifier or second identifier) to the answer sample in the first positive sample or the first negative sample, the first positive sample and the first negative sample can be data cleaned to filter noise samples in the constructed first positive sample and the first negative sample.
[0094] In this case, before adding the identifier to the answer sample in the first positive sample or the first negative sample, the above question and answer method can further include the following steps: generating an evaluation result of each initial sample based on a preset prompt word and each initial sample by the fourth question and answer large model; determining whether the initial sample is a noise sample based on the evaluation result of the initial sample, and constructing a corresponding target sample using the initial sample that belongs to a non-noise sample.
[0095] In this embodiment, the initial sample is the first positive sample or the first negative sample, and the corresponding target sample is the second positive sample or the second negative sample.
[0096] The evaluation result can include a score and an evaluation reason. The score can be used to determine noise samples subsequently. The evaluation reason can be used by the developer to determine whether the score is reasonable. The manner of determining noise samples based on the score can refer to the related embodiments described below.
[0097] The fourth Q&A large model can evaluate the question sample and the answer sample in the initial sample from different dimensions to obtain an evaluation result. Based on the evaluation result, it can be determined whether the first positive sample is a positive sample or whether the first negative sample is a negative sample.
[0098] As an example, the plurality of dimensions can include at least one of relevance of the question and the answer, completeness of the answer, and correctness of the answer.
[0099] As an example, the preset prompt can be: please analyze the quality of the Q&A data from the dimensions of relevance of the question and the answer, completeness of the answer, and correctness of the answer, and give a score between 0 and 1, and give a reasonable explanation.
[0100] In the embodiment, the step of determining whether the initial sample is a noise sample based on the evaluation result of the initial sample can be obtained in the following manner: for the first positive sample, determining a noise sample based on a preset second score threshold and the score in the evaluation result of the first positive sample; for the first negative sample, determining a noise sample based on a preset third score threshold and the score in the evaluation result of the first negative sample, wherein the preset second score threshold is greater than the preset third score threshold, thereby widening the gap between positive and negative samples and making it easier for the model to learn the characteristics of positive and negative samples.
[0101] Figure 3 is a process diagram for determining a noise sample according to an embodiment of the present disclosure. Referring to Figure 3 The preset prompt, the question sample in the initial sample, and the answer sample in the initial sample are spliced, and the spliced result is input into the fourth Q&A large model to obtain the evaluation result "score is XX, reason is XX" output by the fourth Q&A large model.
[0102] For the first positive sample, if the score in the evaluation result of the first positive sample is greater than 0.6 (i.e., the preset second score threshold), the first positive sample is determined to be a non-noise sample and can be used to construct a corresponding second positive sample; if the score in the evaluation result of the first positive sample is less than or equal to 0.6, the first positive sample is determined to be a noise sample and cannot be used to construct a corresponding second positive sample.
[0103] For the first negative sample, if the score in the evaluation result of the first negative sample is greater than or equal to 0.4 (i.e., a preset third score threshold), it is determined that the first negative sample is a noise sample and cannot be used to construct a corresponding second negative sample; if the score in the evaluation result of the first negative sample is less than 0.4, it is determined that the first negative sample is a non-noise sample and can be used to construct a corresponding second negative sample.
[0104] The preset second score threshold is set to 0.6 and the preset third score threshold is set to 0.4. In this way, for the positive sample, the first positive sample with a score in the evaluation result of 0.5-0.6, which is easy to confuse, can be filtered out. Similarly, for the negative sample, the first negative sample with a score in the evaluation result of 0.4-0.5, which is easy to confuse, can be filtered out.
[0105] After the second positive sample and the second negative sample are constructed, the third question and answer large model can be trained using these samples to obtain the first question and answer large model. The training of the third question and answer large model is exemplarily described below in combination with Figure 4 , Figure 5 and Figure 6 .
[0106] First, the first identifier (e.g., “first”) needs to be added to the vocabulary of the third question and answer large model. <ga>], an index of the first identifier, a second identifier (e.g. <ba>) and an index of the second identifier. And initialize <ga>and <ba>vectors of other characters in the lexicon, as an example, can be utilized <ga>and <ba>the initialization vector, for example, the mean of the vector of other characters is taken as <ga>and <ba>initialization vector, then during the training process <ga>and <ba>The vector of the initialization vector is updated, and during the training process, the vectors of other characters in the vocabulary do not need to be updated, thereby retaining the related knowledge of other characters in the pre-training and improving the training efficiency.
[0107] Then, the third question and answer large model is trained to obtain the first question and answer large model in the following manner: generating, by the third question and answer large model, a predicted reply corresponding to a target sample based on a question sample in the target sample, the target sample including the second positive sample or the second negative sample, the predicted reply including a predicted answer corresponding to the question sample and a predicted identifier located after the predicted answer, the predicted identifier being one of the first identifier and the second identifier; and updating the third question and answer large model based on the predicted reply and the target sample.
[0108] With reference to Figure 4 The third question and answer large model can include a tokenizer, a vector encoding module, a decoding module connected to the vector encoding module, and an adapter corresponding to each original weight matrix in the decoding module, with reference to Figure 4 During the forward propagation process, the tokenizer is used to tokenize the question sample in the target sample, and the index of each word is queried from the vocabulary based on the tokenization result to obtain an index sequence corresponding to the question sample. The index sequence of the question sample is input into the vector encoding module for encoding to obtain an encoding vector of the question sample. The decoding module obtains an output based on the encoding vector, the original weight matrix, and the weight matrix of the adapter. As can be known from the above content, the output can represent a confidence vector of the model in the process of outputting each character. The character with the maximum confidence in the confidence vector is selected as the current output. All output characters are combined according to the output time, and the predicted reply can be obtained.
[0109] The predicted reply includes a predicted answer corresponding to the question sample and a predicted identifier located after the predicted answer, that is, through training, the model obtained by training can additionally and stably generate the corresponding <ga>or <ba>.
[0110] With reference to Figure 5 The step of updating the third large-scale question and answer model based on the predicted reply and the target sample can specifically include: determining a loss value based on the predicted reply and an answer sample in the target sample and an identifier added for the answer sample; freezing an original weight matrix of the decoding module and updating parameters in the third large-scale question and answer model based on the loss value, the parameters including a vector of the identifier in the vector encoding module and a weight matrix of the adapter.
[0111] The adapter can include, in sequence, a dimension reduction layer, an activation function, and a dimension increase layer. The dimension reduction layer includes a reduced rank matrix, which is used to compress the original high-dimensional feature (i.e., the encoding vector) to a low dimension. The activation function is used to enhance the expression ability of the low-dimensional feature output by the dimension reduction layer. The dimension increase layer includes an increased rank matrix, which is used to restore the feature output by the activation function to the original high dimension to be compatible with the dimension of the feature obtained by processing the encoding vector with the original weight matrix in the decoding module. The reduced rank matrix and the increased rank matrix are initialized before training. The reduced rank matrix is initialized using a normal distribution with a mean of 0 and a variance of 1, and the increased rank matrix is initialized to all 0. During the training process, the reduced rank matrix and the increased rank matrix are updated. Figure 6 During the forward inference process, the product of the reduced rank matrix and the increased rank matrix can be superimposed on the basis of the original weight matrix of the decoding module, and the encoding vector X is processed using the superimposed weight matrix to obtain H. H can be understood as the product of the superimposed weight matrix and the encoding vector X. H can be understood as an intermediate result of the confidence vector. By processing H, the confidence vector can be obtained.
[0112] The parameters can include the vector of the identifier in the vector encoding module and the weight matrix of the adapter. As known from the foregoing, the vector of the identifier can be the first identifier vector or the second identifier vector, and the weight matrix of the adapter includes the reduced rank matrix and the increased rank matrix.
[0113] The stopping update condition can be that the number of updates of the parameters in the third large-scale question and answer model reaches a preset number threshold, or other conditions. The present embodiment is not limited thereto. When the stopping update condition is reached, the iterative update of the third large-scale question and answer model is stopped, and the current third large-scale question and answer model can be used as the first large-scale question and answer model.
[0114] In the foregoing manner, the universal knowledge learned by the third large-scale question and answer model in the pre-training stage is retained without modifying the original weight matrices in the decoding module, effectively avoiding forgetting old knowledge when the model learns new tasks, and the way of adjusting the reduced rank matrix and the increased rank matrix in the present embodiment can improve the training efficiency compared with adjusting the original weights.
[0115] Based on the same concept, the embodiment of the present disclosure provides a large model-based question answering device, Figure 7 is a block diagram of a large model-based question answering device provided by the embodiment of the present disclosure, referring to Figure 7 The large model-based question answering device 700 comprises:
[0116] The first acquisition module 701 is configured to acquire a first question.
[0117] The first generation module 702 is configured to generate a first reply to the first question based on the first question through a first question answering large model, the first reply comprising a first answer to the first question and an evaluation result of the first answer.
[0118] The first determination module 703 is configured to determine whether the answer to the first question needs to be regenerated based on the evaluation result of the first answer in the first reply.
[0119] The second generation module 704 is configured to generate a second answer to the first question based on the first question through a second question answering large model in the case of determining that the answer to the first question needs to be regenerated, the inference performance of the second question answering large model being higher than that of the first question answering large model.
[0120] Optionally, the first generation module 702 is further configured to:
[0121] generate a first answer to the first question and a target identifier through a first question answering large model based on the first question, the target identifier comprising one of a first identifier and a second identifier, the first identifier being used to represent a positive evaluation result, and the second identifier being used to represent a negative evaluation result.
[0122] determine the evaluation result of the first answer based on a confidence vector corresponding to a last character output by the first question answering large model, the last character being the target identifier, the confidence vector comprising a confidence of the target identifier and a confidence of another identifier, the other identifier being the other one of the first identifier and the second identifier other than the target identifier.
[0123] Optionally, the large model-based question answering device 700 further comprises:
[0124] The second acquisition module is configured to acquire a first positive sample, the first positive sample comprising a question sample and an answer sample corresponding to the question sample.
[0125] The construction module is configured to construct a first negative sample based on the first positive sample.
[0126] a first construction module, configured to add a first identifier at the end of an answer sample in the first positive sample to construct a second positive sample, the first identifier being used to represent a positive evaluation result;
[0127] a second construction module, configured to add a second identifier at the end of an answer sample in the first negative sample to construct a second negative sample, the second identifier being used to represent a negative evaluation result, the second positive sample and the second negative sample being used to train a third large model to obtain the first large model.
[0128] Optionally, the construction module is further configured to construct a first negative sample based on the first positive sample according to a constraint, a sample quality of the first negative sample constructed according to the constraint being higher than a sample quality of a first negative sample not constructed according to the constraint.
[0129] Optionally, the first positive sample includes a first sub-positive sample and a second sub-positive sample, the first sub-positive sample including a first question sample and a first answer sample corresponding to the first question sample, the second sub-positive sample including a second question sample and a second answer sample corresponding to the second question sample, a difficulty of the first large model generating the first answer sample based on the first question sample being higher than a difficulty of the first large model generating the second answer sample based on the second question sample.
[0130] Optionally, the large model-based question and answer device 700 further includes:
[0131] a third generation module, configured to generate an evaluation result of each initial sample by a fourth large model based on a preset prompt word and each initial sample before adding an identifier to an answer sample, the initial sample being the first positive sample or the first negative sample;
[0132] a second determination module, configured to determine whether the initial sample is a noise sample based on the evaluation result of the initial sample, and to construct a corresponding target sample using the initial sample that is not a noise sample, the target sample including the second positive sample or the second negative sample.
[0133] Optionally, the first identifier and the second identifier are added to a word library of the third large model, and the large model-based question and answer device 700 further includes a training module, the third large model being trained by the training module to obtain the first large model, the training module including:
[0134] The generating sub-module is configured to generate, by the third large model, a predicted reply corresponding to a target sample based on a question sample in the target sample, the target sample comprising the second positive sample or the second negative sample, the predicted reply comprising a predicted answer corresponding to the question sample and a predicted identifier located after the predicted answer, the predicted identifier being one of the first identifier and the second identifier.
[0135] The updating sub-module is configured to update the third large model based on the predicted reply and the target sample.
[0136] Optionally, the third large model comprises a vector encoding module, a decoding module connected to the vector encoding module, and an adapter corresponding to each original weight matrix in the decoding module, the vector encoding module being configured to encode a question sample in the target sample to obtain an encoding vector, an output of the decoding module in a forward propagation process being obtained based on the encoding vector, the original weight matrix, and a weight matrix of the adapter, the output being used to determine the predicted reply, and the updating sub-module is further configured to:
[0137] determine a loss value based on the predicted reply and the target sample;
[0138] freeze the original weight matrix of the decoding module based on the loss value, and update a parameter in the third large model, the parameter comprising a vector of an identifier in the vector encoding module and a weight matrix of the adapter.
[0139] In the above large model-based question and answer device 700, the embodiments of each module can refer to the related embodiments described above, and will not be described here.
[0140] Based on the same concept, the embodiments of the present disclosure provide a computer readable medium having a computer program stored thereon, the computer program being executed by a processing device to implement the steps of the above large model-based question and answer method.
[0141] Based on the same concept, the embodiments of the present disclosure provide a computer program product comprising a computer program, the computer program being executed by a processor to implement the steps of the above large model-based question and answer method.
[0142] Based on the same concept, the embodiments of the present disclosure provide an electronic device comprising:
[0143] a storage device having a computer program stored thereon;
[0144] a processing device configured to execute the computer program in the storage device to implement the steps of the above large model-based question and answer method.
[0145] The following is for reference. Figure 8 This diagram illustrates a structural schematic of an electronic device 800 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0146] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0147] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0148] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.
[0149] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable storage medium or carried by a carrier wave in a baseband or as part of a carrier wave. Such a propagated computer-readable signal medium can take various forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can be used to carry or store a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.
[0150] In some embodiments, the electronic device can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications (e.g., a communications network) of any form or medium (e.g., a communications network). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.
[0151] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device and not be assembled into the electronic device.
[0152] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire a first question; generate, by a first question and answer large model, a first reply to the first question based on the first question, the first reply including a first answer to the first question and an evaluation result of the first answer; determine whether the answer to the first question needs to be regenerated based on the evaluation result of the first answer in the first reply; and in a case where it is determined that the answer to the first question needs to be regenerated, generate, by a second question and answer large model, a second answer to the first question based on the first question, an inference performance of the second question and answer large model being higher than an inference performance of the first question and answer large model.
[0153] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0154] The flow and block diagrams in the drawings show architectural, functional, and operational architectures of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0155] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of a module does not constitute a limitation on the module itself. For example, the first obtaining module can also be described as a module for obtaining a first question.
[0156] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that can be used include: Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0157] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage media can include, without limitation, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include, but are not limited to, an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0158] The above description is merely exemplary of the present disclosure and of the application of the principles thereof. The scope of the disclosure is not limited to the specific embodiments described herein, but only by the claims, and even if the disclosed embodiments contain many preferences, the disclosed scope of the disclosure is not limited to these and it encompasses any other technical solutions falling within the concepts described herein, resulting from a combination of the disclosed features or their equivalents. For example, the above-described features can be replaced by other features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0159] Moreover, while operations are depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order, and that some operations can be performed in parallel or in any suitably enabled order. Similarly, while several specific implementation details are included herein, they should not be taken as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.< / ba> < / ga> < / ba> < / ga> < / ba> < / ga> < / ba> < / ga> < / ba> < / ga> < / ba> < / ga> < / ba> < / ga> < / ga> < / ga> < / ba> < / ga>
Claims
1. A question-answering method based on a large model, characterized in that, include: Get the first question; Based on the first question, the first question-answering model generates a first response to the first question. The first response includes a first answer to the first question and an evaluation result of the first answer. The evaluation result is output synchronously when the first question-answering model outputs the first answer. Based on the evaluation result of the first answer in the first response, determine whether it is necessary to regenerate the answer to the first question; If it is determined that the answer to the first question needs to be regenerated, a second answer to the first question is generated based on the first question using a second question-answering big model. The reasoning performance of the second question-answering big model is higher than that of the first question-answering big model. The question-and-answer method also includes: Obtain a first positive sample, which includes a question sample and an answer sample corresponding to the question sample; Based on the first positive sample, a first negative sample is constructed according to constraints. The quality of the first negative sample constructed according to the constraints is higher than that of the first negative sample not constructed according to the constraints. The constraints include a first constraint and a second constraint. The first constraint is used to prevent the disruption of the correlation between the question sample and the answer sample in the first negative sample. The second constraint is used to prevent the disruption of the sentence coherence in the answer sample in the first negative sample. When constructing the first negative sample based on the second constraint, the answer sample in the first positive sample is randomly truncated from the beginning or end. The remaining content after truncation is used as the answer sample of the first negative sample. A first identifier is added to the end of the answer sample in the first positive sample to construct a second positive sample, wherein the first identifier is used to characterize the positive evaluation result; A second identifier is added to the end of the answer sample in the first negative sample to construct a second negative sample. The second identifier is used to characterize the negative evaluation result. The second positive sample and the second negative sample are used to train a third question-answering model to obtain the first question-answering model.
2. The question-and-answer method according to claim 1, characterized in that, The step of generating a first response to the first question based on the first question using a first question-answering model includes: Based on the first question, the first question-answering model generates a first answer and a target identifier. The target identifier includes one of a first identifier and a second identifier. The first identifier is used to characterize a positive evaluation result, and the second identifier is used to characterize a negative evaluation result. Based on the confidence vector corresponding to the last character output by the first question-answering model, the evaluation result of the first answer is determined. The last character is the target identifier. The confidence vector includes the confidence of the target identifier and the confidence of other identifiers. The other identifier is the other one of the first identifier and the second identifier besides the target identifier.
3. The question-and-answer method according to claim 1, characterized in that, The first positive sample includes a first sub-positive sample and a second sub-positive sample. The first sub-positive sample includes a first question sample and a first answer sample corresponding to the first question sample. The second sub-positive sample includes a second question sample and a second answer sample corresponding to the second question sample. The difficulty for the large model to generate the first answer sample based on the first question sample is higher than the difficulty for the large model to generate the second answer sample based on the second question sample.
4. The question-and-answer method according to claim 1, characterized in that, Before adding identifiers to the answer samples, the question-answering method also includes: The fourth question-answering model generates evaluation results for each initial sample based on preset prompts and initial samples, wherein the initial sample is either the first positive sample or the first negative sample. Based on the evaluation results of the initial sample, it is determined whether the initial sample is a noise sample, and the corresponding target sample is constructed using the initial sample that is a non-noise sample. The target sample includes the second positive sample or the second negative sample.
5. The question-and-answer method according to claim 1, characterized in that, The first identifier and the second identifier are added to the vocabulary of the third question-answering model, and the third question-answering model is trained in the following way to obtain the first question-answering model: The third question-answering model generates a predicted response corresponding to the target sample based on the question sample in the target sample. The target sample includes the second positive sample or the second negative sample. The predicted response includes the predicted answer corresponding to the question sample and a predicted identifier located after the predicted answer. The predicted identifier is one of the first identifier and the second identifier. The third question-answering model is updated based on the predicted response and the target sample.
6. The question-and-answer method according to claim 5, characterized in that, The third question-answering model includes a vector encoding module, a decoding module connected to the vector encoding module, and an adapter corresponding to each original weight matrix in the decoding module. The vector encoding module is used to encode question samples in the target sample to obtain an encoding vector. The output of the decoding module during the forward propagation process is obtained based on the encoding vector, the original weight matrix, and the weight matrix of the adapter. The output is used to determine the predicted response. The updating of the third question-answering model based on the predicted response and the target sample includes: Based on the predicted response and the target sample, the loss value is determined; Based on the loss value, the original weight matrix of the decoding module is frozen, and the parameters in the third question-answering model are updated. The parameters include the vector of identifiers in the vector encoding module and the weight matrix of the adapter.
7. A question-answering device based on a large model, characterized in that, include: The first acquisition module is used to acquire the first question; The first generation module is used to generate a first response to the first question based on the first question using a first question-answering model. The first response includes a first answer to the first question and an evaluation result of the first answer. The evaluation result is output synchronously when the first question-answering model outputs the first answer. The first determining module is used to determine, based on the evaluation result of the first answer in the first response, whether it is necessary to regenerate the answer to the first question; The second generation module is used to generate a second answer to the first question based on the first question using a second question-answering big model when it is determined that the answer to the first question needs to be regenerated. The reasoning performance of the second question-answering big model is higher than that of the first question-answering big model. The large model-based question-answering device also includes: The second acquisition module is used to acquire a first positive sample, which includes a question sample and an answer sample corresponding to the question sample; A construction module is used to construct a first negative sample based on the first positive sample, according to constraints. The quality of the first negative sample constructed according to the constraints is higher than that of the first negative sample not constructed according to the constraints. The constraints include a first constraint and a second constraint. The first constraint is used to prevent the disruption of the correlation between the question samples and answer samples in the first negative sample. The second constraint is used to prevent the disruption of the sentence coherence in the answer samples in the first negative sample. When constructing the first negative sample based on the second constraint, the answer samples in the first positive sample are randomly truncated from the beginning or end. The remaining content after truncation is used as the answer sample of the first negative sample. A first construction module is used to add a first identifier to the end of the answer sample in the first positive sample to construct a second positive sample, wherein the first identifier is used to characterize the positive evaluation result; The second construction module is used to add a second identifier to the end of the answer sample in the first negative sample to construct a second negative sample. The second identifier is used to characterize the negative evaluation result. The second positive sample and the second negative sample are used to train a third question-answering big model to obtain the first question-answering big model.
8. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by a processing device, the computer program implements the steps of the question-and-answer method according to any one of claims 1-6.
9. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the question-answering method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the question-and-answer method according to any one of claims 1-6.
Citation Information
Patent Citations
Large model training method and device, equipment and storage medium
CN120747672A