Training data construction method and device, large language model and electronic equipment

By constructing training data of sample questions and answers containing error information, fine-tuning of the large language model is solved, and the problem that the model output answer is logically self-consistent but inconsistent with the facts is improved, and the authenticity and accuracy of the model are improved.

CN120069083APending Publication Date: 2025-05-30ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510222735.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Large language models are prone to hallucinations when outputting answers, which are logically self-consistent but do not match objective facts.

Method used

By constructing training data, a sample question containing error information is generated, and corresponding sample answers are generated for it, and added to the training data to fine-tune the fine-tuning model and improve its authenticity.

Benefits of technology

It effectively improves the authenticity of the model to be fine-tuned, avoids the occurrence of hallucinations, ensures that the output answers are consistent with objective facts, and avoids the introduction of new hallucinations in the fine-tuning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069083A_ABST
    Figure CN120069083A_ABST
Patent Text Reader

Abstract

The invention provides a training data construction method and device, a large language model and electronic equipment. The training data construction method comprises the steps of generating a sample problem containing error information based on pre-training data of a model to be finely tuned; and generating a corresponding sample answer for the sample question, and adding the sample question and the corresponding sample answer as a training sample to training data. The sample problem in the training sample in the training data constructed by the method contains error information, so that the authenticity of the to-be-fine-tuned model can be improved after the to-be-fine-tuned model is fine-tuned, and the illusion problem of the to-be-fine-tuned model is avoided; moreover, the sample problem in the training sample in the training data constructed by the method is constructed based on the pre-training data of the model to be finely adjusted, so that the fine adjustment of the model to be finely adjusted does not involve knowledge beyond the pre-training stage, and a new illusion problem can be prevented from being introduced after fine adjustment. Therefore, the authenticity of the to-be-fine-tuned model after fine tuning is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of artificial intelligence technology, and in particular, to a method and device for constructing training data, a large language model, and an electronic device. Background Art

[0002] With the development of artificial intelligence technology and mobile Internet technology, users encounter more and more artificial intelligence products in their lives and work, and these products bring users an efficient and convenient usage experience. For example, current large language models have very strong conversation and reasoning abilities and can output answers based on the questions input by users.

[0003] However, in related technologies, the authenticity of large language models needs to be improved, and the hallucination problem is likely to occur, that is, the output answers are logically consistent but do not conform to objective facts. Summary of the Invention

[0004] In view of this, one or more embodiments of this specification provide a method and device for constructing training data, a large language model, and an electronic device.

[0005] To achieve the above object, one or more embodiments of this specification provide the following technical solutions:

[0006] According to a first aspect of one or more embodiments of this specification, a method for constructing training data is proposed. The training data is used for fine-tuning a model to be fine-tuned, and the method includes:

[0007] Generating a sample question containing error information based on the pre-training data of the model to be fine-tuned;

[0008] Generating a corresponding sample answer for the sample question, and adding the sample question and the corresponding sample answer as a training sample to the training data.

[0009] In a possible embodiment of this specification, the generating a sample question containing error information based on the pre-training data of the model to be fine-tuned includes:

[0010] Generating an original question based on the pre-training data of the model to be fine-tuned, and inputting the original question into the model to be fine-tuned to obtain the original answer output by the model to be fine-tuned;

[0011] Determining whether the original answer is correct based on the pre-training data;

[0012] If the original answer is correct, adding error information to the original question to obtain a sample question.

[0013] In a possible embodiment of this specification, generating the original question based on the pre-training data of the model to be fine-tuned includes:

[0014] Inputting the pre-training data into a trained auxiliary model outside the model to be trained, and instructing the auxiliary model to generate the original question based on the pre-training data of the model to be fine-tuned;

[0015] Determining whether the original answer is correct based on the pre-training data includes:

[0016] Inputting the pre-training data, the original question, and the original answer into a trained auxiliary model outside the model to be trained, and instructing the auxiliary model to determine whether the original answer is correct;

[0017] Adding error information to the original question to obtain a sample question includes:

[0018] Inputting the original question and the pre-training data into a trained auxiliary model outside the model to be trained, and instructing the auxiliary model to add error information to the original question to obtain a sample question.

[0019] In a possible embodiment of this specification, generating a corresponding sample answer for the sample question includes:

[0020] Inputting the sample question into the model to be trained to obtain a predicted answer output by the model to be trained;

[0021] Determining whether the predicted answer is correct based on the error information in the sample question and the pre-training data;

[0022] If the predicted answer is incorrect, generating a corresponding sample answer for the sample question.

[0023] In a possible embodiment of this specification, determining whether the predicted answer is correct based on the error information in the sample question and the pre-training data includes:

[0024] Inputting the pre-training data, the sample question, the error information in the sample question, and the predicted answer into a trained auxiliary model outside the model to be trained, and instructing the auxiliary model to determine whether the predicted answer is correct;

[0025] Generating a corresponding sample answer for the sample question includes:

[0026] Inputting the pre-training data, the sample question, and the error information in the sample question into a trained auxiliary model outside the model to be trained, and instructing the auxiliary model to generate a corresponding sample answer for the sample question.

[0027] In a possible embodiment of this specification, if the predicted answer is incorrect, generating a corresponding sample answer for the sample question includes:

[0028] If the predicted answer output by the model to be fine-tuned for the sample question or the prompt information is incorrect, generating prompt information based on the predicted answer, the error information in the sample question, and the pre-trained data;

[0029] Inputting the prompt information into the model to be fine-tuned to obtain the predicted answer output by the model to be fine-tuned, and determining whether the predicted answer is correct based on the error information in the sample question and the pre-trained data;

[0030] If the predicted answer output by the model to be fine-tuned is correct, taking the record composed of the predicted answer output each time and the prompt information input each time as the sample answer for the sample question.

[0031] In a possible embodiment of this specification, generating prompt information based on the predicted answer, the error information in the sample question, and the pre-trained data includes:

[0032] Inputting the pre-trained data, the sample question, the error information in the sample question, and the predicted answer into a trained auxiliary model outside the model to be trained, and instructing the auxiliary model to generate prompt information;

[0033] Determining whether the predicted answer is correct based on the error information in the sample question and the pre-trained data includes:

[0034] Inputting the pre-trained data, the sample question, the error information in the sample question, and the predicted answer into a trained auxiliary model outside the model to be trained, and instructing the auxiliary model to determine whether the predicted answer is correct.

[0035] According to the second aspect of one or more embodiments of this specification, a model fine-tuning method is proposed, and the method includes:

[0036] Fine-tuning the model to be fine-tuned that has completed pre-training based on the fine-tuning data until convergence, where the fine-tuning data is the training data constructed based on the training data construction method described in any embodiment of the first aspect.

[0037] According to the third aspect of one or more embodiments of this specification, a training data construction device is proposed, where the training data is used for fine-tuning the model to be fine-tuned, and the device includes:

[0038] A problem construction module, configured to generate sample problems containing error information based on the pre-training data of the to-be-fine-tuned model;

[0039] An answer construction module, configured to generate corresponding sample answers for the sample problems, and add the sample problems and the corresponding sample answers as a training sample to the training data.

[0040] According to a fourth aspect of one or more embodiments of the present specification, a model fine-tuning device is provided. The device includes a fine-tuning module, configured to:

[0041] Fine-tune the pre-trained to-be-fine-tuned model based on the fine-tuning data until convergence, where the fine-tuning data is the training data constructed based on the training data construction method described in any one of the embodiments in the first aspect.

[0042] According to a fifth aspect of one or more embodiments of the present specification, a large language model is provided. The model is trained to convergence based on the model fine-tuning method described in the second aspect. The model is configured to output a prompt message and an answer for a problem containing error information, where the prompt message is used to prompt the error information in the problem, and the answer is the answer generated after correcting the error information in the problem.

[0043] According to a sixth aspect of one or more embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.

[0044] According to a seventh aspect of one or more embodiments of the present specification, an electronic device is provided, including:

[0045] A processor;

[0046] A memory for storing executable instructions of the processor;

[0047] Wherein, the processor runs the executable instructions to implement the method described in the first aspect or the second aspect.

[0048] According to an eighth aspect of one or more embodiments of the present specification, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.

[0049] The technical solutions provided by the embodiments of the present specification may include the following beneficial effects:

[0050] The training data construction method provided by the embodiments of this specification first generates sample questions containing error information based on the pre-training data of the model to be fine-tuned; then generates corresponding sample answers for the sample questions, and adds the sample questions and the corresponding sample answers as a training sample to the training data. The sample questions in the training samples of the training data constructed by this method contain error information, that is, information that does not conform to objective facts. Therefore, after fine-tuning the model to be fine-tuned based on this, the authenticity of the model to be fine-tuned can be improved, and the hallucination problem of the model to be fine-tuned can be avoided, that is, it can identify the error information contained in the question and generate an answer accordingly, ensuring that the input answer conforms to objective facts; moreover, the sample questions in the training samples of the training data constructed by this method are constructed based on the pre-training data of the model to be fine-tuned. Therefore, fine-tuning the model to be fine-tuned based on this does not involve knowledge outside the pre-training stage, so new hallucination problems can be avoided after fine-tuning, thereby further improving the authenticity of the fine-tuned model to be fine-tuned. Brief Description of the Drawings

[0051] Figure 1 is a flowchart of a training data construction method provided by an exemplary embodiment.

[0052] Figure 2 is a flowchart of a sample answer generation method provided by an exemplary embodiment.

[0053] Figure 3 is a schematic diagram of a training data construction process provided by an exemplary embodiment.

[0054] Figure 4 is a schematic structural diagram of a device provided by an exemplary embodiment.

[0055] Figure 5 is a block diagram of a training data construction device provided by an exemplary embodiment. Detailed Description of the Embodiments

[0056] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0057] It should be noted that: In other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0058] First, some concepts involved in this specification are explained.

[0059] Large Language Model (LLM): A deep learning model trained on a vast amount of data, capable of generating and understanding natural language text.

[0060] LLM authenticity: Also known as the model hallucination problem, that is, the large model generates some content that seems correct but may violate the facts.

[0061] SFT (Supervised Fine-Tuning): A method of further adjusting a pre-trained model by using labeled data to improve its performance on specific tasks.

[0062] Model Instruction Attack: An attack method on a language model, which induces the model to generate incorrect or harmful outputs by inserting specific instructions into the input.

[0063] In the related art, the authenticity of large language models needs to be improved, and the hallucination problem is likely to occur, that is, the output answers are logically consistent but do not conform to objective facts. For example, when a user inputs the question "After the RAID memory was proposed by the Massachusetts Institute of Technology in 1987, which hard disk systems is it mainly applicable to?" to the large language model, the large language model outputs the answer "The RAID memory was born in 1987 and was proposed by the Massachusetts Institute of Technology; initially, the RAID technology was mainly targeted at SCSI hard disk systems. At that time, the system cost was relatively high, but with the introduction of the RAID control chip supporting IDE hard disks by HighPoint in 1993, RAID began to be applicable to relatively inexpensive IDE hard disks, which enabled more individual users to use the RAID technology. In the current personal computer market, IDE-RAID chips that mainly support specifications such as RAID0, RAID1, and RAID10 are commonly found in companies such as HighPoint and Promise." The information "the RAID memory was proposed by the Massachusetts Institute of Technology" included in the question input by the user does not conform to objective facts. In fact, the RAID memory was proposed by the University of California, Berkeley, USA. However, the large language model did not recognize the error in this information and output an answer based on this information, resulting in the error still existing in the output answer and even further introducing other incorrect information.

[0064] The hallucination problem of large language models is likely to generate incorrect or even wrong information, which has an adverse impact on the use of users.

[0065] Based on the above technical problems, at least one embodiment of this specification provides a training data construction method. The training data constructed by this method can be used for fine-tuning a model to be fine-tuned, such as supervised fine-tuning. The model to be fine-tuned is a large language model, and the model to be fine-tuned is a model that has completed pre-training based on pre-training data. Through fine-tuning, the model to be fine-tuned can be adapted to a specific field or a specific task. After the model to be fine-tuned is fine-tuned based on the training data constructed by this method, its authenticity can be significantly improved, and the hallucination problem can be avoided, that is, it can effectively identify the information in the question that does not conform to objective facts, and correct and process it.

[0066] Please refer to the appendix Figure 1 , which exemplarily shows the flow of this method, including step S101 to step S102.

[0067] In step S101, sample questions containing incorrect information are generated based on the pre-training data of the model to be fine-tuned.

[0068] Among them, the pre-training data is the training data used by the model to be fine-tuned in the pre-training stage. For example, the pre-training data can be a knowledge base. The model to be fine-tuned learns and masters all or part of the knowledge contained in the pre-training data in the pre-training stage. This step generates sample questions based on the pre-training data, which can avoid introducing new knowledge when the model to be fine-tuned is fine-tuned based on the sample questions, thereby avoiding introducing new hallucination problems.

[0069] Among them, the error information refers to the information that does not conform to the objective facts. For example, in the sample question "Introduce the life of Wu Chengen, the author of Romance of the Three Kingdoms", it contains the error information that "the author of Romance of the Three Kingdoms is Wu Chengen", which does not conform to the objective facts that "the author of Romance of the Three Kingdoms is Luo Guanzhong" and "the author of Journey to the West is Wu Chengen".

[0070] Exemplarily, sample questions containing error information can be generated in the following manner:

[0071] First, generate an original question based on the pre-training data of the model to be fine-tuned, and input the original question into the model to be fine-tuned to obtain the original answer output by the model to be fine-tuned.

[0072] Among them, the original question is a question that does not contain error information. For example, the original question can be an encyclopedia knowledge question. Such as "Who is the author of Romance of the Three Kingdoms?", "Who is the author of Journey to the West?", "Who proposed the RAID memory?", etc. In other words, the objective facts, objective knowledge, etc. contained in the pre-training data can be constructed in the form of questions as the original questions.

[0073] Optionally, input the pre-training data into a trained auxiliary model outside the model to be trained, and instruct the auxiliary model to generate original questions based on the pre-training data of the model to be fine-tuned. Among them, the auxiliary model can be other open-source large language models, etc. For example, the pre-training data can be input into the auxiliary model, and at the same time, the following prompt content "Please generate a batch of encyclopedia knowledge questions based on the objective facts and objective knowledge contained in the above data. Note that the questions cannot contain information that does not conform to the objective facts and objective knowledge, and enable the questions to examine whether the respondent masters the objective facts and objective knowledge contained in the above data and the degree of mastery; if an objective knowledge is that 'the creator of work A is B', then questions such as 'Who is the creator of work A?' and 'Who created work A?' can be generated" can be input into the auxiliary model to enable the auxiliary model to output original questions in batches.

[0074] It should be understood that the original questions can also be generated by professionals based on the pre-training data.

[0075] Preferably, the original questions can be input into the model to be fine-tuned in batches, so that the model to be fine-tuned outputs the original answers to each original question in batches.

[0076] Next, based on the pre-training data, it is determined whether the original answer is correct.

[0077] Among them, whether the original answer is correct can measure whether the model to be fine-tuned has mastered the objective facts and objective knowledge involved in the original question; that is: if the original answer is correct, the model to be fine-tuned has mastered the objective facts and objective knowledge involved in the original question; if the original answer is wrong, the model to be fine-tuned has not mastered the objective facts and objective knowledge involved in the original question.

[0078] Optionally, the pre-training data, the original question, and the original answer are input into a trained auxiliary model outside the model to be trained, and the auxiliary model is instructed to determine whether the original answer is correct. Among them, the auxiliary model can be other open-source large language models, etc. For example, the pre-training data, the original question, and the original answer can be input into the auxiliary model, and the following prompt content "Please judge whether the above answer is correct for the above question based on the objective facts and objective knowledge contained in the above data, and the judgment result can be directly output" can be input into the auxiliary model at the same time, so that the auxiliary model judges whether the original answer is correct. Preferably, the original question and the corresponding original answer can be input into the auxiliary model in batches, so that the auxiliary model judges whether the original answer is correct in batches.

[0079] It should be understood that professionals can also judge whether the original answer is correct based on the pre-training data.

[0080] Finally, if the original answer is correct, error information is added to the original question to obtain a sample question.

[0081] Among them, the correctness of the original answer can indicate that the model to be fine-tuned has mastered the objective facts and objective knowledge involved in the original question. Generating a sample question based on the objective facts and objective knowledge that have been mastered can strengthen the model to be fine-tuned's understanding of the objective facts and objective knowledge that have been mastered during the subsequent fine-tuning process.

[0082] Optionally, input the original question and the pre-trained data into a trained auxiliary model outside the model to be trained, and instruct the auxiliary model to add error information to the original question to obtain a sample question. The auxiliary model can be other open-source large language models, etc. For example, the pre-trained data and the original question can be input into the auxiliary model, and at the same time, the following prompt content "Please add information that does not conform to objective facts and objective knowledge in the above question based on the objective facts and objective knowledge contained in the above data to obtain a question containing error information. The question should be such that the answerer cannot identify the error information as much as possible; for example, replace one or more correct information, that is, information that conforms to objective facts and objective knowledge, in the question with incorrect information, or directly add incorrect information to the question, etc." can be input into the auxiliary model to make the auxiliary model output a sample question. Preferably, the original questions can be input into the auxiliary model in batches so that the auxiliary model outputs sample questions in batches.

[0083] It should be understood that professionals can also add error information to the original question based on the pre-trained data to obtain a sample question.

[0084] It should also be understood that when adding error information to the original question, the error information contained in each original question can be recorded for subsequent use, such as subsequently judging whether the predicted answer is correct, generating hint information, generating sample answers, etc.

[0085] In addition, if the original answer is incorrect, it means that the model to be fine-tuned does not master the objective facts and objective knowledge involved in the original question, so no sample question is generated for this original question.

[0086] In step S102, generate a corresponding sample answer for the sample question, and use the sample question and the corresponding sample answer as a training sample to add to the training data.

[0087] Among them, the sample answer can include hint information and answer content. The hint information is used to point out the error information contained in the sample question; the answer content is used to indicate the correct information corresponding to the error information and the answer generated by replacing the error information in the sample question with the correct information.

[0088] Exemplarily, the corresponding sample answer can be generated for the sample question in the following manner:

[0089] First, input the sample question into the model to be trained to obtain the predicted answer output by the model to be trained.

[0090] Preferably, a large number of obtained sample questions are input into the model to be trained in batches so that the model to be trained processes the sample questions in batches and outputs the predicted answers corresponding to each sample question in batches.

[0091] It should be understood that if the error information contained in each sample question is recorded when generating the sample question, this record is not input into the model to be fine-tuned along with the sample question.

[0092] Next, based on the error information in the sample question and the pre-training data, determine whether the predicted answer is correct.

[0093] Among them, whether the predicted answer is correct can measure whether the model to be fine-tuned is biased by the sample question, that is: if the predicted answer is incorrect, the model to be fine-tuned is biased by the sample question, and the model to be fine-tuned cannot identify the error information in the sample question. Although the fine-tuning model has seen the objective facts and objective knowledge involved in the sample question, its robustness to this objective fact and objective knowledge is insufficient, and it will produce hallucinations under induction and cannot guarantee authenticity; if the predicted answer is correct, the model to be fine-tuned is not biased by the sample question, and the model to be fine-tuned can identify the error information in the sample question.

[0094] Optionally, input the pre-training data, the sample question, the error information in the sample question, and the predicted answer into a trained auxiliary model outside the model to be trained, and instruct the auxiliary model to determine whether the predicted answer is correct. Among them, the auxiliary model can be other open-source large language models, etc. For example, the pre-training data, the sample question, the error information in the sample question (that is, the error information in the question recorded when generating the sample question), and the predicted answer can be input into the auxiliary model, and at the same time, the following prompt content prompt "Please judge whether the above answer points out the error information in the above question and whether the above answer is correct based on the objective facts, objective knowledge contained in the above data, and the error information in the above question, and the judgment result can be directly output" is input into the auxiliary model, or the pre-training data, the sample question, and the predicted answer can be input into the auxiliary model, and at the same time, the following prompt content prompt "Please detect the error information in the above question based on the objective facts and objective knowledge contained in the above data, that is, the information that does not conform to the objective facts and objective knowledge; and judge whether the above answer points out the error information in the above question and whether the above answer is correct based on the objective facts, objective knowledge contained in the above data, and the error information in the above question, and the judgment result can be directly output" is input into the auxiliary model to enable the auxiliary model to judge whether the predicted answer is correct. Preferably, the sample question and the corresponding predicted answer can be input into the auxiliary model in batches to enable the auxiliary model to judge whether the predicted answer is correct in batches.

[0095] It should be understood that professionals can also judge whether the predicted answer is correct based on the pre-training data.

[0096] Finally, if the predicted answer is incorrect, generate a corresponding sample answer for the sample question.

[0097] Among them, if the predicted answer output by the model to be fine-tuned for a sample question is incorrect, it can be shown that the model to be fine-tuned is misled by the sample question, and the model to be fine-tuned cannot recognize the error information in the sample question. Although the fine-tuning model has seen the objective facts and objective knowledge involved in the sample question, its robustness to this objective fact and objective knowledge is insufficient and it will have hallucinations under induction, unable to guarantee authenticity. Therefore, this sample question can be used to enhance the authenticity of the model to be fine-tuned, that is, the robustness to objective facts and objective knowledge. Therefore, when the predicted answer corresponding to a sample question is incorrect in this step, a corresponding sample answer is generated for the sample question to form a piece of sample data. If the predicted answer corresponding to a sample question is correct, no sample data is generated based on this sample question, that is, this sample question is discarded.

[0098] Optionally, input the pre-trained data, the sample question, and the error information in the sample question into a trained auxiliary model outside the model to be trained, and instruct the auxiliary model to generate a corresponding sample answer for the sample question. Among them, the auxiliary model can be other open-source large language models, etc. For example, the pre-trained data, the sample question, and the error information in the sample question (that is, the error information in the question recorded when generating the sample question) can be input into the auxiliary model, and at the same time, the following prompt content "Please generate an answer for the above question based on the objective facts, objective knowledge contained in the above data, and the error information contained in the above question. The answer should at least include the prompt information and the answer content. The prompt information is used to point out the error information contained in the sample question; the answer content is used to indicate the correct information corresponding to the error information and the answer generated by replacing the error information in the sample question with the correct information" is input into the auxiliary model, so that the auxiliary model outputs the sample answer. Preferably, the sample questions can be input into the auxiliary model in batches so that the auxiliary model outputs the sample answers in batches.

[0099] Optionally, generate the sample answer in the following manner shown in Attachment Figure 2 as follows:

[0100] Step S201: If the predicted answer output by the model to be fine-tuned for the sample question or the prompt information is incorrect, generate prompt information based on the predicted answer, the error information in the sample question, and the pre-trained data.

[0101] For example, input the pre-trained data, the sample question, the error information in the sample question, and the predicted answer into a trained auxiliary model outside the model to be trained, and instruct the auxiliary model to generate prompt information. Among them, the auxiliary model can be other open-source large language models, etc. For example, the pre-trained data, the sample question, the error information in the sample question (i.e., the error information in the question recorded when generating the sample question), and the predicted answer can be input into the auxiliary model, and at the same time, the following prompt content "prompt" is input into the auxiliary model: "Please generate prompt information for the above incorrect predicted answer based on the objective facts, objective knowledge contained in the above data, and the error information contained in the above question. This prompt information is used to prompt the answerer to discover the error information in the sample question and re-output the correct predicted answer", so that the auxiliary model outputs the prompt information.

[0102] Step S202: Input the prompt information into the model to be fine-tuned to obtain the predicted answer output by the model to be fine-tuned, and determine whether the predicted answer is correct based on the error information in the sample question and the pre-trained data.

[0103] For example, input the pre-trained data, the sample question, the error information in the sample question, and the predicted answer into a trained auxiliary model outside the model to be trained, and instruct the auxiliary model to determine whether the predicted answer is correct. Among them, the auxiliary model can be other open-source large language models, etc. For example, the pre-trained data, the sample question, the error information in the sample question (i.e., the error information in the question recorded when generating the sample question), and the predicted answer can be input into the auxiliary model, and at the same time, the following prompt content "prompt" is input into the auxiliary model: "Please judge whether the above answer points out the error information in the above question and whether the above answer is correct based on the objective facts, objective knowledge contained in the above data, and the error information in the above question, and the judgment result can be directly output", or the pre-trained data, the sample question, and the predicted answer can be input into the auxiliary model, and at the same time, the following prompt content "prompt" is input into the auxiliary model: "Please detect the error information in the above question based on the objective facts, objective knowledge contained in the above data, that is, the information that does not conform to the objective facts and objective knowledge; and judge whether the above answer points out the error information in the above question and whether the above answer is correct based on the objective facts, objective knowledge contained in the above data, and the error information in the above question, and the judgment result can be directly output", so that the auxiliary model judges whether the predicted answer is correct.

[0104] If it is judged in this step that the predicted answer is correct, then execute step S203; if it is judged in this step that the predicted answer is incorrect, then repeat steps S201 and S202.

[0105] Step S203: If the predicted answer output by the model to be fine-tuned is correct, record the predicted answer output each time and the hint information input each time as the sample answer to the sample question.

[0106] That is, use the chat record generated by the model to be fine-tuned based on the input sample question and hint information as the sample answer.

[0107] In this optional example, at least one round of hint information is given for the incorrect predicted answer output by the model to be fine-tuned until all the predicted answers output by the model to be fine-tuned are correct. The sample data constructed in this way is sample data with multi-round feedback. Fine-tuning based on this can enable the model to be fine-tuned to improve the authenticity, reflection ability, and multi-round dialogue ability of the model.

[0108] Please refer to the appendix Figure 3 which exemplarily shows the process of constructing training data for this method.

[0109] First, generate an original question based on the pre-trained data, input the original question into the model to be fine-tuned, and determine whether the model to be fine-tuned answers correctly.

[0110] Next, if the answer is correct and there is no refusal to answer, generate a sample question and input the sample question into the model to be fine-tuned to determine whether the attack on the model to be fine-tuned by the sample question is successful; if the answer is wrong, input the original question into other data sets.

[0111] Finally, if it is successful, generate a sample answer and add the sample question and the sample answer as sample data to the training data.

[0112] For example, a training sample generated based on this method is as follows:

[0113] Question: After the RAID memory was proposed by the Massachusetts Institute of Technology in 1987, which hard disk systems is it mainly applicable to?

[0114] Answer: The RAID memory was born in 1987 and was proposed by the University of California, Berkeley, USA, rather than the Massachusetts Institute of Technology. Initially, RAID technology was mainly for SCSI hard disk systems. At that time, the system cost was relatively high. However, with the release of the RAID control chip supporting IDE hard disks by HighPoint in 1993, RAID began to be applicable to relatively inexpensive IDE hard disks, which enabled more individual users to use RAID technology. In the current personal computer market, IDE-RAID chips that mainly support specifications such as RAID0, RAID1, and RAID10 are commonly found in companies such as HighPoint and Promise.

[0115] It can be seen that this method can generate specific training data for different models to be fine-tuned, so that the models to be fine-tuned can improve authenticity without introducing new hallucinations.

[0116] For the training data construction method provided in the embodiments of this specification, first, sample questions containing error information are generated based on the pre-training data of the model to be fine-tuned; then, corresponding sample answers are generated for the sample questions, and the sample questions and the corresponding sample answers are added to the training data as a training sample. The sample questions in the training samples in the training data constructed by this method contain error information, that is, information that does not conform to objective facts. Therefore, after fine-tuning the model to be fine-tuned based on this, the authenticity of the model to be fine-tuned can be improved, and the hallucination problem of the model to be fine-tuned can be avoided, that is, it can identify the error information contained in the question and generate an answer based on this, ensuring that the input answer conforms to objective facts; moreover, the sample questions in the training samples in the training data constructed by this method are constructed based on the pre-training data of the model to be fine-tuned. Therefore, fine-tuning the model to be fine-tuned based on this does not involve knowledge outside the pre-training stage, so new hallucination problems can be avoided after fine-tuning, thereby further improving the authenticity of the fine-tuned model to be fine-tuned.

[0117] At least one embodiment of this specification further provides a model fine-tuning method, which includes: fine-tuning the pre-trained model to be fine-tuned based on fine-tuning data until convergence, where the fine-tuning data is the training data constructed based on the training data construction method described in any of the above embodiments.

[0118] For the model fine-tuned based on this model fine-tuning method, the authenticity is relatively high, and it can identify the errors in the counterfactual questions, correct them and then give the correct answers; the model has strong reflection ability. For user questions, the model will first review the content in the questions, identify the errors or contradictions in them, and give its own thinking, and finally output the answers; the model can identify and resist model instruction attacks.

[0119] Figure 4 It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 4 , at the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, a memory 408, and a non-volatile memory 410. Of course, there may also be other hardware required for other tasks. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into the memory 408 and then runs it. Of course, in addition to the software implementation manner, one or more embodiments of this specification do not exclude other implementation manners, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or logical devices.

[0120] Please refer to Figure 5 , the training data construction device can be applied to devices such as Figure 4 shown to implement the technical solutions of this specification. The training data construction device is used for fine-tuning the model to be fine-tuned, and may include:

[0121] A question construction module 501, configured to generate a sample question containing error information based on the pre-training data of the model to be fine-tuned;

[0122] An answer construction module 502, configured to generate a corresponding sample answer for the sample question, and add the sample question and the corresponding sample answer as a training sample to the training data.

[0123] In a possible embodiment of this specification, the question construction module is configured to:

[0124] Generate an original question based on the pre-training data of the model to be fine-tuned, and input the original question into the model to be fine-tuned to obtain the original answer output by the model to be fine-tuned;

[0125] Determine whether the original answer is correct based on the pre-training data;

[0126] If the original answer is correct, add error information to the original question to obtain a sample question.

[0127] In a possible embodiment of this specification, when the question construction module is used to generate an original question based on the pre-training data of the model to be fine-tuned, it is configured to:

[0128] Input the pre-training data into a trained auxiliary model outside the model to be trained, and instruct the auxiliary model to generate an original question based on the pre-training data of the model to be fine-tuned;

[0129] The determination of whether the original answer is correct based on the pre-training data includes:

[0130] Input the pre-training data, the original question, and the original answer into a trained auxiliary model outside the model to be trained, and instruct the auxiliary model to determine whether the original answer is correct;

[0131] The addition of error information to the original question to obtain a sample question includes:

[0132] Input the original question and the pre-training data into a trained auxiliary model outside the model to be trained, and instruct the auxiliary model to add error information to the original question to obtain a sample question.

[0133] In a possible embodiment of this specification, when the answer construction module is used to generate a corresponding sample answer for the sample question, it is used for:

[0134] Input the sample question into the model to be trained to obtain the predicted answer output by the model to be trained;

[0135] Based on the error information in the sample question and the pre-training data, determine whether the predicted answer is correct;

[0136] If the predicted answer is incorrect, generate a corresponding sample answer for the sample question.

[0137] In a possible embodiment of this specification, when the answer construction module is used to determine whether the predicted answer is correct based on the error information in the sample question and the pre-training data, it is used for:

[0138] Input the pre-training data, the sample question, the error information in the sample question, and the predicted answer into an auxiliary model that has been trained outside the model to be trained, and instruct the auxiliary model to determine whether the predicted answer is correct;

[0139] The generation of the corresponding sample answer for the sample question includes:

[0140] Input the pre-training data, the sample question, and the error information in the sample question into an auxiliary model that has been trained outside the model to be trained, and instruct the auxiliary model to generate a corresponding sample answer for the sample question.

[0141] In a possible embodiment of this specification, when the answer construction module is used to generate a corresponding sample answer for the sample question if the predicted answer is incorrect, it is used for:

[0142] If the predicted answer output by the model to be fine-tuned for the sample question or the prompt information is incorrect, generate prompt information based on the predicted answer, the error information in the sample question, and the pre-training data;

[0143] Input the prompt information into the model to be fine-tuned to obtain the predicted answer output by the model to be fine-tuned, and determine whether the predicted answer is correct based on the error information in the sample question and the pre-training data;

[0144] If the predicted answer output by the model to be fine-tuned is correct, use the record composed of the predicted answer output each time and the prompt information input each time as the sample answer for the sample question.

[0145] In one possible embodiment of this specification, when the answer construction module is used to generate prompt information based on the predicted answer, the error information in the sample question, and the pre-trained data, it is used to:

[0146] Input the pre-trained data, the sample question, the error information in the sample question, and the predicted answer into an auxiliary model that has been trained and is outside the model to be trained, and instruct the auxiliary model to generate prompt information;

[0147] Determining whether the predicted answer is correct based on the error information in the sample question and the pre-trained data includes:

[0148] Input the pre-trained data, the sample question, the error information in the sample question, and the predicted answer into an auxiliary model that has been trained and is outside the model to be trained, and instruct the auxiliary model to determine whether the predicted answer is correct.

[0149] The model fine-tuning device can be applied to a device as shown in Figure 4 to implement the technical solutions of this specification. The model fine-tuning device can include a fine-tuning module, which is used to:

[0150] Fine-tune the model to be fine-tuned that has completed pre-training based on the fine-tuning data until convergence, where the fine-tuning data is the training data constructed based on the training data construction method described in any embodiment of the first aspect.

[0151] One or more embodiments of this specification also propose a large language model. The model is trained to convergence based on the model fine-tuning method described in the second aspect. The model is used to output prompt information and an answer for a question containing error information, where the prompt information is used to prompt the error information in the question, and the answer is the answer generated after correcting the error information in the question.

[0152] One or more embodiments of this specification also propose a computer program product, including computer programs / instructions. When the computer program / instructions are executed by a processor, the steps of the methods provided in the first aspect or the second aspect are implemented.

[0153] One or more embodiments of this specification also propose a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the methods described in the first aspect or the second aspect are implemented.

[0154] The systems, devices, modules or units illustrated in the above embodiments may be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer may be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0155] In a typical configuration, a computer includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0156] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0157] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage, quantum memory, graphene-based storage media, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0158] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0159] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0160] The terms used in one or more embodiments of this specification are for the purpose of describing particular embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "that" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0161] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.

[0162] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "upon" or "in response to determining".

[0163] The above is only the preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope protected by one or more embodiments of this specification.

Claims

1. A method for constructing training data, wherein the training data is used for fine-tuning a model to be fine-tuned, the method comprising: Generating sample questions containing error information based on the pre-training data of the model to be fine-tuned; A corresponding sample answer is generated for the sample question, and the sample question and the corresponding sample answer are added to the training data as a training sample.

2. According to the training data construction method of claim 1, the step of generating sample questions containing error information based on the pre-training data of the model to be fine-tuned comprises: Generate an original question based on the pre-training data of the model to be fine-tuned, and input the original question into the model to be fine-tuned to obtain an original answer output by the model to be fine-tuned; Determining whether the original answer is correct based on the pre-training data; If the original answer is correct, then error information is added to the original question to obtain a sample question.

3. The training data construction method according to claim 2, wherein the generating the original question based on the pre-training data of the model to be fine-tuned comprises: Inputting the pre-trained data into a trained auxiliary model other than the model to be trained, and instructing the auxiliary model to generate an original question based on the pre-trained data of the model to be fine-tuned; The determining whether the original answer is correct based on the pre-training data comprises: Inputting the pre-training data, the original question and the original answer into a trained auxiliary model other than the model to be trained, and instructing the auxiliary model to determine whether the original answer is correct; Adding error information to the original question to obtain a sample question includes: The original question and the pre-training data are input into a trained auxiliary model other than the model to be trained, and the auxiliary model is instructed to add error information to the original question to obtain a sample question.

4. The training data construction method according to claim 2, wherein generating a corresponding sample answer for the sample question comprises: Inputting the sample question into the model to be trained to obtain a predicted answer output by the model to be trained; Determining whether the predicted answer is correct based on error information in the sample question and the pre-training data; If the predicted answer is wrong, a corresponding sample answer is generated for the sample question.

5. The training data construction method according to claim 4, wherein the step of determining whether the predicted answer is correct based on the error information in the sample question and the pre-training data comprises: Inputting the pre-training data, the sample questions, the error information in the sample questions and the predicted answers into a trained auxiliary model other than the model to be trained, and instructing the auxiliary model to determine whether the predicted answers are correct; Generating a corresponding sample answer for the sample question includes: The pre-training data, the sample questions, and error information in the sample questions are input into a trained auxiliary model outside the model to be trained, and the auxiliary model is instructed to generate a corresponding sample answer for the sample question.

6. The training data construction method according to claim 4, wherein if the predicted answer is wrong, generating a corresponding sample answer for the sample question comprises: If the predicted answer output by the to-be-fine-tuned model for the sample question or prompt information is wrong, generating prompt information based on the predicted answer, the error information in the sample question and the pre-training data; Inputting the prompt information into the model to be fine-tuned to obtain a predicted answer output by the model to be fine-tuned, and determining whether the predicted answer is correct based on the error information in the sample question and the pre-training data; If the predicted answer output by the model to be fine-tuned is correct, a record consisting of the predicted answer output each time and the prompt information input each time is taken as a sample answer to the sample question.

7. The training data construction method according to claim 6, wherein the generating prompt information based on the predicted answer, the error information in the sample question and the pre-training data comprises: Inputting the pre-training data, the sample questions, the error information in the sample questions and the predicted answers into a trained auxiliary model other than the model to be trained, and instructing the auxiliary model to generate prompt information; The determining whether the predicted answer is correct based on the error information in the sample question and the pre-training data comprises: The pre-training data, the sample questions, the error information in the sample questions and the predicted answers are input into a trained auxiliary model other than the model to be trained, and the auxiliary model is instructed to determine whether the predicted answers are correct.

8. A model fine-tuning method, the method comprising: Fine-tune the pre-trained model to be fine-tuned based on the fine-tuning data until convergence, wherein the fine-tuning data is training data constructed based on the training data construction method according to any one of claims 1 to 7.

9. A training data construction device, wherein the training data is used for fine-tuning a model to be fine-tuned, the device comprising: A question construction module, used for generating sample questions containing error information based on the pre-training data of the model to be fine-tuned; The answer construction module is used to generate a corresponding sample answer for the sample question, and add the sample question and the corresponding sample answer as a training sample to the training data.

10. The training data construction device according to claim 9, wherein the question construction module is used for: Generate an original question based on the pre-training data of the model to be fine-tuned, and input the original question into the model to be fine-tuned to obtain an original answer output by the model to be fine-tuned; Determining whether the original answer is correct based on the pre-training data; If the original answer is correct, then error information is added to the original question to obtain a sample question.

11. The training data construction device according to claim 10, wherein when the answer construction module is used to generate a corresponding sample answer for the sample question, it is used to: Inputting the sample question into the model to be trained to obtain a predicted answer output by the model to be trained; Determining whether the predicted answer is correct based on error information in the sample question and the pre-training data; If the predicted answer is wrong, a corresponding sample answer is generated for the sample question.

12. The training data construction device according to claim 11, wherein the answer construction module is used to generate a corresponding sample answer for the sample question if the predicted answer is wrong, and is used to: If the predicted answer output by the to-be-fine-tuned model for the sample question or prompt information is wrong, generating prompt information based on the predicted answer, the error information in the sample question and the pre-training data; Inputting the prompt information into the model to be fine-tuned to obtain a predicted answer output by the model to be fine-tuned, and determining whether the predicted answer is correct based on the error information in the sample question and the pre-training data; If the predicted answer output by the model to be fine-tuned is correct, a record consisting of the predicted answer output each time and the prompt information input each time is taken as a sample answer to the sample question.

13. A model fine-tuning device, the device comprising a fine-tuning module, for: Fine-tune the pre-trained model to be fine-tuned based on the fine-tuning data until convergence, where: The fine-tuning data is training data constructed based on the training data construction method according to any one of claims 1 to 7.

14. A large language model, the model is trained to convergence based on the model fine-tuning method according to claim 8, the model is used to output prompt information and answers for questions containing error information, wherein: The prompt information is used to prompt the wrong information in the question, and the answer is the answer generated after correcting the wrong information in the question.

15. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.

16. An electronic device, comprising: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 8 by running the executable instructions.

17. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 8.