Method and device for finely adjusting large language model
By designing special fine-tuning instructions and supervising text in large language models and fine-tuning them, the problems of noise and interference in natural language processing tasks are solved, and the model's context extraction and reasoning capabilities in specific tasks are improved, achieving higher accuracy and efficiency.
Patent Information
- Application Number
- CN202510126462.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-30
AI Technical Summary
In natural language processing tasks, with the growth of external knowledge base data, the content of the document collection retrieved using RAG technology is noise or interference, which is difficult to effectively utilize, resulting in understanding bias and inefficiency in large language models when processing specific tasks.
By designing special fine-tuning instructions and supervising text, fine-tuning large language models will be improved to improve their context extraction and reasoning capabilities in specific text processing tasks. Specific methods include: determining the content of the document collection related to a specific problem, building fine-tuning instructions, obtaining answers through step-by-step reasoning and filtering, and fine-tuning the model with supervised text.
Through this method, the large language model can significantly improve its understanding and logical reasoning ability when processing natural language tasks, reduce noise impact, and improve task processing accuracy and efficiency.
Smart Images

Figure CN120068851A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the technical field of large language models and instruction fine-tuning, and in particular, to a method and device for fine-tuning large language models. Background Art
[0002] With the continuous development of artificial intelligence technology, the training and fine-tuning of large language models (LLMs) have become one of the important research directions in the field of natural language processing. As a natural language processing model based on deep learning, LLMs are widely used in various scenarios. For example, natural language text generation, text content classification, information retrieval, etc. The development of these scenario applications usually relies on the Retrieval-Augmented Generation (RAG) technology. RAG can retrieve external knowledge bases and provide document content related to text processing tasks for LLMs to enrich the knowledge scope of LLMs in specific application scenarios.
[0003] However, with the explosive growth of external knowledge base data, the sources of external information have become increasingly diverse. The content of the document collections retrieved using RAG technology sometimes contains noise or interference, and sometimes the information is so extensive that it is difficult to utilize. Therefore, there is a need for a solution that can, through technical means, improve the ability of large language models to extract and reason about context fragments closely related to specific text processing tasks, and thereby improve the accuracy and efficiency of large language models in processing natural language tasks. Summary of the Invention
[0004] One or more embodiments of this specification describe a method and device for fine-tuning large language models, which can improve the context extraction and reasoning ability of large language models, so as to better complete natural language processing tasks.
[0005] According to a first aspect, there is provided a method for fine-tuning a large language model, including:
[0006] Determine a first question and a first answer to the first question, where the first answer is obtained through a first reasoning based on the content of a document collection, and the first reasoning includes document filtering, document combination, and recursive reasoning.
[0007] Input a first fine-tuning instruction into the large language model, where the first fine-tuning instruction includes the first question, the document collection, and instructs the large language model to perform step-by-step reasoning and output a reasoning process marked with a first tag and a reasoning answer marked with a second tag.
[0008] Fine-tune the large language model according to the inference process and inference answer output by the large language model, and the supervision text, where the supervision text includes a first inference text marked with the first tag and a first answer marked with the second tag.
[0009] According to one implementation, the supervision text further includes a target corpus marked with a third tag. The target corpus contains multiple corpus segments extracted from multiple documents in the document set; the first fine-tuning instruction is further used to instruct the large language model to retrieve several context segments marked with the third tag from the document set.
[0010] In a scenario of the above implementation, the supervision text further includes source information of the target corpus marked with a fourth tag. The first fine-tuning instruction is further used to instruct the large language model to mark the source information corresponding to each of the several context segments with the fourth tag.
[0011] According to one implementation, the first answer is obtained by document filtering based on the content of the document set, including:
[0012] Based on the document set, determine target corpus segments related to the first question.
[0013] Filter the content of the target corpus segments to obtain the first answer.
[0014] According to one implementation, the first answer is obtained by document combination based on the content of the document set, including:
[0015] Based on the document set, determine several target corpus segments related to the first question.
[0016] Combine the content of the several target corpus segments to obtain the first answer.
[0017] According to one implementation, the first answer is obtained by recursive reasoning based on the content of the document set, including:
[0018] Based on the document set, determine several target corpus segments on which answering the first question depends, and there is an inference sequence relationship between the target corpus segments.
[0019] Based on the content of the several target corpus segments, recursively obtain the first answer.
[0020] According to one implementation, the fine-tuning of the large language model includes:
[0021] Determine the prediction loss according to the inference process and inference answer output by the large language model, and the supervision text.
[0022] Fine-tune the large language model according to the predicted loss.
[0023] According to one implementation, the fine-tuned large language model is used to answer target questions according to a target document set.
[0024] According to a second aspect, there is provided an apparatus for fine-tuning a large language model, including:
[0025] A determination module configured to determine a first question and a first answer to the first question, wherein the first answer is obtained through a first inference based on the content of a document set, and the first inference includes document filtering, document combination, and recursive inference;
[0026] An input module configured to input a first fine-tuning instruction into the large language model, the first fine-tuning instruction including the first question, the document set, and instructing the large language model to perform step-by-step inference and output an inference process marked with a first tag and an inference answer marked with a second tag.
[0027] A fine-tuning module configured to fine-tune the large language model according to the inference process and inference answer output by the large language model, and supervision text, where the supervision text includes first inference text marked with the first tag and the first answer marked with the second tag.
[0028] According to a third aspect, there is provided a computer program product including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0029] According to a fourth aspect, there is provided a computing device including a memory and a processor, characterized in that an executable code is stored in the memory, and when the processor executes the executable code, the method described in the first aspect is implemented.
[0030] In summary, in the method and apparatus provided in the embodiments of this specification, a method for fine-tuning a large language model is designed, which can infer and generate question-and-answer pairs according to the content of a document set; and based on the generated questions and the document set, construct fine-tuning instructions. At the same time, using the generated answers as supervision text, perform instruction fine-tuning on the large language model, gradually improving the performance of the large language model, so that the fine-tuned large language model has better understanding and logical reasoning capabilities for natural language processing tasks. Description of the Drawings
[0031] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required in the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0032] Figure 1 It is a schematic diagram of a method framework for fine-tuning a large language model disclosed in the embodiments of this specification;
[0033] Figure 2 It is a flowchart of a method for fine-tuning a large language model provided according to the embodiments of this specification;
[0034] Figure 3 It is a schematic diagram of a device for fine-tuning a large language model given according to the embodiments of this specification. Specific embodiments
[0035] The following will describe the solutions provided in the embodiments of this specification with reference to the accompanying drawings.
[0036] As mentioned above, thanks to massive text training, large language models have shown certain potential in the understanding, retrieval, analysis, and generation of natural language texts, and have been widely applied to various application scenarios.
[0037] In some application scenarios where information changes rapidly and domain knowledge is updated frequently, due to the long training and evaluation cycles of large language models, the text paradigms and domain knowledge learned by the models during the training phase usually cannot keep up with the rapid knowledge updates. This lag will cause large language models to have understanding biases, hallucinations, and answer irrelevantly when dealing with specific natural language text tasks. To address this issue, the Retrieval-Augmented Generation (RAG) technology has emerged, which can assist large language models in retrieving documents from external knowledge bases, obtaining relevant document sets required during the processing of specific natural language text tasks, thereby enriching the knowledge of large language models in related fields and effectively improving the accuracy of their processing of specific text tasks. For example, when a large language model performs the task of answering a question query, the RAG technology can guide the large language model to first retrieve a document set related to the question query from a large external knowledge base; subsequently, the large language model can use these documents to answer the input question.
[0038] Although the RAG technology can provide a collection of documents related to task processing for large language models, due to various reasons (such as the heterogeneity of information sources, the fragmentation of information content, etc.), there are often noises in the document collection retrieved by the RAG technology. That is to say, large language models cannot directly obtain the correct answers that meet the requirements of text processing tasks through these documents in the document collection. Therefore, how to enable large language models to accurately extract context fragments related to specific text tasks from the document collection and reasonably utilize these context fragments to reason and answer text tasks is a technical problem that needs to be solved currently.
[0039] The inventors propose that during the process of fine-tuning a large language model, the ability of the large language model to extract context fragments can be improved directionally by designing special fine-tuning instructions, instruction tags, and supervised texts. Instruction fine-tuning means using fine-tuning instructions in the form of natural language to adjust the parameters of the large language model. During the instruction fine-tuning process, according to the output of the large language model, it is evaluated based on the supervised text, and the model is continuously fine-tuned to expect the model to output reasoning text close to the expectation. Through instruction fine-tuning, the large language model can gradually improve its ability to understand input instructions in relevant field scenarios and better adapt to the needs of specific text processing tasks.
[0040] Based on this concept, to solve the above technical problems, the inventors propose a method for fine-tuning a large language model, which can reason and generate questions and corresponding answers according to the content of the document collection; and construct fine-tuning instructions based on the generated questions and the document collection, and at the same time use the generated answers as supervised texts to perform instruction fine-tuning on the large language model, gradually improving the performance of the large language model, so that the fine-tuned large language model has better understanding and logical reasoning abilities for natural language processing tasks.
[0041] Figure 1Disclosed is a method framework for fine-tuning a large language model in an embodiment. In this framework, according to the document content included in the document collection, for a question, corresponding answers are inferred and generated (shown as the first question and the first answer in the attached figure). Thus, the latent association between the document collection content and the first question is incorporated into the reasoning process of deriving the answer from the question. The combination of the first question and the document collection is used as a fine-tuning instruction (shown as the first fine-tuning instruction in the attached figure), and the first answer and its reasoning process are used to determine the supervised text. The large language model is fine-tuned with instructions. According to the output of the large language model, the model is adjusted based on the supervised text. Since the specific document collection content, the first question, the first answer, and the reasoning relationship between them are embedded in the process of inferring the answer corresponding to the question, the large language model fine-tuned with instructions based on this question-answer pair can learn how to effectively extract specific document collection content and use this content to reason and answer questions, thereby improving the ability and accuracy of the large language model in processing related tasks.
[0042] Following the above technical concept, in Figure 2 is shown a flowchart of a method for fine-tuning a large language model according to an embodiment of this specification. It can be understood that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities. Referring to Figure 2 , in one embodiment, the method at least includes the following steps: S201: Determine a first question and a first answer for the first question, where the first answer is obtained through a first reasoning based on the content of the document collection, and the first reasoning includes document filtering, document combination, and recursive reasoning. S203: Input the first fine-tuning instruction into the large language model, where the first fine-tuning instruction includes the first question, the document collection, and instructs the large language model to perform step-by-step reasoning and output a reasoning process marked with a first tag and a reasoning answer marked with a second tag. S205: Fine-tune the large language model according to the reasoning process and reasoning answer output by the large language model, and the supervised text, where the supervised text includes the first reasoning text marked with the first tag and the first answer marked with the second tag.
[0043] As mentioned above, the large language model LLM is trained on a large amount of natural language training data sets, learns natural language features, and shows extremely strong semantic understanding and performance capabilities for natural language. It is widely used in processing various natural language-related tasks such as language translation, text refinement and understanding, and knowledge Q&A. When applying the large language model to a specific natural language task scenario, the model can be fine-tuned to improve its understanding and processing capabilities for specific task scenarios, domain data, and external knowledge.
[0044] In one or more embodiments of this specification, taking the application of a large language model to natural language question-answering tasks as an example, the embodiments will be elaborated. However, it should be understood that the embodiments of this specification aim to provide a method for fine-tuning a large language model, and are not limited to the application scenarios of the large language model and the types of tasks processed. Any scenario related to the technical concept provided by the embodiments of the present invention can apply the method provided by the embodiments of this specification.
[0045] The specific implementation manners of the above steps will be described in detail below with reference to the accompanying drawings.
[0046] In step S201, a first question and a first answer to the first question are determined, where the first answer is obtained through a first reasoning based on the content of the document set, and the first reasoning includes document filtering, document combination, and recursive reasoning.
[0047] In the document set, there are several natural language text materials. These materials can be retrieved by the large language model using the RAG technology from an external knowledge base according to a specific question-answering task, or can be pre-collected text materials for performing instruction fine-tuning on the large language model. In addition, the sources of these materials can be several external knowledge bases, which can be online or offline databases on any topic such as common sense, technology, entertainment, etc. Moreover, the data storage formats of the text materials can be diverse. For example, they can be in the lightweight markup Markdown format, the hypertext markup HTML format, the extensible markup XML format, etc., which are data storage formats that can be read by machines. In this embodiment, no special limitations are imposed on the source, data storage format, and acquisition method of each text material in the document set.
[0048] In a specific example, the document set may include the exemplary natural language text materials shown in Table 1 below.
[0049] Table 1 Example of Document Set
[0050]
[0051] It can be understood that the "document number" shown in Table 1 is exemplified by natural number coding and is used to mark the corresponding document. However, the document numbers in the example do not represent limitations on document marking. In specific practice, other types of strings with unique identification characteristics such as UUID and document attribute coding can also be used as document numbers. The embodiments of this specification do not make specific limitations on this.
[0052] There is a vast amount of document content (natural language text) in the document collection. Although these document contents are independent in terms of storage form, there may be some logical reasoning associations among them inherently. And this association relationship is an essential dependence factor for the large language model to infer the correct answer when performing the question-answering task. That is to say, in order to ensure that the large language model can perform accurate content retrieval, reasoning and output the correct answer for the question-answering task, it is necessary to train the large language model to deeply explore and learn the potential logical reasoning associations among these document contents. In view of this, this step aims to identify and extract the key document collection contents that can support the effective reasoning of the large language model in the instruction fine-tuning stage, and provide reasoning content support for constructing the question-answer pairs for instruction fine-tuning.
[0053] It should be understood that the above-mentioned logical reasoning associations can be manifested in various types. For example, document filtering, document combination, and recursive reasoning, etc. Specific elaborations will be given in combination with examples below and will not be listed one by one here.
[0054] Next, for the first question, based on the content of the document collection, the first answer corresponding to the first question can be inferred.
[0055] According to one implementation, the first answer can be inferred through the following steps: Based on the document collection, determine the target corpus segment related to the first question. Screen the content of the target corpus segment to obtain the first answer.
[0056] In this implementation, the first answer comes from the original content in the document collection. That is to say, by locating the target corpus segment closely related to answering the first question and screening its content, the first answer can be directly obtained. In subsequent steps, based on such question-answer pairs, the denoising ability of the large language model for the document collection can be improved in the instruction fine-tuning stage, and the large language model can be trained to filter out the documents irrelevant to answering the first question from the document collection, accurately locate the target corpus segment, and extract the first answer corresponding to the first question from it.
[0057] Corresponding to the example of the document set shown in Table 1, in a specific example, the first question is: "On which day was Einstein born?". Corresponding to this question, the relevant target corpus fragment can be: "Albert Einstein was born on March 14, 1879", and this corpus fragment contains information about Einstein's birth date. Thus, by filtering the content of this target corpus fragment, the first answer corresponding to the first question can be obtained: "Einstein's birth date is March 14, 1879". It can be seen that the answer to this question uniquely exists in the content of the document numbered 1 in the document set and is not associated with any date information in the content of other documents. Therefore, in the subsequent instruction fine-tuning process, the large language model can only obtain this correct answer by accurately locating this document and screening its text content.
[0058] It should be noted that as can be seen from the example, in addition to the core answer text "March 14, 1879", the first answer can also include some open-ended text "Einstein's birth date is...", so as to construct the first answer into a complete text that is more in line with natural language habits and improve readability. These open-ended texts do not represent a limitation on the content of the first answer, and they can be flexibly replaced in different embodiments. Moreover, in the instruction fine-tuning process, the large language model can be specifically fine-tuned according to specific application requirements, so that the large language model can output open-ended text that matches a specific context or instruction. Such fine-tuning helps to ensure that when the large language model outputs an answer, it not only accurately conveys the core information but also presents the first answer in a more language-habitual form, enhancing the user experience. The embodiments of this specification do not specifically limit the text organization form of the first answer and the open-ended text therein.
[0059] According to one implementation, the first answer can be inferred through the following steps: Based on the document set, determine several target corpus fragments related to the first question. Combine the content of the several target corpus fragments to obtain the first answer.
[0060] In this implementation, the first answer is derived from the content combination of several target corpus segments. That is to say, by locating several target corpus segments closely related to answering the first question and integrating their content, the first answer can be obtained. The combination relationship existing among the target corpus segments emphasizes the cooperation and complementarity between different corpus segments. Several target corpus segments together constitute a comprehensive understanding of the first answer. In subsequent steps, based on such question-answer pairs, the depth and breadth of the extraction of the large language model from the document set according to the first question can be improved in the instruction fine-tuning stage; train the large language model to extract different perspective information points about the same question from multiple documents and fuse these information points to form a complete and comprehensive answer.
[0061] Corresponding to the example of the document set shown in Table 1, in a specific example, the first question is: "What theories has Einstein published?". Corresponding to this question, the relevant target corpus segments can be: "Special relativity and the famous equation E=MC 2 ", "A paper of epoch-making significance", and "General relativity was proposed". The content of these corpus segments is all related to answering the first question. After appropriate text integration, it can accurately and comprehensively answer the first question. Therefore, these corpus segments can be combined to obtain the first answer corresponding to the first question, for example: "What Einstein published includes: special relativity, general relativity and the famous equation E=MC 2 ".
[0062] It is not difficult to see that in the above example, the design of the first question and the first answer also includes considerations of the information filtering ability. Among the several target corpus segments extracted, although "A paper of epoch-making significance" is related to the solution of the first question like other corpus segments, when answering the first question given in the example, compared with other target corpus segments, the relevance of this corpus segment to the question answer is not obvious and can be filtered. Therefore, this corpus segment is not included in the design of the first answer. Designing question-answer pairs in this targeted manner can not only improve the information integration ability of the large language model in the instruction fine-tuning stage, but also improve its ability to screen and exclude weakly related information.
[0063] According to one implementation, the first answer is derived through the following steps: Based on the document set, determine several target corpus segments on which answering the first question depends, and there is an inference sequence relationship among the target corpus segments. Based on the content of the several target corpus segments, recursively derive the first answer.
[0064] In this implementation, several extracted target corpus segments can be arranged in a certain order of inference sequence to form a chain-like derivation structure. Each of the target corpus segments can be regarded as a node in the inference process, which not only depends on the information of the previous target corpus segment but also provides a logical inference basis for subsequent other target corpus segments. Based on these target corpus segments, recursive reasoning can be carried out to obtain the first answer. In subsequent steps, based on such question-and-answer pairs, the logical reasoning ability of the large language model can be enhanced in the instruction fine-tuning stage, so that when the large language model answers questions that require multi-step reasoning, it can more accurately locate relevant context segments and infer the correct answer.
[0065] Corresponding to the example of the document set shown in Table 1, in a specific example, the first question is: "What was Einstein's age at the time of his death?". Corresponding to this question, the relevant target corpus segments can be: "Albert Einstein was born on March 14, 1879" and "He died on April 18, 1955". These target corpus segments respectively describe Einstein's birth date and death date. According to the inference sequence relationship between these two dates, Einstein's lifespan can be derived. Therefore, the first answer corresponding to the first question is: "Einstein was 76 years old at the time of his death". It can be understood that the age in the first answer needs to be obtained by analyzing the target corpus segments, subtracting the year in the death date from the year in the birth date, which reflects the inference sequence relationship between the target corpus segments. Such a design of the first question and the first answer can focus on demonstrating the logical continuity and dependence between corpus information. During the instruction fine-tuning process, the model can be required to follow a series of logical derivation steps when answering the first question and deduce the final conclusion from one fact.
[0066] It should be noted that the above example gives two target corpus segments with logical dependence. In specific practice, the target corpus segments can be several, and the inference sequence relationship between them is not limited to the mathematical relationship given in the example, but can also be a causal relationship, a time series relationship, a conditional hypothesis relationship, etc. These are not listed one by one in the embodiments of this specification.
[0067] Next, in steps S203 and S205, the first fine-tuning instruction is input into the large language model. The first fine-tuning instruction includes the first question, the document set, and instructs the large language model to perform step-by-step reasoning and output the reasoning process marked with the first tag and the reasoning answer marked with the second tag. According to the reasoning process and reasoning answer output by the large language model, and the supervision text, the large language model is fine-tuned. The supervision text includes the first reasoning text marked with the first tag and the first answer marked with the second tag.
[0068] The constructed first fine-tuning instruction can direct the large language model to extract context fragments related to answering the first question from a document collection according to the given first question, and reason to answer the first question based on these context fragments.
[0069] The constructed supervised text can provide the correct answer paradigm and answer content for evaluating the output of the large language model. That is to say, the supervised text can not only be limited to providing one or more expected output texts, but can also include structural information of the expected output text, reasoning process information, etc., content related to the answer paradigm. The supervised text can be used as a benchmark for evaluating the output of the large language model. By comparing the differences between the output text generated by the model's reasoning and the supervised text, the performance of the large language model can be quantified, and the model parameters can be fine-tuned accordingly to improve the accuracy and logic of text generation in a specific field (or specific task).
[0070] In a specific example, the first fine-tuning instruction can be as shown in Table 2 below:
[0071] Table 2 Example of Fine-Tuning Instruction
[0072]
[0073]
[0074] The supervised text can be as shown in Table 3 below:
[0075] Table 3 Example of Supervised Text
[0076]
[0077] As shown in Table 2, the document set and the first question can form a fine-tuning instruction and be input into the large language model. Among them, the "(document set)" and "(first question)" in the example are only used as display indicators for the corresponding content in the fine-tuning instruction to facilitate the understanding of this embodiment. In actual practice, these indicator information do not need to be included in the fine-tuning instruction. It should be noted that the above examples only show the basic components of the fine-tuning instruction in the embodiments of this specification. In actual applications, the fine-tuning instruction can also include other open natural language texts to clearly indicate the tasks that the large language model needs to perform, such as: "Please answer the question according to the following materials", "The above is an introduction to... May I ask...", and so on. In addition, in the example, each document in the document set is marked with a double tag "<DOC#document number>...", but this does not represent a specific limitation on the document marking form. In other embodiments, it can also be other types of marks, such as prefix marks "Material 1:", "Material 2:", or special delimiters (such as line breaks), and so on. The embodiments of this specification do not make specific limitations on this.
[0078] In this embodiment, the supervision text may further include the first reasoning process on which the first answer depends. The supervision text not only provides the correct answer corresponding to the first question, but may also include a descriptive text showing the reasoning process between the first question and the first answer. The first reasoning process details the reasoning steps from the question to the answer. In a specific example, the supervision text further includes a first tag and a second tag; the first tag is used to mark the first reasoning process, and the second tag is used to mark the first answer. Correspondingly, the first fine-tuning instruction is also used to instruct the large language model to mark the reasoning process with the first tag and mark the reasoning answer with the second tag.
[0079] In the supervision text, in order to distinguish the reasoning process description from the final answer text, the first reasoning process can be marked with the first tag, and the first answer can be marked with the second tag.
[0080] Correspondingly, the first fine-tuning instruction can also instruct the large language model to perform step-by-step reasoning when answering the first question, and mark and output the reasoning information related to the reasoning process with the first tag, and mark and output the reasoning answer with the second tag. In this way, during the instruction fine-tuning process, the large language model not only learns how to answer questions correctly, but can also display the entire reasoning process and be evaluated according to the supervision text, gradually improving the model's reasoning ability in deep logic and enhancing its logical reasoning ability and reasoning interpretability in answering complex questions.
[0081] In a specific example, the first tag is <reason>, the second mark is <answer>The first fine-tuning instruction can be as shown in Table 4 below:
[0082] Table 4 Example of Fine-Tuning Instruction
[0083]
[0084] The supervision text can be as shown in Table 5 below:
[0085] Table 5 Example of Supervision Text
[0086]
[0087] In another specific example, the large language model can also output specific reasoning steps. The first fine-tuning instruction can be as shown in Table 6 below:
[0088] Table 6 Example of Fine-Tuning Instruction
[0089]
[0090] The supervision text can be as shown in Table 7 below:
[0091] Table 7 Example of Supervision Text
[0092]
[0093] According to one implementation, the supervision text further includes the target corpus marked with a third tag. In this implementation, in addition to providing the correct answer corresponding to the first question, the supervision text can also contain the target corpus required to obtain the correct answer. The corpus fragments contained in the target corpus are closely related to the solution of the first question and are crucial for generating the correct answer. To distinguish the target corpus from the final answer, the corpus fragments in the target corpus can be marked with a third tag. Correspondingly, the first fine-tuning instruction is also used to instruct the large language model to mark with the third tag a number of context fragments retrieved from the document set.
[0094] During the process of answering the first question, the large language model can retrieve context fragments related to answering the first question from the document set, and these context fragments can be marked and output through the third tag. In this way, during the instruction fine-tuning process, the large language model not only learns how to answer questions correctly, but also can output the context fragments that support it to obtain the reasoning result, and is evaluated according to the supervision text, gradually improving the model's ability to perform relevant context positioning and retrieval citation on the document set.
[0095] In a specific example, the third tag is <quote> ……< / quote> The first fine-tuning instruction can be as shown in Table 8 below:
[0096] Table 8 Example of Fine-Tuning Instruction
[0097]
[0098] The supervised text can be as shown in Table 9 below:
[0099] Table 9 Example of supervised text
[0100]
[0101] In addition, in this implementation, the ability of the large language model to trace the context fragments can also be considered. According to one practice, the supervised text further includes the source information of the multiple corpus fragments marked with a fourth marker. The first fine-tuning instruction is also used to instruct the large language model to mark the source information corresponding to each of the several context fragments with the fourth marker.
[0102] In order to clearly indicate the origin of each corpus fragment (corresponding to the context fragment output by the large language model), the source information of the corpus fragment can be marked with a fourth marker in the supervised text.
[0103] Correspondingly, the first fine-tuning instruction can also be used to instruct the large language model to mark and output the source information of the context fragments retrieved and related to answering the first question with the fourth marker when answering the first question. In this way, during the process of fine-tuning the large language model with instructions, the context fragment access and retrieval paths of the model can be evaluated, and the tracing ability of the large language model can be improved.
[0104] In a specific example, the fourth marker is <cite>……< / cite> . The first fine-tuning instruction can be as shown in Table 10 below:
[0105] Table 10 Example of fine-tuning instruction
[0106]
[0107] The supervised text can be as shown in Table 11 below:
[0108] Table 11 Example of supervised text
[0109]
[0110] It should be understood that the source information of the corpus fragments given in the examples is illustrated by taking the document numbers corresponding to the documents as an example. In specific applications, the source information is not limited to this, and information such as network addresses and bibliographic numbers can also be used as the source information of the corpus fragments. The embodiments of this specification do not make specific limitations in this regard.
[0111] The above embodiments disclose the composition of fine-tuning instructions and corresponding supervised texts in different scenarios. It should be noted that the above embodiments can be applied separately in practice or combined freely for application to comprehensively improve multiple capabilities of the large language model. In a specific example, the first fine-tuning instruction can be as shown in Table 12 below:
[0112] Table 12 Example of Fine-Tuning Instructions
[0113]
[0114] The supervised text can be as shown in Table 13 below:
[0115] Table 13 Example of Supervised Text
[0116]
[0117] In the above steps, according to the document set, multiple corpus fragments are extracted from it, and the first question and the corresponding first answer are inferred and generated specifically, and based on this, the fine-tuning instruction and the supervised text are constructed.
[0118] Next, the first fine-tuning instruction can be input into the large language model for step-by-step inference to output the inference process and the inference answer; and according to the inference process and the inference answer output by the large language model, and the supervised text, the large language model is fine-tuned to obtain a fine-tuned large language model.
[0119] That is to say, the prediction loss can be determined according to the inference process and the inference answer output by the large language model, and the supervised text; and the large language model is fine-tuned according to the prediction loss.
[0120] Generally speaking, in the process of instruction fine-tuning, the large language model predicts the next token according to the tokens already generated in the first fine-tuning instruction in the token order in the first fine-tuning instruction until all tokens are generated. In this text generation process, based on the generated tokens, continuous evaluation is carried out according to the supervised text, a suitable loss function can be designed, and the calculated prediction loss is used as the evaluation result. The evaluation result can be used to guide the fine-tuning of the large language model to obtain a fine-tuned large language model. There are various ways to fine-tune the large language model. For example, full parameter adjustment is performed on the large language model to make the output of the large language model gradually approach the supervised text. In a specific example, the large language model can also be fine-tuned by means of low-rank decomposition. That is, the parameter matrix of the large language model is decomposed into the product of two or more low-rank matrices, and these low-rank matrices have lower dimensions, which can effectively reduce the complexity of model adjustment. According to the evaluation result (for example, the prediction loss), the element values in the low-rank matrix are adjusted specifically to improve the performance of the large language model on specific tasks and make its output more in line with the expected inference text.
[0121] According to one implementation, the fine-tuned large language model can be used to answer a target question based on a target document set. After instruction fine-tuning according to the above steps, the large language model has a significant improvement in the logical reasoning ability for the document set. Applying this large language model to downstream task processing can more accurately and effectively extract the task-related corpus fragments in the document set, mine the potential logical relationships therein, perform reasoning, and accurately answer the target question.
[0122] The above text elaborated in detail a method for fine-tuning a large language model according to one or more embodiments. By using the above method provided in the embodiments of this specification, question-answer pairs can be inferred and generated based on the contents of multiple documents in the document set, and special tags can be used to construct fine-tuning instructions and supervised texts to perform instruction fine-tuning on the large language model. Through iterative instruction fine-tuning, the understanding and reasoning ability of the large language model for the logical association relationships in natural language texts can be continuously enhanced; at the same time, the ability of the large language model to perform relevant context retrieval in the document set for the question to be answered can be improved.
[0123] In this specification, the "first" in terms such as the first tag and the first fine-tuning instruction, and the corresponding "second", "third" (if any) in the text are only for the convenience of distinction and description, and do not have any limiting meaning.
[0124] The above content describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the embodiments, and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily have to be executed in the specific order or continuous order shown to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0125] Figure 3 FIG. is a schematic diagram of a device for fine-tuning a large language model according to an embodiment of this specification. The device 300 is deployed in a computing device, and the computing device can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities. This device embodiment corresponds to Figure 2 the method embodiment shown, and the device 300 includes:
[0126] A determination module 301, configured to determine a first question and a first answer for the first question, where the first answer is obtained through a first inference based on the content of the document set, and the first inference includes document filtering, document combination, and recursive inference.
[0127] An input module 302 is configured to input a first fine-tuning instruction into the large language model. The first fine-tuning instruction includes the first question, the document set, and instructs the large language model to perform step-by-step reasoning and output a reasoning process marked with a first tag and a reasoning answer marked with a second tag.
[0128] A fine-tuning module 303 is configured to fine-tune the large language model according to the reasoning process and reasoning answer output by the large language model, and the supervision text. The supervision text includes a first reasoning text marked with the first tag and the first answer marked with the second tag.
[0129] According to an embodiment of another aspect, this specification also provides a computer program product, including computer programs / instructions, which when executed by a processor, implement the steps of the foregoing method in combination with Figure 2 the method.
[0130] According to an embodiment of yet another aspect, this specification also provides a computing device, including a memory and a processor. It is characterized in that executable code is stored in the memory, and when the processor executes the executable code, the steps of the foregoing method in combination with Figure 2 the method are implemented.
[0131] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0132] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the embodiments of the present invention. It should be understood that the above is only the specific embodiments of the embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included within the protection scope of the present invention.< / answer> < / reason>
Claims
1. A method for fine-tuning a large language model, comprising: Determine a first question and a first answer to the first question, wherein the first answer is obtained through first reasoning based on the content of the document collection, and the first reasoning includes document filtering, document combination and recursive reasoning; Inputting a first fine-tuning instruction into the large language model, the first fine-tuning instruction includes the first question and the document set, and instructs the large language model to perform step-by-step reasoning, and output a reasoning process marked with a first tag and a reasoning answer marked with a second tag; The large language model is fine-tuned according to the reasoning process and the reasoning answer output by the large language model, as well as the supervision text, wherein the supervision text includes the first reasoning text marked with the first tag and the first answer marked with the second tag.
2. The method according to claim 1, wherein: The supervised text also includes a target corpus annotated with a third tag; the target corpus includes multiple corpus fragments extracted from multiple documents in the document collection; the first fine-tuning instruction is also used to instruct the large language model to annotate several context fragments retrieved from the document collection with the third tag.
3. The method according to claim 2, wherein: The supervised text also includes source information of the target corpus marked with a fourth marker; the first fine-tuning instruction is also used to instruct the large language model to mark the source information corresponding to each of the plurality of context segments with the fourth marker.
4. The method according to claim 1, wherein: The first answer is obtained by filtering the documents according to the content of the document collection, including: Based on the document collection, determining a target corpus segment related to the first question; The content of the target corpus segment is screened to obtain the first answer.
5. The method according to claim 1, wherein: The first answer is obtained by combining documents according to the content of the document collection, including: Based on the document collection, determining a number of target corpus segments related to the first question; The contents of the plurality of target corpus segments are combined to obtain the first answer.
6. The method according to claim 1, wherein: The first answer is obtained through recursive reasoning based on the content of the document collection, including: Based on the document set, determining a number of target corpus segments that are relied upon to answer the first question, wherein there is an inference sequence relationship between the target corpus segments; The first answer is obtained recursively based on the contents of the plurality of target corpus segments.
7. The method according to claim 1, wherein: The fine-tuning of the large language model comprises: Determining a prediction loss according to the reasoning process and the reasoning answer output by the large language model and the supervision text; The large language model is fine-tuned based on the prediction loss.
8. The method according to claim 1, wherein: The fine-tuned large language model is used to answer the target question based on the target document set.
9. A device for fine-tuning a large language model, comprising: A determination module is configured to determine a first question and a first answer to the first question, wherein the first answer is obtained through a first reasoning based on the content of a document collection, and the first reasoning includes document filtering, document combination, and recursive reasoning; An input module is configured to input a first fine-tuning instruction into a large language model, wherein the first fine-tuning instruction includes the first question and the document set, and instructs the large language model to perform step-by-step reasoning and output a reasoning process marked with a first tag and a reasoning answer marked with a second tag; The fine-tuning module is configured to fine-tune the large language model according to the reasoning process and the reasoning answer output by the large language model, and the supervision text, wherein the supervision text includes the first reasoning text marked with the first tag and the first answer marked with the second tag.
10. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.
11. A computing device comprising a memory and a processor, characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Instruction fine tuning data set generation method, electronic equipment and storage medium
CN121579077A
An instruction fine-tuning dataset generation method, an electronic device, and a storage medium
CN121579077B