Large model fine tuning method and device, storage medium and electronic equipment
By decomposing complex instructions into simple instructions and fine-tuning using multi-step reasoning, the problem of difficult understanding of large language models when facing complex instructions is solved, and better fine-tuning effects are achieved.
Patent Information
- Application Number
- CN202510080708.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The existing large language model is difficult to accurately understand the intent of the instructions when facing complex instructions, resulting in poor fine-tuning when using complex instructions to fine-tune the instructions.
By decomposing complex sample fine-tuning instructions into simple decomposed instructions sequences and using multi-step inference methods, we gradually optimize the target model so that it can better understand and execute complex instructions.
Effective fine-tuning of the target model is achieved, allowing it to better receive and understand complex instructions, and improve the effect of fine-tuning of instructions.
Smart Images

Figure CN120010923A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a large model fine-tuning method, device, storage medium and electronic device. Background Art
[0002] Today, Large Language Model (LLM) is one of the most important technologies in the field of neural network research. Thanks to its powerful language understanding and generation capabilities, LLM has outstanding performance in a variety of application scenarios, including but not limited to machine translation, text summarization, question-answering systems, dialogue generation, and automatic coding assistants.
[0003] In addition to traditional training, instruction tuning is an emerging large language model tuning method that inserts specific "hints" or "instructions" into the input so that the model can adapt to specific tasks without changing the model structure. However, currently, for complex instructions that require comprehensive reasoning, such as instructions involving professional domain knowledge, with multiple conditions and multiple steps, the existing LLM cannot accurately understand the intention of the instructions. Therefore, when using complex instructions for instruction fine-tuning, it is often difficult to achieve satisfactory fine-tuning results.
[0004] Therefore, how to fine-tune the LLM using complex instructions is an urgent problem to be solved. Summary of the invention
[0005] This specification provides a large model fine-tuning method, device, storage medium and electronic device to at least partially solve the above-mentioned problems existing in the prior art.
[0006] This manual adopts the following technical solutions:
[0007] This manual provides a large model fine-tuning method, including:
[0008] Get sample fine-tuning instructions for fine-tuning the target large model;
[0009] Decomposing the sample fine-tuning instruction according to the function realized by the sample fine-tuning instruction to obtain a decomposed instruction sequence;
[0010] Determine a step answer corresponding to each decomposition instruction included in the decomposition instruction sequence to form a step answer sequence;
[0011] Determining the number of fine-tuning rounds according to the number of decomposed instructions;
[0012] For each round of fine-tuning, determining the target decomposition instruction required in the round of fine-tuning, and determining the target decomposition instruction and all decomposition instructions before the target decomposition instruction in the decomposition instruction sequence as the fine-tuning decomposition instruction;
[0013] The target large model is fine-tuned using the fine-tuning decomposition instructions and the step answers corresponding to the fine-tuning decomposition instructions until the number of fine-tuning rounds reaches the number of fine-tuning rounds.
[0014] Optionally, the fine-tuning instruction is decomposed according to the function implemented by the fine-tuning instruction to obtain a decomposed instruction sequence, specifically including:
[0015] Determine the specialized field knowledge involved in the fine-tuning instruction;
[0016] The fine-tuning instruction is decomposed according to the professional domain knowledge and the general domain knowledge to obtain a decomposed instruction sequence, wherein each decomposed instruction contained in the decomposed instruction sequence can be executed only by relying on the general domain knowledge.
[0017] Optionally, determining a step answer corresponding to each decomposition instruction included in the decomposition instruction sequence specifically includes:
[0018] For each decomposition instruction included in the decomposition instruction sequence, each decomposition instruction before the decomposition instruction in the decomposition instruction sequence is used as a preceding decomposition instruction of the decomposition instruction;
[0019] The decomposition instruction, the preceding decomposition instruction, and each step answer corresponding to the preceding decomposition instruction are input into the target large model, and the step answer corresponding to the decomposition instruction output by the target large model is obtained.
[0020] Optionally, determining the target decomposition instructions required in this round of fine-tuning specifically includes:
[0021] According to the total number of fine-tuning performed during this round of fine-tuning, the target decomposition instructions required in this round of fine-tuning are determined.
[0022] Optionally, fine-tuning the target large model by using the fine-tuning decomposition instruction and each step answer corresponding to each fine-tuning decomposition instruction, and fine-tuning the target large model specifically includes:
[0023] Inputting the fine-tuning decomposition instruction into the target large model to obtain an answer to be optimized obtained by the target large model;
[0024] The target large model is fine-tuned according to the difference between the answer to be optimized and the step answer corresponding to the fine-tuning decomposition instruction.
[0025] Optionally, fine-tuning the target large model by using the fine-tuning decomposition instruction and each step answer corresponding to each fine-tuning decomposition instruction, and fine-tuning the target large model specifically includes:
[0026] Inputting the fine-tuning decomposition instruction into a target large model that has not undergone any fine-tuning, and obtaining an output of the target large model as a low-quality answer;
[0027] The fine-tuning decomposition instruction is used as a sample question, and the step answer corresponding to the fine-tuning decomposition instruction is used as a high-quality answer;
[0028] The sample questions, the high-quality answers and the low-quality answers are used together to fine-tune the target large model.
[0029] Optionally, fine-tuning the target large model using the sample input, the high-quality answer and the low-quality answer together specifically includes:
[0030] Inputting the sample question into the target large model to obtain the answer to be optimized output by the target large model;
[0031] The target large model is fine-tuned with the optimization goal of minimizing the difference between the answer to be optimized and the high-quality answer, and maximizing the difference between the answer to be optimized and the low-quality answer.
[0032] This specification provides a large model fine-tuning device, including:
[0033] An acquisition module, used for acquiring sample fine-tuning instructions for fine-tuning a target large model;
[0034] A decomposition module, used for decomposing the sample fine-tuning instruction according to the function realized by the sample fine-tuning instruction to obtain a decomposed instruction sequence;
[0035] A construction module, used to determine a step answer corresponding to each decomposition instruction included in the decomposition instruction sequence to form a step answer sequence;
[0036] A determination module, used to determine the number of fine-tuning rounds according to the number of decomposed instructions;
[0037] A first fine-tuning module is used to determine, for each round of fine-tuning, a target decomposition instruction required in the round of fine-tuning, and determine the target decomposition instruction and all decomposition instructions before the target decomposition instruction in the decomposition instruction sequence as fine-tuning decomposition instructions;
[0038] The second fine-tuning module is used to fine-tune the target large model using the fine-tuning decomposition instructions and the step answers corresponding to the fine-tuning decomposition instructions until the number of fine-tuning rounds reaches the number of fine-tuning rounds.
[0039] Optionally, the decomposition module is specifically used to determine the professional domain knowledge involved in the fine-tuning instruction; based on the professional domain knowledge and general domain knowledge, the fine-tuning instruction is decomposed to obtain a decomposition instruction sequence, and each decomposition instruction contained in the decomposition instruction sequence can be executed only relying on the general domain knowledge.
[0040] Optionally, the construction module is specifically used to, for each decomposition instruction contained in the decomposition instruction sequence, use the decomposition instructions that precede the decomposition instruction in the decomposition instruction sequence as predecessor decomposition instructions of the decomposition instruction; input the decomposition instruction and the predecessor decomposition instruction, as well as the step answers corresponding to the predecessor decomposition instruction into the target large model, to obtain the step answers corresponding to the decomposition instruction output by the target large model.
[0041] Optionally, the first fine-tuning module is specifically configured to determine the target decomposition instructions required in this round of fine-tuning according to a total number of fine-tunings performed during this round of fine-tuning.
[0042] Optionally, the second fine-tuning module is specifically used to input the fine-tuning decomposition instruction into the target large model to obtain the answer to be optimized obtained by the target large model; and fine-tune the target large model according to the difference between the answer to be optimized and the step answer corresponding to the fine-tuning decomposition instruction.
[0043] Optionally, the second fine-tuning module is specifically used to input the fine-tuning decomposition instruction into a target large model that has not undergone any fine-tuning, and obtain the output of the target large model as a low-quality answer; use the fine-tuning decomposition instruction as a sample question, and use the step answer corresponding to the fine-tuning decomposition instruction as a high-quality answer; and use the sample question, the high-quality answer and the low-quality answer to jointly fine-tune the target large model.
[0044] Optionally, the second fine-tuning module is specifically used to input the sample question into the target large model to obtain the answer to be optimized output by the target large model; and fine-tune the target large model with the optimization goal of minimizing the difference between the answer to be optimized and the high-quality answer, and maximizing the difference between the answer to be optimized and the low-quality answer.
[0045] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned large model fine-tuning method is implemented.
[0046] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned large model fine-tuning method when executing the program.
[0047] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0048] In the large model fine-tuning method provided in the present specification, sample fine-tuning instructions for fine-tuning a target large model are obtained; the sample fine-tuning instructions are decomposed according to the functions implemented by the sample fine-tuning instructions to obtain a decomposition instruction sequence; the step answers corresponding to each decomposition instruction contained in the decomposition instruction sequence are determined to form a step answer sequence; the number of fine-tuning rounds is determined according to the number of the decomposition instructions; for each round of fine-tuning, the target decomposition instruction required in the round of fine-tuning is determined, and the target decomposition instruction and all decomposition instructions before the target decomposition instruction in the decomposition instruction sequence are determined as fine-tuning decomposition instructions; the target large model is fine-tuned using the fine-tuning decomposition instructions and the step answers corresponding to the fine-tuning decomposition instructions until the number of fine-tuning rounds reaches the number of fine-tuning rounds.
[0049] When using this method to fine-tune the target large model using complex instructions, the fine-tuning of the target large model can be completed by decomposing the complex sample fine-tuning instructions into simple decomposition instructions through progressive optimization and multi-step reasoning. This method enables the target large model to be fine-tuned by complex instructions, providing a new idea for fine-tuning the instructions of the large model while effectively strengthening the ability of the large model to receive and understand complex instructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings:
[0051] Figure 1 A schematic diagram of a large model fine-tuning method provided in this specification;
[0052] Figure 2 A schematic diagram of a specific embodiment of decomposing a sample fine-tuning instruction into decomposed instructions provided in this specification;
[0053] Figure 3 A schematic diagram of a single reasoning provided for this specification;
[0054] Figure 4 A schematic diagram of a multi-step reasoning provided for this specification;
[0055] Figure 5 A schematic diagram of a large model fine-tuning device provided for this specification;
[0056] Figure 6 A method corresponding to the Figure 1Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0057] In the traditional instruction fine-tuning method, when fine-tuning the large model, the more complex instructions are usually difficult to be understood by the large model, which leads to the failure to achieve the expected fine-tuning effect. Among them, complex instructions can usually refer to some instructions involving professional knowledge and including multiple conditions or requiring multiple steps to complete. For example, assuming that there is a set of transaction data known to the large model, if the instruction "how much is the transaction amount of each transaction" is input to the large model, then the large model can easily call the search function to complete this relatively simple instruction. However, if the instruction "please determine whether there are risky transactions in this set of transaction data" is input to the large model, then if the large model wants to execute this complex instruction, it must first be able to correctly understand what the risks are in the financial field, and be familiar with the judgment conditions of each risk, and perform multiple steps to determine whether each transaction has various risks. It can be seen that the rigid requirements for the knowledge reserve of the large model and the "illusion" problem that all large models will have at this stage make it difficult for the current instruction fine-tuning method to complete the fine-tuning of complex instructions for the large model. To solve this technical problem, this specification provides a large model fine-tuning method that can achieve complex instruction fine-tuning.
[0058] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0059] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0060] Figure 1 A flow chart of a large model fine-tuning method provided in this specification includes the following steps:
[0061] S100: Obtain sample fine-tuning instructions for fine-tuning a target large model.
[0062] In this specification, the execution subject for implementing the large model fine-tuning method may refer to a designated device such as a server set up on the business platform. For the sake of convenience, this specification only takes the server as the execution subject as an example to illustrate a large model fine-tuning method provided in this specification. At the same time, in this method, the large model refers to a large language model.
[0063] This method is mainly used in the scenario of fine-tuning a large model using complex instructions. Based on this, in this step, sample fine-tuning instructions for fine-tuning the target large model can be first obtained. The sample fine-tuning instructions are complex instructions involving professional knowledge, with multiple conditions, and requiring the execution of multiple steps.
[0064] S102: Decompose the sample fine-tuning instruction according to the function realized by the sample fine-tuning instruction to obtain a decomposed instruction sequence.
[0065] After the sample fine-tuning instructions are obtained in step S100, the obtained sample fine-tuning instructions can be decomposed in this step, and the obtained decomposed instruction sequence is used to fine-tune the target large model in subsequent steps.
[0066] Specifically, the sample fine-tuning instruction can be decomposed as follows: the original complex sample fine-tuning instruction I is decomposed into multiple steps I1, I2, ..., I n , each step obtained is a decomposition instruction. The decomposition instructions obtained are sorted from front to back according to the execution order, and the corresponding decomposition instruction sequence [I1, I2, ..., I n ]. Among them, I i The i-th step to be executed when completing the sample complex instruction is the i-th decomposition instruction in the decomposition instruction sequence. When decomposing, the decomposition is mainly based on the functions that can be achieved by the sample fine-tuning instruction. In other words, the functions that can be achieved by the decomposition instruction sequence obtained by decomposition are the same as those of the sample fine-tuning instruction.
[0067] Still taking the large model processing transaction record data as an example, Figure 2 This is a schematic diagram of a specific embodiment of decomposing a sample fine-tuning instruction into decomposed instructions provided in this specification. Figure 2 As shown, in a specific embodiment, assuming that in a transaction record with a known large model, the sample fine-tuning instruction that needs to fine-tune the target large model is "Please determine whether this transaction record has the risk of splitting money laundering." For this complex sample fine-tuning instruction I, it can be divided into the following three relatively simple steps I1, I2, I3:
[0068] I1=Find the transfer (receipt) records in the transaction records whose transaction amount is greater than the specified amount, and output the transaction sequence number;
[0069] I2 = Find the outgoing (payment) record whose transaction time is close to the transaction output in the previous step and whose transaction amount is greater than the specified amount, and output the transaction sequence number;
[0070] I3=For the transfer-in and transfer-out transaction records outputted in the first two steps, determine whether there are several pairs of transfer-in records and transfer-out records with similar transaction amounts. If so, it is determined that there is a risk of split money laundering.
[0071] In the above embodiment, for the original sample fine-tuning instruction "Please determine whether this transaction record has the risk of splitting and money laundering", it is difficult for the large model to understand and complete this complex instruction without sufficient relevant knowledge in the professional field, and thus it is impossible to achieve the fine-tuning effect. However, under the method provided in this application, the complex sample fine-tuning instruction can be split into a series of simple instructions that can be completed only by basic functions such as search and judgment, which enables most large models to easily complete the corresponding tasks.
[0072] It is not difficult to see that whether or not a large model has relevant knowledge in a professional field is a key factor in its ability to understand complex instructions. In actual applications, users' question-and-answer requirements for large models often cover knowledge in various different fields. However, for large models to systematically learn knowledge in many different professional fields, on the one hand, there is a huge implementation cost, and on the other hand, the model itself will have the problem of forgetting, so it is difficult to achieve at this stage.
[0073] Based on this, in order to make the decomposed instructions obtained by decomposing complex instructions better understood by the large model, when decomposing the sample fine-tuning instructions, this method can specifically determine the professional domain knowledge involved in the fine-tuning instructions; according to the professional domain knowledge and general domain knowledge, the fine-tuning instructions are decomposed to obtain a decomposed instruction sequence, and each decomposed instruction contained in the decomposed instruction sequence can be executed only relying on the general domain knowledge.
[0074] Through the above method, complex instructions that originally require professional domain knowledge to understand and execute can be converted into a series of simple instructions that can be understood and executed only by relying on general domain knowledge. This method can greatly reduce the possibility that the large model cannot understand the decomposed instructions, ensuring that the implementation of this method can achieve better results.
[0075] S104: Determine a step answer corresponding to each decomposition instruction included in the decomposition instruction sequence to form a step answer sequence.
[0076] After decomposing the sample fine-tuning instruction to obtain a decomposition instruction sequence for fine-tuning the target large model in step S102, in this step, the step answer corresponding to the decomposition instruction required for fine-tuning the target large model can be further obtained. i A corresponding step can be determined for each answer A i , the determined step answers are arranged in the same order as the corresponding decomposition instructions in the decomposition instruction sequence, and the step answer sequence A = [A1, A2, ..., A n ].
[0077] In this method, the step answer is the correct answer to the corresponding decomposed instruction, and can be used as a label when fine-tuning the instruction of the target large model. When determining the step answer, it can be achieved by, for example, manual labeling. However, it can be imagined that the manual labeling method requires a lot of energy. Therefore, in actual use, it is often hoped that the large model itself can be used to obtain the step answer in a machine-generated manner.
[0078] In this method, since the decomposition instructions are related to each other, when obtaining the step answer corresponding to a decomposition instruction, this decomposition instruction and all the decomposition instructions before it in the decomposition instruction sequence need to be used as the input of the large model. However, in this case, due to the long context and insufficient information grasped by the large model, the answer given by the large model through a single reasoning is often not accurate enough and cannot be used as the step answer required in this method.
[0079] To solve the above problem, the method can use a multi-step reasoning method to replace a single reasoning, thereby obtaining the step answers required in the method through a large model. Specifically, for each decomposition instruction included in the decomposition instruction sequence, each decomposition instruction before the decomposition instruction in the decomposition instruction sequence can be used as a predecessor decomposition instruction of the decomposition instruction; the decomposition instruction and the predecessor decomposition instruction, as well as each step answer corresponding to the predecessor decomposition instruction, are input into the target large model to obtain the step answer corresponding to the decomposition instruction output by the target large model.
[0080] Figure 3 A schematic diagram of a single inference provided for this specification, Figure 4 A schematic diagram of the multi-step reasoning provided for this specification, such as Figure 3 and Figure 4 As shown, C is the relevant information of the instruction (such as the transaction record mentioned in the above embodiment), I i is the i-th decomposition instruction in the decomposition instruction sequence, A i is the answer to the i-th step in the final step answer sequence, I i With A i Correspondingly, n is the total number of decomposed instructions.
[0081] In the case of multi-step reasoning, each step of reasoning can get a step answer. In the i-th step of reasoning, the input to the large model is I1~I in the decomposition instruction sequence. i , and A1~A i-1 , the answer of the large model is A i In simple terms, each pair of I i With A iAs a question-answer pair, the content input into the large model in the i-th step of reasoning is the 1st to i-1th question-answer pairs and I i , the output of the large model is A i .
[0082] The principle of the above multi-step reasoning is that when asking a question to the big model at each step, the big model only needs to give an answer based on the decomposition instructions and the results of all the previous steps. In this way of answering, the big model has sufficient information to give a high-quality and accurate answer, which can be used as the step answer in this method.
[0083] S106: Determine the number of fine-tuning rounds according to the number of decomposed instructions.
[0084] After the decomposition instruction sequence and the step answer sequence are determined in step S102 and step S104 respectively, the data for fine-tuning the target large model is sufficient and fine-tuning of the target large model can be started.
[0085] In this method, the fine-tuning method used is to complete a complete fine-tuning through multiple rounds of progressive fine-tuning. Therefore, before starting to fine-tune the target large model, it is necessary to first determine the number of fine-tuning rounds required for fine-tuning the target large model in this fine-tuning.
[0086] Based on the way of decomposing the sample fine-tuning instructions in the present method, the number of fine-tuning rounds can be determined according to the number of decomposed instructions in this step. For example, the number of fine-tuning rounds can be set to be the same as the number of decomposed instructions.
[0087] S108: For each round of fine-tuning, determine the target decomposition instruction required in the round of fine-tuning, and determine the target decomposition instruction and all decomposition instructions before the target decomposition instruction in the decomposition instruction sequence as the fine-tuning decomposition instruction.
[0088] In this method, the decomposition instructions required for each round of fine-tuning are not exactly the same, but because the decomposition instructions are different steps that jointly realize the same function and have a strong correlation in timing, the decomposition instructions used in each round are all continuous decomposition instructions in the decomposition instruction sequence. Based on this, in each round of fine-tuning, it is necessary to determine the target decomposition instruction, and determine the target decomposition instruction and all decomposition instructions before the target decomposition instruction as the fine-tuning decomposition instruction used in this round of fine-tuning.
[0089] Specifically, the target decomposition instructions determined in each round can be determined according to the sequence of the decomposition instructions in the decomposition instructions. Specifically, the target decomposition instructions required in the fine-tuning can be determined according to the total number of fine-tunings performed in the fine-tuning of the round. In simple terms, in the i-th round of fine-tuning, the i-th decomposition instruction in the decomposition instruction sequence is determined as the target instruction. In this case, the 1st to i-th decomposition instructions in the decomposition instruction sequence are the fine-tuning decomposition instructions required for the current round of fine-tuning.
[0090] S110: fine-tune the target large model using the fine-tuning decomposition instructions and the step answers corresponding to the fine-tuning decomposition instructions, until the number of fine-tuning reaches the number of fine-tuning rounds.
[0091] After determining the fine-tuning decomposition instructions required for this round of fine-tuning in step S108, the fine-tuning decomposition instructions together with the step answers corresponding to the fine-tuning decomposition instructions can be used to fine-tune the target large model. Specifically, the fine-tuning decomposition instructions can be input into the target large model to obtain the answer to be optimized obtained by the target large model; and the target large model is fine-tuned according to the difference between the answer to be optimized and the step answer corresponding to the fine-tuning decomposition instructions.
[0092] Specifically, in the i-th round of fine-tuning, the fine-tuning decomposition instructions input into the target large model are the 1st to ith decomposition instructions. At this time, the answer to be optimized given by the target large model is the overall answer to the 1st to ith decomposition instructions, and the corresponding positive label is the 1st to ith step answer corresponding to the 1st to ith decomposition instructions. It is worth mentioning that since this method ultimately wants to improve the single-shot reasoning ability of the target large model, the 1st to ith decomposition instructions input are a complete instruction connected together, and the 1st to ith step answers as positive labels are also a complete answer connected together, rather than separate small steps.
[0093] According to the difference between the answer to be optimized given by the target large model for the input fine-tuning decomposition instruction and the corresponding step answer, the target large model can adjust the parameters by itself to complete the fine-tuning of the target large model. The specific operation can be that after inputting the fine-tuning decomposition instruction and obtaining the answer to be optimized output by the target large model, the step answer is input to the target large model again, and it is informed that it is the correct answer to the fine-tuning decomposition instruction, so that the target large model can complete the optimization adjustment of the parameters by itself.
[0094] Furthermore, in the process of fine-tuning the target large model, in addition to the positive labels, negative labels can be added to further enhance the reasoning level and output capability of the target large model. Specifically, the fine-tuning decomposition instruction can be input into the target large model without any fine-tuning to obtain the output of the target large model as a low-quality answer; the fine-tuning decomposition instruction is used as a sample question, and the step answer corresponding to the fine-tuning decomposition instruction is used as a high-quality answer; the sample question, the high-quality answer and the low-quality answer are used to fine-tune the target large model.
[0095] In the process of fine-tuning the target large model, the target large model before fine-tuning will be retained to facilitate restoration when fine-tuning fails. Based on this, in each round of fine-tuning, the fine-tuning decomposition instructions can be input into the target large model before fine-tuning to obtain the answer given by the unfine-tuned target large model through a single reasoning. That is, in the i-th round of fine-tuning, the 1st to i-th decomposition instructions are input into the unfine-tuned target large model to obtain the answer given by it. At this time, the answer given by the target large model is a low-quality answer of poor quality. The step answer corresponding to the fine-tuning decomposition instruction is taken as a high-quality answer, that is, a positive label, and the low-quality answer, that is, a negative label, is used together to complete the fine-tuning of the target large model.
[0096] Under this method, when fine-tuning the target large model, specifically, the sample question can be input into the target large model to obtain the answer to be optimized output by the target large model; the target large model is fine-tuned with the optimization goal of minimizing the difference between the answer to be optimized and the high-quality answer, and maximizing the difference between the answer to be optimized and the low-quality answer.
[0097] The sample question is also the sample fine-tuning instruction. During the fine-tuning process, we hope that the answer of the target large model is as close to the positive label as possible and away from the negative label. Therefore, the difference between the answer to be optimized of the target large model and the positive label can be minimized, and the difference between the answer to be optimized and the negative label can be maximized as the optimization target, so as to achieve the parameter optimization of the target large model and complete the fine-tuning of the target large model.
[0098] It should be noted that steps S108 to S110 are the operation steps required to perform a round of fine-tuning. In the process of executing this method, multiple rounds (fine-tuning rounds) of fine-tuning are required to complete a complete fine-tuning, that is, steps S108 to S110 are repeatedly executed multiple times.
[0099] Additionally, after completing all rounds of fine-tuning, the original complex instructions, that is, the sample fine-tuning instructions, can be input into the target large model, and the target large model can be informed that the decomposed instructions are obtained based on the decomposition of the sample fine-tuning instructions, thereby further enhancing the target large model's ability to understand the sample fine-tuning instructions of this fine-tuning.
[0100] When using this method to fine-tune the target large model using complex instructions, the fine-tuning of the target large model can be completed by decomposing the complex sample fine-tuning instructions into simple decomposition instructions through progressive optimization and multi-step reasoning. This method enables the target large model to be fine-tuned by complex instructions, providing a new idea for fine-tuning the instructions of the large model while effectively strengthening the ability of the large model to receive and understand complex instructions.
[0101] The above are one or more methods for implementing large model fine-tuning in this specification. Based on the same idea, this specification also provides a corresponding large model fine-tuning device, such as Figure 5 shown.
[0102] Figure 5 A schematic diagram of a large model fine-tuning device provided for this specification includes:
[0103] An acquisition module 200 is used to acquire a sample fine-tuning instruction for fine-tuning a target large model;
[0104] A decomposition module 202, configured to decompose the sample fine-tuning instruction according to the function implemented by the sample fine-tuning instruction to obtain a decomposed instruction sequence;
[0105] A construction module 204 is used to determine a step answer corresponding to each decomposition instruction included in the decomposition instruction sequence to form a step answer sequence;
[0106] A determination module 206, configured to determine the number of fine-tuning rounds according to the number of decomposed instructions;
[0107] A first fine-tuning module 208 is used to determine, for each round of fine-tuning, a target decomposition instruction required in the round of fine-tuning, and determine the target decomposition instruction and all decomposition instructions before the target decomposition instruction in the decomposition instruction sequence as fine-tuning decomposition instructions;
[0108] The second fine-tuning module 210 is used to fine-tune the target large model using the fine-tuning decomposition instructions and the step answers corresponding to the fine-tuning decomposition instructions until the number of fine-tuning reaches the number of fine-tuning rounds.
[0109] Optionally, the decomposition module 202 is specifically used to determine the professional domain knowledge involved in the fine-tuning instruction; based on the professional domain knowledge and general domain knowledge, the fine-tuning instruction is decomposed to obtain a decomposed instruction sequence, and each decomposed instruction contained in the decomposed instruction sequence can be executed only relying on the general domain knowledge.
[0110] Optionally, the construction module 204 is specifically used to, for each decomposition instruction contained in the decomposition instruction sequence, use the decomposition instructions before the decomposition instruction in the decomposition instruction sequence as predecessor decomposition instructions of the decomposition instruction; input the decomposition instruction and the predecessor decomposition instruction, as well as the step answers corresponding to the predecessor decomposition instruction into the target large model, to obtain the step answers corresponding to the decomposition instruction output by the target large model.
[0111] Optionally, the first fine-tuning module 208 is specifically configured to determine the target decomposition instructions required in the round of fine-tuning according to the total number of fine-tunings performed during the round of fine-tuning.
[0112] Optionally, the second fine-tuning module 210 is specifically used to input the fine-tuning decomposition instruction into the target large model to obtain the answer to be optimized obtained by the target large model; and fine-tune the target large model according to the difference between the answer to be optimized and the step answer corresponding to the fine-tuning decomposition instruction.
[0113] Optionally, the second fine-tuning module 210 is specifically used to input the fine-tuning decomposition instruction into a target large model that has not undergone any fine-tuning, and obtain the output of the target large model as a low-quality answer; use the fine-tuning decomposition instruction as a sample question, and use the step answer corresponding to the fine-tuning decomposition instruction as a high-quality answer; and use the sample question, the high-quality answer and the low-quality answer to jointly fine-tune the target large model.
[0114] Optionally, the second fine-tuning module 210 is specifically used to input the sample question into the target large model to obtain the answer to be optimized output by the target large model; and fine-tune the target large model with the optimization goal of minimizing the difference between the answer to be optimized and the high-quality answer, and maximizing the difference between the answer to be optimized and the low-quality answer.
[0115] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A large model fine-tuning method is provided.
[0116] This manual also provides Figure 6 The one shown corresponds to Figure 1 A schematic diagram of the electronic device. Figure 6 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1Of course, in addition to the software implementation, this specification does not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0117] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0118] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0119] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0120] For the convenience of description, the above device is described by dividing it into various units according to its functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0121] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0122] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0123] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0125] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0126] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0127] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0128] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0129] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0130] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0131] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0132] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. A large model fine-tuning method, comprising: Get sample fine-tuning instructions for fine-tuning the target large model; Decomposing the sample fine-tuning instruction according to the function realized by the sample fine-tuning instruction to obtain a decomposed instruction sequence; Determine a step answer corresponding to each decomposition instruction included in the decomposition instruction sequence to form a step answer sequence; Determining the number of fine-tuning rounds according to the number of decomposed instructions; For each round of fine-tuning, determining the target decomposition instruction required in the round of fine-tuning, and determining the target decomposition instruction and all decomposition instructions before the target decomposition instruction in the decomposition instruction sequence as the fine-tuning decomposition instruction; The target large model is fine-tuned using the fine-tuning decomposition instructions and the step answers corresponding to the fine-tuning decomposition instructions until the number of fine-tuning rounds reaches the number of fine-tuning rounds.
2. The method according to claim 1, decomposing the fine-tuning instruction according to the function implemented by the fine-tuning instruction to obtain a decomposed instruction sequence, specifically comprising: Determine the specialized field knowledge involved in the fine-tuning instruction; The fine-tuning instruction is decomposed according to the professional domain knowledge and the general domain knowledge to obtain a decomposed instruction sequence, wherein each decomposed instruction contained in the decomposed instruction sequence can be executed only by relying on the general domain knowledge.
3. The method of claim 1, wherein determining a step answer corresponding to each decomposed instruction included in the decomposed instruction sequence comprises: For each decomposition instruction included in the decomposition instruction sequence, each decomposition instruction before the decomposition instruction in the decomposition instruction sequence is used as a preceding decomposition instruction of the decomposition instruction; The decomposition instruction, the preceding decomposition instruction, and each step answer corresponding to the preceding decomposition instruction are input into the target large model, and the step answer corresponding to the decomposition instruction output by the target large model is obtained.
4. The method according to claim 1, determining the target decomposition instructions required in the round of fine-tuning, specifically comprising: According to the total number of fine-tuning performed during this round of fine-tuning, the target decomposition instructions required in this round of fine-tuning are determined.
5. The method according to claim 1, wherein the fine-tuning decomposition instruction and each step answer corresponding to each fine-tuning decomposition instruction are used to fine-tune the target large model, and the fine-tuning of the target large model specifically includes: Inputting the fine-tuning decomposition instruction into the target large model to obtain an answer to be optimized obtained by the target large model; The target large model is fine-tuned according to the difference between the answer to be optimized and the step answer corresponding to the fine-tuning decomposition instruction.
6. The method according to claim 1, wherein the fine-tuning decomposition instruction and each step answer corresponding to each fine-tuning decomposition instruction are used to fine-tune the target large model, and the fine-tuning of the target large model specifically includes: Inputting the fine-tuning decomposition instruction into a target large model that has not undergone any fine-tuning, and obtaining an output of the target large model as a low-quality answer; The fine-tuning decomposition instruction is used as a sample question, and the step answer corresponding to the fine-tuning decomposition instruction is used as a high-quality answer; The sample questions, the high-quality answers and the low-quality answers are used together to fine-tune the target large model.
7. The method according to claim 6, wherein the sample input, the high-quality answer and the low-quality answer are used together to fine-tune the target large model, specifically comprising: Inputting the sample question into the target large model to obtain the answer to be optimized output by the target large model; The target large model is fine-tuned with the optimization goal of minimizing the difference between the answer to be optimized and the high-quality answer, and maximizing the difference between the answer to be optimized and the low-quality answer.
8. A large model fine-tuning device, comprising: An acquisition module, used for acquiring sample fine-tuning instructions for fine-tuning a target large model; A decomposition module, used for decomposing the sample fine-tuning instruction according to the function realized by the sample fine-tuning instruction to obtain a decomposed instruction sequence; A construction module, used to determine a step answer corresponding to each decomposition instruction included in the decomposition instruction sequence to form a step answer sequence; A determination module, used to determine the number of fine-tuning rounds according to the number of decomposed instructions; A first fine-tuning module is used to determine, for each round of fine-tuning, a target decomposition instruction required in the round of fine-tuning, and determine the target decomposition instruction and all decomposition instructions before the target decomposition instruction in the decomposition instruction sequence as fine-tuning decomposition instructions; The second fine-tuning module is used to fine-tune the target large model using the fine-tuning decomposition instructions and the step answers corresponding to the fine-tuning decomposition instructions until the number of fine-tuning rounds reaches the number of fine-tuning rounds.
9. In the device as described in claim 8, the decomposition module is specifically used to determine the professional domain knowledge involved in the fine-tuning instruction; based on the professional domain knowledge and general domain knowledge, the fine-tuning instruction is decomposed to obtain a decomposition instruction sequence, and each decomposition instruction contained in the decomposition instruction sequence can be executed only relying on the general domain knowledge.
10. The device as described in claim 8, wherein the construction module is specifically used to, for each decomposition instruction contained in the decomposition instruction sequence, use the decomposition instructions before the decomposition instruction in the decomposition instruction sequence as the predecessor decomposition instructions of the decomposition instruction; input the decomposition instruction and the predecessor decomposition instruction, as well as the step answers corresponding to the predecessor decomposition instruction into the target large model, to obtain the step answers corresponding to the decomposition instruction output by the target large model.
11. The device according to claim 8, wherein the first fine-tuning module is specifically used to determine the target decomposition instructions required in the round of fine-tuning according to the total number of fine-tunings performed during the round of fine-tuning.
12. In the device as described in claim 8, the second fine-tuning module is specifically used to input the fine-tuning decomposition instruction into the target large model to obtain the answer to be optimized obtained by the target large model; and fine-tune the target large model according to the difference between the answer to be optimized and the step answer corresponding to the fine-tuning decomposition instruction.
13. In the device as described in claim 8, the second fine-tuning module is specifically used to input the fine-tuning decomposition instruction into the target large model that has not undergone any fine-tuning, and obtain the output of the target large model as a low-quality answer; use the fine-tuning decomposition instruction as a sample question, and use the step answer corresponding to the fine-tuning decomposition instruction as a high-quality answer; use the sample question, the high-quality answer and the low-quality answer to jointly fine-tune the target large model.
14. In the device as described in claim 13, the second fine-tuning module is specifically used to input the sample question into the target large model to obtain the answer to be optimized output by the target large model; and fine-tune the target large model with the optimization goal of minimizing the difference between the answer to be optimized and the high-quality answer, and maximizing the difference between the answer to be optimized and the low-quality answer.
15. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 7.
16. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.