Fine adjustment method and device of large language model, storage medium and electronic equipment

The DPO training method generates virtual samples and labels standard response corpus, and fine-tunes the large language model are fine-tuned, solving the problem of high complexity of fine-tuning training in the existing technology, improving the efficiency of fine-tuning training and output accuracy.

CN120373401APending Publication Date: 2025-07-25ANT ZHIXIN HANGZHOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510363817.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing large language models have challenges in following user instructions and providing factual information. Although existing fine-tuning methods such as SFT and RL have certain success, RL training is complex and costly, making it difficult to improve fine-tuning training efficiency.

Method used

Direct preference optimization DPO training method is adopted, and by generating virtual samples and labeling standard response corpus, large language models are fine-tuned and trained, simplifying reinforcement learning problems into binary classification, reducing training complexity.

Benefits of technology

The fine-tuning training process of large language models is simplified, the efficiency of fine-tuning training and the accuracy of output results are improved, and the training complexity and cost are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373401A_ABST
    Figure CN120373401A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a fine tuning method of a large language model, which improves the fine tuning training in the traditional reinforcement learning stage into DPO training, and can simplify the complex reinforcement learning problem of how to select the response with the maximum reward into the relatively simple dichotomy problem of whether the reward of the response is high or low. Therefore, the training complexity of the large language model in the reinforcement learning stage is reduced, the fine tuning training of the large language model is simplified, and the efficiency of the fine tuning training is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular, to a method, apparatus, storage medium, and electronic device for fine-tuning large language models. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have made significant progress in the field of natural language processing. The powerful reasoning ability of large language models enables them to play an important role in complex language tasks such as text generation, translation, summarization, and question answering. However, although LLMs perform well in understanding and generating natural language, they still face challenges in following user instructions and providing factual information.

[0003] In real-world applications, users expect LLMs to accurately execute instructions and provide true and reliable information. For example, in intelligent assistant or customer service scenarios, the questions raised by users require the model to provide accurate answers. However, existing LLMs may produce "hallucinations" when processing these tasks, that is, generate information that does not conform to reality or is completely fabricated. This phenomenon will reduce the reliability of the model and the trust of users.

[0004] To solve this problem, two main methods are commonly used to fine-tune LLMs: supervised fine-tuning (SFT) and reinforcement learning (RL). SFT trains the model using a labeled dataset to make it better follow instructions. RL guides the model to generate better outputs through reward signals. Although these methods have achieved some success in improving model performance, the complexity of RL training is relatively high, and the training cost is relatively large.

[0005] Therefore, how to improve the efficiency of fine-tuning training for large language models has become an urgent problem to be solved. Summary of the Invention

[0006] Embodiments of this specification provide a method, apparatus, storage medium, and electronic device for fine-tuning large language models to partially solve the problems existing in the above-mentioned prior art.

[0007] Embodiments of this specification adopt the following technical solutions:

[0008] A method for fine-tuning a large language model provided in this specification, the method includes:

[0009] Obtain an advanced large language model that has undergone supervised fine-tuning training (SFT);

[0010] Generate virtual training samples using the advanced large language model as the first virtual samples; wherein each first virtual sample includes a virtual instruction corpus and at least two virtual response corpora corresponding to the virtual instruction corpus.

[0011] For each first virtual sample, select a standard response corpus from the at least two virtual response corpora included in the first virtual sample and perform annotation.

[0012] Use each of the annotated first virtual samples to perform direct preference optimization (DPO) fine-tuning training on the advanced large language model.

[0013] A fine-tuning device for a large language model provided in this specification, the device includes:

[0014] An acquisition module for acquiring an advanced large language model that has undergone supervised fine-tuning training (SFT).

[0015] A sample synthesis module for generating virtual training samples using the advanced large language model as the first virtual samples; wherein each first virtual sample includes a virtual instruction corpus and at least two virtual response corpora corresponding to the virtual instruction corpus.

[0016] An annotation module for, for each first virtual sample, selecting a standard response corpus from the at least two virtual response corpora included in the first virtual sample and performing annotation.

[0017] A fine-tuning module for using each of the annotated first virtual samples to perform direct preference optimization (DPO) fine-tuning training on the advanced large language model.

[0018] A computer-readable storage medium of an electronic device provided in this specification, the storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned fine-tuning method of the large language model.

[0019] An electronic device provided in this specification, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the above-mentioned fine-tuning method of the large language model.

[0020] The above-mentioned at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:

[0021] The embodiments of this specification disclose a fine-tuning method for large language models, which improves the fine-tuning training in the traditional reinforcement learning stage to DPO training. It can simplify the complex reinforcement learning problem of "how to select the response with the maximum reward" into a relatively simple binary classification problem of "is the reward of this response high or low", thereby reducing the training complexity of the large language model in the reinforcement learning stage, simplifying the fine-tuning training of the large language model, and improving the efficiency of fine-tuning training. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings described herein are used to provide a further understanding of this specification and form a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation to this specification. In the drawings:

[0023] Figure 1 is a flowchart of the fine-tuning method for the large language model provided by the embodiments of this specification;

[0024] Figure 2 is a detailed flowchart of the fine-tuning method for the large language model provided by the embodiments of this specification;

[0025] Figure 3 is a schematic diagram of a fine-tuning device for a large language model provided by the embodiments of this specification;

[0026] Figure 4 is a schematic diagram of the structure of an electronic device provided by the embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] LLM fine-tuning methods in the prior art, such as SFT and RL. SFT fine-tuning training relies on a large amount of high-quality labeled data, which is difficult to obtain in many cases. Moreover, SFT focuses on improving the model's ability to follow instructions while ignoring the factual accuracy of the generated text. For example, when a user needs the large language model to describe the uses of a mobile phone and inputs the instruction "Please describe the uses of a mobile phone" into the large language model, the response of the large language model is "One can make calls with others through a mobile phone". It can be seen that the output result given by the large language model is not perfect. Although it follows the user's instruction "Describe the uses of a mobile phone", due to reasons such as training samples and the training process, the uses of the mobile phone described by the large language model are obviously too simple and plain (the uses of a mobile phone also include watching videos, playing games, surfing the Internet, etc.), which is the so-called low factual accuracy. And RL fine-tuning training requires designing and training an explicit reward model, with high training complexity, an unstable training process, and RL being sensitive to hyperparameter adjustment and having a high computational cost.

[0028] Based on this, the embodiments of this specification provide a fine-tuning method for large language models, aiming to improve the traditional fine-tuning training in the RL stage to DPO fine-tuning training, so as to reduce the complexity of the fine-tuning training of large language models and improve the fine-tuning training efficiency. DPO is also a reinforcement learning method for training large language models. It optimizes the model through human preference data without using complex reinforcement learning algorithms. Its core idea is to directly use preference data to adjust model parameters, avoiding the fitting of explicit reward models and complex reinforcement learning optimization processes.

[0029] To make the purpose, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0030] The following will detail the technical solutions provided by each embodiment of this specification in conjunction with the drawings.

[0031] Figure 1 The following is a flowchart of the fine-tuning method for the large language model provided by the embodiments of this specification, including the following steps:

[0032] S100: Obtain an advanced large language model that has undergone supervised fine-tuning training (SFT).

[0033] In the embodiments of this specification, the device for fine-tuning the large language model can be a server or a distributed system composed of multiple servers, or any electronic device that can perform model training. The following will take a server as an example for illustration.

[0034] When the server performs fine-tuning training on the large language model, it needs to first obtain a pre-trained large language model. Specifically, an open-source large language model with good language understanding and generation capabilities, such as GPT4, can be selected. Then, perform SFT fine-tuning training on the pre-trained large language model to obtain an advanced large language model. How to specifically perform SFT fine-tuning training on the pre-trained large language model will not be described here for now and will be elaborated later in this text.

[0035] S102: Use the advanced large language model to generate virtual training samples as the first virtual samples.

[0036] Among them, each first virtual sample includes a virtual instruction corpus and at least two virtual response corpora corresponding to the virtual instruction corpus.

[0037] In the embodiments of this specification, the so-called virtual training samples refer to those that are not the instruction corpus actually input by the user during the execution of a certain service and the response corpus actually obtained, but are generated by the advanced large language model.

[0038] Specifically, the server can first obtain the prompt information in the specified type of service, input the prompt information into the advanced large language model, and based on the prompt information, use the advanced large language model to generate virtual training samples. That is, taking the prompt information as prior knowledge, the advanced large language model extracts knowledge points from the prompt information, and based on these knowledge points, the advanced large language model generates "self-questioning and self-answering" virtual training samples.

[0039] Among them, the first virtual sample generated by the advanced large language model contains a virtual instruction corpus and more than two corresponding virtual response corpora. That is, in the process of the advanced large language model generating "self-questioning and self-answering" virtual training samples according to the above knowledge points, it can, based on these knowledge points, infer the instruction corpus that the user may input during the execution of the above specified type of service, and can also, based on the possible input instruction corpus and these knowledge points, generate at least two optional response corpora required for the possible input instruction corpus. Then, the above possible input instruction corpus inferred by the advanced large language model and the at least two optional response corpora required for the possible input instruction corpus can be used as virtual training samples, that is, the first virtual sample.

[0040] For example, the server can directly input the product description of the mobile phone as a product on the e-commerce platform into the advanced large language model as the prompt information. Then, based on the trained understanding and reasoning abilities, the advanced large language model can extract various functions such as the mobile phone can take pictures, access the Internet, play games, etc. as knowledge points. Then, the advanced large language model can generate the instruction corpus that the user may input, "Describe the uses of the mobile phone", and four corresponding optional response corpora, "The mobile phone can make calls", "The mobile phone can take pictures", "The mobile phone can access the Internet", "The mobile phone can play games", etc. as the first virtual sample.

[0041] S104: For each first virtual sample, select a standard response corpus from the at least two virtual response corpora included in the first virtual sample and perform annotation.

[0042] In the embodiments of this specification, since each first virtual sample contains a virtual instruction corpus and more than two virtual response corpora, in step S104, for each first virtual sample, one or several preferred response corpora can be selected from the at least two virtual response corpora included in the first virtual sample as the standard response corpus and perform annotation.

[0043] Continuing with the above example, the first virtual sample includes a virtual instruction corpus "Describe the purpose of a mobile phone" and four corresponding optional virtual response corpora "The mobile phone can make calls", "The mobile phone can take pictures", "The mobile phone can surf the Internet", and "The mobile phone can play games". If you want to fine-tune the trained large language model to be more inclined to output "The mobile phone can play games" when encountering questions like the virtual instruction corpus "Describe the purpose of a mobile phone", the virtual response corpus "The mobile phone can play games" in the first virtual sample can be marked as the standard response corpus.

[0044] Among them, the method of annotating the standard response corpus can be manual annotation, or it can be automatic annotation through other machine learning models, and the embodiments of this specification do not limit this.

[0045] S106: Using the labeled first virtual samples to perform direct preference optimization (DPO) fine-tuning training on the advanced large language model.

[0046] When the advanced large language model is subjected to DPO fine-tuning training using the annotated first virtual samples, a preset reward function can be used to perform DPO fine-tuning training on the advanced large language model, wherein the reward function is used to determine a reward value based on the instruction corpus input into the large language model and the response corpus output by the large language model. The reward function is set to: after the virtual instruction corpus in the first virtual sample is input into the large language model, the reward value corresponding to the response corpus output by the large language model being the standard response corpus is greater than the reward value corresponding to the output response corpus being not the standard response corpus. Thus, the reward value determined by the reward function can be maximized as the goal of fine-tuning training to adjust the model parameters of the advanced large language model.

[0047] In the embodiments of this specification, the large language model that has been fine-tuned by the DPO can be called the final large language model. After the final large language model is obtained, the final large language model can be used to execute the above-mentioned specified type of business. The specified type of business is the business that the large language model after fine-tuning training needs to provide, including all language-related businesses that require text generation, such as chatbots, intelligent customer service, translation, question and answer, summary, etc.

[0048] From the above method, it can be seen that the DPO fine-tuning training method can simplify the complex reinforcement learning problem of "how to select the response with the largest reward" into a relatively simple binary classification problem of "is the reward of this response high or low", thereby reducing the training complexity of the large language model in the reinforcement learning stage, simplifying the fine-tuning training of the large language model, and improving the efficiency of fine-tuning training.

[0049] exist Figure 1In step S100 shown above, the server can directly use the large language model fine-tuned by SFT in the prior art as the obtained advanced large language model. However, the factual accuracy of the large language model fine-tuned by SFT in the prior art is disclosed. Therefore, in order to further improve the accuracy of the output results of the fine-tuned large language model, the embodiments of this specification can perform two SFT fine-tuning trainings on the pre-trained large language model to improve the accuracy of the output results of the large prediction model, as Figure 2 shown.

[0050] Figure 2 The detailed flowchart of the fine-tuning method for the large language model provided by the embodiments of this specification includes the following steps:

[0051] S1000: Obtain a pre-trained large language model and real training samples.

[0052] When the server performs fine-tuning training on the large language model, it needs to first obtain a pre-trained large language model. Specifically, it can select an open-source large language model with good language understanding and generation capabilities, such as GPT4, etc. It also needs to obtain real training samples.

[0053] Different from virtual training samples, the real training samples described in the embodiments of this specification refer to the instruction corpus actually input by users when performing a certain business in history and the response corpus returned to users for this instruction corpus in this business. This response corpus can be output by other large language models according to the instruction corpus actually input by users, or can be returned manually according to the instruction corpus actually input by users. The embodiments of this specification do not limit this.

[0054] It should be noted that the purpose of fine-tuning training the large language model is to enable the fine-tuned large language model to provide services in a specific scenario, that is, to provide a specified type of business. Therefore, when obtaining real training samples, the server can obtain the instruction corpus actually input by users when performing the specified type of business in history, and the response corpus returned to the users for the actually input instruction corpus, as real training samples. The specified type of business is the business required by the fine-tuned large language model, including all language-based services that require text generation, such as chatbots, intelligent customer service, translation, question answering, summarization, etc. The methods for obtaining real training samples include but are not limited to extracting the Q&A corpus of users performing the specified type of business in history from the database, or crawling the Q&A corpus of users performing the specified type of business in history through tools such as web crawlers on the network.

[0055] S1002: Based on the real training samples, perform preliminary fine-tuning training on the pre-trained large language model to obtain a roughly tuned large language model.

[0056] After obtaining the pre-trained large language model and the real training samples through step S1000, when the server performs fine-tuning training on the large language model using the real training samples, it can first determine the annotations of the real training samples, and then use the supervised training method to preliminarily adjust the model parameters of the pre-trained large language model based on the real training samples and the annotations, and the resulting large language model is the roughly-tuned large language model.

[0057] Among them, the method for annotating the real training samples can be manual annotation or any existing automated annotation method, and the embodiments of this specification do not limit this.

[0058] It should be noted that the effect of the preliminary fine-tuning training of the large language model in step S1002 is not entirely satisfactory. Therefore, in the embodiments of this specification, only the large language model obtained at this time is used as the roughly-tuned large language model, and further fine-tuning still needs to be performed later.

[0059] S1004: Use the roughly-tuned large language model to generate virtual training samples as the second virtual samples.

[0060] In the embodiments of this specification, in order to minimize the training cost, further fine-tuning training of the roughly-tuned large language model no longer requires obtaining real training samples, and naturally does not require any annotation operations. Instead, the roughly-tuned large language model is directly used to generate virtual training samples. Similar to step S102, the second virtual samples here also refer to the instruction corpus and the response corpus that are not truly input by the user during the execution of a certain service, but are "fabricated" by the roughly-tuned large language model. Different from step S102, a second virtual sample only includes one virtual instruction corpus and a corresponding virtual response corpus.

[0061] Specifically, the server can first obtain the prompt information in the above-specified type of service, input the prompt information into the roughly-tuned large language model, and based on the prompt information, use the roughly-tuned large language model to generate virtual training samples. That is, taking the prompt information as prior knowledge, the roughly-tuned large language model extracts knowledge points from the prompt information, and based on these knowledge points, the roughly-tuned large language model generates "self-questioning and self-answering" virtual training samples. In this process, the roughly-tuned large language model can both infer the instruction corpus that the user may input during the execution of the above-specified type of service according to these knowledge points, and generate a response corpus required for these possible input instruction corpora according to these possible input instruction corpora and these knowledge points. Then, the above-inferred possible input instruction corpora and the response corpora required for these possible input instruction corpora can be used as the second virtual samples.

[0062] For example, the server can directly input the product description of a mobile phone as a commodity on the e-commerce platform into the coarsely-tuned large language model as prompt information. Then, based on the already trained understanding and reasoning abilities, the coarsely-tuned large language model can extract various functions of the mobile phone, such as taking pictures, surfing the Internet, playing games, etc., as knowledge points. Then, the coarsely-tuned large language model can generate the instruction corpus that the user may input, such as "Describe the uses of a mobile phone", and the corresponding response corpus, such as "A mobile phone can make calls, take pictures, surf the Internet, play games, etc.", as virtual training samples.

[0063] S1006: Based on the second virtual sample, perform advanced fine-tuning training on the coarsely-tuned large language model to obtain an advanced large language model.

[0064] After obtaining the second virtual sample through step S1004, the server can, based on this second virtual sample, also use the method of supervised training to perform further fine-tuning training on the coarsely-tuned large language model, which is hereinafter referred to as advanced fine-tuning training in this specification, so as to obtain an advanced large language model. From the QA pair in the above example, "Describe the uses of a mobile phone" and "A mobile phone can make calls, take pictures, surf the Internet, play games, etc.", it can be seen that after using such virtual QA pairs as training samples to perform advanced fine-tuning training on the large language model, if the large language model receives an instruction like "Please describe the uses of a mobile phone" again, it will no longer give a response with relatively low factual accuracy like "You can make calls with others through a mobile phone", but will give a more practical response like "A mobile phone can make calls, take pictures, surf the Internet, play games, etc.".

[0065] After performing preliminary fine-tuning training on the pre-trained large language model using real training samples, use the large language model after preliminary fine-tuning training to generate virtual training samples, and then use the second virtual sample to perform advanced fine-tuning training on the large language model, so that the large language model can further consolidate the knowledge it has learned, and can improve the accuracy of the output results of the large language model without additional manual annotation costs.

[0066] S102: Use the advanced large language model to generate virtual training samples as the first virtual samples.

[0067] S104: For each first virtual sample, select the standard response corpus and perform annotation among at least two virtual response corpora included in the first virtual sample.

[0068] S106: Use the annotated first virtual samples to perform direct preference optimization (DPO) fine-tuning training on the advanced large language model to obtain the final large language model.

[0069] Steps S102 - S106 have been described in detail in Figure 1 and will not be elaborated here.

[0070] Of course, after obtaining the final large language model through the Figure 2 method shown, the performance of the final large language model can also be tested. If the test result reaches the preset expected performance index, the final large language model is used to provide services of a specified type for users, including language services that require text generation. If the test result does not reach the preset expected performance index, the final large language model can be used as the pre-trained large language model again, and return to Figure 2 the step S1000 shown, that is, use the Figure 2 method shown to continue fine-tuning the re-determined pre-trained large language model.

[0071] Specifically, when testing the performance of the final large language model, a test instruction corpus for testing (this specification does not limit whether the test instruction corpus is a virtual instruction corpus) can be input into the final large language model, and the response corpus output by the final large language model based on the test instruction corpus is obtained, and the test result of testing the performance of the final large language model is determined according to the response corpus.

[0072] The above is a fine-tuning method for a large language model provided by an embodiment of this specification. Based on the same idea, this specification also provides corresponding devices, storage media, and electronic devices.

[0073] Figure 3 The following is a schematic diagram of a fine-tuning device for a large language model provided by an embodiment of this specification. The device includes:

[0074] An acquisition module 301, configured to acquire an advanced large language model that has undergone supervised fine-tuning training (SFT);

[0075] A sample synthesis module 302, configured to use the advanced large language model to generate virtual training samples as first virtual samples; wherein, each first virtual sample includes a virtual instruction corpus and at least two virtual response corpora corresponding to the virtual instruction corpus;

[0076] A labeling module 303, configured to, for each first virtual sample, select a standard response corpus from the at least two virtual response corpora included in the first virtual sample and perform labeling;

[0077] A fine-tuning module 304, configured to perform direct preference optimization (DPO) fine-tuning training on the advanced large language model by using each labeled first virtual sample.

[0078] Optionally, the sample synthesis module 302 is specifically configured to input the prompt information in the specified type of service into the advanced large language model; and generate virtual training samples by using the advanced large language model based on the prompt information.

[0079] Optionally, the sample synthesis module 302 is specifically configured to use the prompt information as prior knowledge, generate instruction corpora that a user may input in the specified type of business through the advanced large language model; generate at least two optional response corpora required to be returned for the instruction corpora that may be input according to the instruction corpora that may be input and the prior knowledge, and use the instruction corpora that may be input and the at least two optional response corpora required to be returned as virtual training samples.

[0080] Optionally, the fine-tuning module 304 is specifically configured to adjust the model parameters of the advanced large language model with the goal of maximizing the reward value determined by the reward function according to a preset reward function; wherein, the reward function is used to determine the reward value according to the instruction corpus input to the advanced large language model and the response corpus output by the advanced large language model; the reward function is set such that after the virtual instruction corpus in the first virtual sample is input to the advanced large language model, the reward value corresponding to the response corpus output by the advanced large language model being the standard response corpus is greater than the reward value corresponding to the response corpus output not being the standard response corpus.

[0081] Optionally, the acquisition module 301 is specifically configured to acquire a pre-trained large language model and real training samples; based on the real training samples, perform preliminary fine-tuning training on the pre-trained large language model to obtain a roughly tuned large language model; use the roughly tuned large language model to generate virtual training samples as the second virtual sample; based on the second virtual sample, perform advanced fine-tuning training on the roughly tuned large language model to obtain an advanced large language model.

[0082] Optionally, the acquisition module 301 is specifically configured to acquire the instruction corpora actually input by users in the past when performing the specified type of business, and the response corpora returned to the users for the actually input instruction corpora as real training samples.

[0083] Optionally, the acquisition module 301 is specifically configured to input the prompt information in the specified type of business into the roughly tuned large language model; use the prompt information as prior knowledge, generate instruction corpora that a user may input in the specified type of business through the roughly tuned large language model; generate response corpora required to be returned for the instruction corpora that may be input according to the instruction corpora that may be input and the prior knowledge, and use the instruction corpora that may be input and the response corpora required to be returned as virtual training samples.

[0084] This specification also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to execute the above Figure 1 orFigure 2 A fine-tuning method for a large language model is provided.

[0085] Based on Figure 1 or Figure 2 the fine-tuning method for the large language model shown, the embodiments of this specification also provide Figure 4 a schematic structural diagram of an electronic device shown. As Figure 4 , at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 or Figure 2 fine-tuning method for the large language model.

[0086] The above are only the embodiments of this specification and are not used to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A fine-tuning method for a large language model, the method comprising: Obtaining an advanced large language model trained by supervised fine-tuning (SFT); Using the advanced large language model to generate virtual training samples as first virtual samples; wherein each first virtual sample includes a virtual instruction corpus and at least two virtual response corpora corresponding to the virtual instruction corpus; For each first virtual sample, select a standard response corpus from the at least two virtual response corpora included in the first virtual sample and perform annotation; Using the annotated first virtual samples to perform direct preference optimization (DPO) fine-tuning training on the advanced large language model.

2. The method according to claim 1, wherein using the advanced large language model to generate virtual training samples specifically includes: Inputting prompt information in a specified type of business into the advanced large language model; Based on the prompt information, using the advanced large language model to generate virtual training samples.

3. The method according to claim 2, wherein based on the prompt information, using the advanced large language model to generate virtual training samples specifically includes: Using the prompt information as prior knowledge, and using the advanced large language model to generate instruction corpora that a user may input in the specified type of business; According to the instruction corpora that may be input and the prior knowledge, generating at least two optional response corpora required to be returned for the instruction corpora that may be input, and using the instruction corpora that may be input and the at least two optional response corpora required to be returned as virtual training samples.

4. The method according to claim 1, wherein using the annotated first virtual samples to perform DPO fine-tuning training on the advanced large language model specifically includes: According to a preset reward function, taking maximizing the reward value determined by the reward function as the objective of fine-tuning training, and adjusting the model parameters of the advanced large language model; wherein the reward function is used to determine a reward value according to the instruction corpus input into the advanced large language model and the response corpus output by the advanced large language model; the reward function is set such that when the response corpus output by the advanced large language model after inputting the virtual instruction corpus in the first virtual sample is the standard response corpus, the corresponding reward value is greater than the reward value corresponding to the case where the output response corpus is not the standard response corpus.

5. The method according to any one of claims 1 to 4, wherein obtaining an advanced large language model trained by SFT specifically includes: Obtaining a pre-trained large language model and real training samples; Based on the real training samples, performing preliminary fine-tuning training on the pre-trained large language model to obtain a roughly-tuned large language model; Using the roughly-tuned large language model to generate virtual training samples as second virtual samples; Based on the second virtual samples, performing advanced fine-tuning training on the roughly-tuned large language model to obtain an advanced large language model.

6. The method according to claim 5, wherein obtaining real training samples specifically includes: Obtain the instruction corpus actually input by the user when performing a specified type of business in history, as well as the response corpus returned to the user for the actually input instruction corpus, as real training samples.

7. The method according to claim 5, using the coarsely tuned large language model to generate virtual training samples, specifically including: Input the prompt information in the specified type of business into the coarsely tuned large language model; Using the prompt information as prior knowledge, generate the instruction corpus that the user may input in the specified type of business through the coarsely tuned large language model; According to the instruction corpus that may be input and the prior knowledge, generate the response corpus required to be returned for the instruction corpus that may be input, and use the instruction corpus that may be input and the required response corpus as virtual training samples.

8. A fine-tuning device for a large language model, the device includes: An acquisition module for acquiring an advanced large language model that has undergone supervised fine-tuning training (SFT); A sample synthesis module for using the advanced large language model to generate virtual training samples as the first virtual samples; wherein, each first virtual sample includes a virtual instruction corpus and at least two virtual response corpora corresponding to the virtual instruction corpus; A labeling module for selecting and labeling the standard response corpus from at least two virtual response corpora included in each first virtual sample; A fine-tuning module for performing direct preference optimization (DPO) fine-tuning training on the advanced large language model using each labeled first virtual sample.

9. A computer-readable storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-7 above.

10. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method according to any one of claims 1-7 above.