Fine adjustment method and device of large language model, storage medium and electronic equipment

Through preliminary fine-tuning of real training samples, advanced fine-tuning of virtual training samples and DPO methods, the problem of insufficient output accuracy of large language models is solved, and the accuracy of model output is improved while reducing costs.

CN120373402APending Publication Date: 2025-07-25ANT ZHIXIN HANGZHOU INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510365599.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing large language models lack accuracy in following user instructions and providing factual information. Existing fine-tuning methods such as SFT and RL ignore factual accuracy when improving model performance, resulting in inaccurate output results.

Method used

The real training samples are used for preliminary fine-tuning, the virtual training samples are generated for advanced fine-tuning, and the final fine-tuning is combined with the Direct Preference Optimization (DPO) method to reduce manual labeling costs and improve model output accuracy.

Benefits of technology

On the basis of no additional manual annotation costs, the output accuracy of large language models is significantly improved, the fine-tuning training process is simplified, and the computational cost and complexity is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373402A_ABST
    Figure CN120373402A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a fine tuning method for a large language model, and the method comprises the steps: carrying out the preliminary fine tuning training of a pre-trained large language model through employing a real training sample, and generating a virtual training sample through employing the large language model which is subjected to the preliminary fine tuning training; and then the virtual training sample is used for carrying out advanced fine tuning training on the large language model, so that the large language model further consolidates the learned knowledge, and the accuracy of the output result of the large language model can be improved on the basis of not needing additional manual annotation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and particularly to a method, apparatus, storage medium, and electronic device for fine-tuning large language models. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have made significant progress in the field of natural language processing. The powerful reasoning ability of large language models enables them to play an important role in complex language tasks such as text generation, translation, summarization, and question answering. However, despite the excellent performance of LLMs in understanding and generating natural language, they still face challenges in following user instructions and providing factual information.

[0003] In real-world applications, users expect LLMs to accurately execute instructions and provide true and reliable information. For example, in intelligent assistant or customer service scenarios, the questions raised by users require the model to provide accurate answers. However, existing LLMs may produce "hallucinations" when processing these tasks, that is, generate information that does not conform to reality or is completely fabricated. This phenomenon will reduce the reliability of the model and the trust of users.

[0004] To address this issue, two main methods are commonly used to fine-tune LLMs: supervised fine-tuning (SFT) and reinforcement learning (RL). SFT trains the model using a labeled dataset to make it better follow instructions. RL guides the model to generate better outputs through reward signals. Although these methods have achieved some success in improving model performance, they usually ignore the importance of factual accuracy, resulting in the model may sacrifice the truth of facts when generating detailed and helpful responses.

[0005] Therefore, how to improve the accuracy of the output results of large language models has become an urgent problem to be solved. Summary of the Invention

[0006] Embodiments of this specification provide a method, apparatus, storage medium, and electronic device for fine-tuning large language models to partially solve the problems existing in the above-mentioned prior art.

[0007] Embodiments of this specification adopt the following technical solutions:

[0008] A method for fine-tuning a large language model provided in this specification, the method includes:

[0009] Obtain a pre-trained large language model and real training samples;

[0010] Based on the real training samples, perform preliminary fine-tuning training on the pre-trained large language model to obtain a roughly fine-tuned large language model;

[0011] Use the roughly fine-tuned large language model to generate virtual training samples;

[0012] Based on the virtual training samples, perform advanced fine-tuning training on the roughly fine-tuned large language model to obtain an advanced large language model.

[0013] A fine-tuning device for a large language model provided in this specification, the device includes:

[0014] An acquisition module, configured to acquire a pre-trained large language model and real training samples;

[0015] A preliminary fine-tuning module, configured to perform preliminary fine-tuning training on the pre-trained large language model based on the real training samples to obtain a roughly fine-tuned large language model;

[0016] A sample synthesis module, configured to generate virtual training samples using the roughly fine-tuned large language model;

[0017] An advanced fine-tuning module, configured to perform advanced fine-tuning training on the roughly fine-tuned large language model based on the virtual training samples to obtain an advanced large language model.

[0018] A computer-readable storage medium of an electronic device provided in this specification, the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned fine-tuning method of the large language model is implemented.

[0019] An electronic device provided in this specification, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the above-mentioned fine-tuning method of the large language model is implemented.

[0020] The above-mentioned at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:

[0021] The embodiments of this specification disclose a fine-tuning method for a large language model. After performing preliminary fine-tuning training on a pre-trained large language model using real training samples, the large language model after the preliminary fine-tuning training can be used to generate virtual training samples, and then the virtual training samples can be used to perform advanced fine-tuning training on the large language model, so that the large language model can further consolidate the knowledge it has learned, and can improve the accuracy of the output results of the large language model without additional manual annotation costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings described herein are used to provide a further understanding of the present specification and form a part of the present specification. The schematic embodiments of the present specification and their descriptions are used to explain the present specification and do not constitute an improper limitation of the present specification. In the drawings:

[0023] Figure 1 It is a flowchart of the fine-tuning method for the large language model provided by the embodiment of the present specification;

[0024] Figure 2 It is a detailed flowchart of the fine-tuning method for the large language model provided by the embodiment of the present specification;

[0025] Figure 3 It is a schematic diagram of a fine-tuning device for a large language model provided by the embodiment of the present specification;

[0026] Figure 4 It is a schematic diagram of the structure of an electronic device provided by the embodiment of the present specification. Detailed implementation manners

[0027] LLM fine-tuning methods in the prior art, such as SFT and RL. SFT fine-tuning training relies on a large amount of high-quality labeled data, which is difficult to obtain in many cases. Moreover, SFT focuses on improving the model's ability to follow instructions and ignores the factual accuracy of the generated text. For example, when a user needs the large language model to describe the uses of a mobile phone and inputs the instruction "Please describe the uses of a mobile phone" into the large language model, the response of the large language model is "One can make calls with others through a mobile phone". It can be seen that the output result given by the large language model is not perfect. Although it follows the user's instruction "Describe the uses of a mobile phone", due to reasons such as training samples and the training process, the uses of the mobile phone described by the large language model are obviously too simple and plain (the uses of a mobile phone also include watching videos, playing games, surfing the Internet, etc.), which is the so-called low factual accuracy. And RL fine-tuning training requires designing and training an explicit reward model, with high training complexity, a potentially unstable training process, and RL being sensitive to hyperparameter adjustment and having a high computational cost.

[0028] Based on this, the embodiment of the present specification provides a fine-tuning method for a large language model, aiming to improve the accuracy of the output result of the large language model after fine-tuning training while minimizing the training cost.

[0029] To make the purpose, technical solution, and advantages of the present specification clearer, the technical solution of the present specification will be clearly and completely described below in conjunction with the specific embodiments of the present specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present specification.

[0030] The following will, in conjunction with the accompanying drawings, elaborate on the technical solutions provided by each embodiment of this specification in detail.

[0031] Figure 1 The following is a flowchart of the fine-tuning method for the large language model provided by the embodiments of this specification, including the following steps:

[0032] S100: Obtain a pre-trained large language model and real training samples.

[0033] In the embodiments of this specification, the device for fine-tuning and training the large language model can be a server or a distributed system composed of multiple servers, etc., any electronic device that can execute model training. The following will take a server as an example for illustration.

[0034] When the server performs fine-tuning training on the large language model, it needs to first obtain a pre-trained large language model. Specifically, a large language model with good language understanding and generation capabilities and open source can be selected, such as GPT4, etc. It also needs to obtain real training samples.

[0035] The real training samples described in the embodiments of this specification refer to the instruction corpus that users actually input when performing a certain business in history and the response corpus returned to the users for this instruction corpus in this business. This response corpus can be output by other large language models according to the instruction corpus actually input by users, or can be returned manually according to the instruction corpus actually input by users. The embodiments of this specification do not limit this.

[0036] It should be noted that the purpose of fine-tuning and training the large language model is to enable the fine-tuned large language model to provide services in a specific scenario, that is, to provide a specified type of business. Therefore, when obtaining real training samples, the server can obtain the instruction corpus that users actually input when performing the specified type of business in history, and the response corpus returned to the users for the actual input instruction corpus, as real training samples. The specified type of business is the business that the fine-tuned large language model needs to provide, including all language-based services that require text generation, such as chatbots, intelligent customer service, translation, question answering, summarization, etc. The methods for obtaining real training samples include, but are not limited to, extracting the Q&A corpus when users perform the specified type of business in history from a database, or crawling the Q&A corpus when users perform the specified type of business in history through tools such as web crawlers on the network.

[0037] S102: Based on the real training samples, perform preliminary fine-tuning training on the pre-trained large language model to obtain a roughly tuned large language model.

[0038] After obtaining the pre-trained large language model and the real training samples through step S100, when the server fine-tunes the large language model using the real training samples, it can first determine the annotations of the real training samples, and then use the supervised training method to preliminarily adjust the model parameters of the pre-trained large language model based on the real training samples and the annotations, and the resulting model is the roughly-tuned large language model.

[0039] Among them, the method for annotating the real training samples can be manual annotation or any existing automated annotation method, and the embodiments of this specification do not limit this.

[0040] It should be noted that the effect of the preliminary fine-tuning training of the large language model in step S102 is not entirely satisfactory. Therefore, in the embodiments of this specification, only the large language model obtained at this time is used as the roughly-tuned large language model, and further fine-tuning is still required in the future.

[0041] S104: Generate virtual training samples using the roughly-tuned large language model.

[0042] In the embodiments of this specification, in order to minimize the training cost, further fine-tuning training of the roughly-tuned large language model no longer requires obtaining real training samples, and naturally there is no need to perform any annotation operations. Instead, the roughly-tuned large language model is directly used to generate virtual training samples. The so-called virtual training samples refer to the instruction corpus that is not actually input by the user during the execution of a certain service and the response corpus that is actually obtained, but are generated by the roughly-tuned large language model "fabricating".

[0043] Specifically, the server can first obtain the prompt information in the above-mentioned specified type of service, input the prompt information into the roughly-tuned large language model, and based on the prompt information, use the roughly-tuned large language model to generate virtual training samples. That is, taking the prompt information as a priori knowledge, the roughly-tuned large language model extracts knowledge points from the prompt information, and based on these knowledge points, the roughly-tuned large language model generates "self-questioning and self-answering" virtual training samples. In this process, the roughly-tuned large language model can not only infer the instruction corpus that the user may input during the execution of the above-mentioned specified type of service based on these knowledge points, but also generate the response corpus required for these possible input instruction corpora based on these possible input instruction corpora and these knowledge points. Then, the above-mentioned possible input instruction corpora inferred by the roughly-tuned large language model and the response corpus required for these possible input instruction corpora can be used as virtual training samples.

[0044] For example, the server can directly input the product description of a mobile phone as a commodity on the e-commerce platform into the coarsely-tuned large language model as prompt information. Then, based on the already trained understanding and reasoning abilities, the coarsely-tuned large language model can extract various functions of the mobile phone, such as taking pictures, accessing the Internet, playing games, etc., as knowledge points. Then, the coarsely-tuned large language model can generate the instruction corpus that the user may input, such as "Describe the uses of a mobile phone", and the corresponding response corpus, such as "A mobile phone can make calls, take pictures, access the Internet, play games, etc.", as virtual training samples.

[0045] S106: Based on the virtual training samples, perform advanced fine-tuning training on the coarsely-tuned large language model to obtain an advanced large language model.

[0046] After obtaining the virtual training samples through step S104, the server can, based on the virtual training samples, also use the supervised training method to perform further fine-tuning training on the coarsely-tuned large language model, which is hereinafter referred to as advanced fine-tuning training in this specification, so as to obtain an advanced large language model. From the QA pair in the above example, "Describe the uses of a mobile phone" and "A mobile phone can make calls, take pictures, access the Internet, play games, etc.", it can be seen that after using such virtual QA pairs as training samples to perform advanced fine-tuning training on the large language model, if the large language model receives an instruction similar to "Please describe the uses of a mobile phone" again, it will no longer give a response with relatively low factual accuracy, such as "You can make calls with others through a mobile phone", but will give a more practical response similar to "A mobile phone can make calls, take pictures, access the Internet, play games, etc.".

[0047] After performing preliminary fine-tuning training on the pre-trained large language model using real training samples, use the large language model after the preliminary fine-tuning training to generate virtual training samples, and then use the virtual training samples to perform advanced fine-tuning training on the large language model, so that the large language model can further consolidate the knowledge it has learned, and can improve the accuracy of the output results of the large language model without additional manual annotation costs.

[0048] In the embodiments of this specification, in order to further improve the accuracy of the output results of the large language model after fine-tuning training, after obtaining the advanced large language model through the above method, the reinforcement learning method can also be used to perform final fine-tuning training on the advanced large language model to obtain a final large language model.

[0049] However, since traditional reinforcement learning requires the design and training of an explicit reward model, the training complexity is relatively high, the training process may be unstable, and RL is sensitive to hyperparameter tuning with high computational costs. Therefore, the embodiments of this specification adopt the Direct Preference Optimization (DPO) method to replace traditional reinforcement learning. DPO is also a reinforcement learning method for training large language models. It optimizes the model through human preference data without using complex reinforcement learning algorithms. Its core idea is to directly use preference data to adjust model parameters, avoiding the fitting of explicit reward models and complex reinforcement learning optimization processes.

[0050] Specifically, after the server obtains the advanced large language model through the method as Figure 1 shown, it can use the advanced large language model again to generate virtual training samples (hereinafter, the virtual training samples generated by the advanced large language model are referred to as second virtual samples, and the virtual training samples generated by the coarsely tuned large language model are referred to as first virtual samples). Among them, one second virtual sample generated by the advanced large language model contains a virtual instruction corpus and at least two virtual response corpora corresponding to the virtual instruction corpus. After obtaining the second virtual samples, the server can, for each second virtual sample, select a standard response corpus from the at least two virtual response corpora included in the second virtual sample for annotation. Finally, use the annotated second virtual samples to perform DPO fine-tuning training on the advanced large language model, so that the fine-tuned large language model outputs the annotated standard response corpora in the second virtual samples as much as possible, as Figure 2 shown.

[0051] Figure 2 The detailed flowchart of the fine-tuning method for the large language model provided by the embodiments of this specification includes the following steps:

[0052] S200: Obtain a pre-trained large language model and real training samples.

[0053] S202: Based on the real training samples, perform preliminary fine-tuning training on the pre-trained large language model to obtain a coarsely tuned large language model.

[0054] S204: Use the coarsely tuned large language model to generate virtual training samples as the first virtual samples.

[0055] S206: Based on the first virtual samples, perform advanced fine-tuning training on the coarsely tuned large language model to obtain an advanced large language model.

[0056] Among them, steps S200 to S206 are the same as Figure 1 the steps S100 to S106 shown, and this specification will not elaborate here.

[0057] S208: Generate virtual training samples using the advanced large language model as the second virtual samples.

[0058] Each of the second virtual samples includes a virtual instruction corpus and at least two virtual response corpora corresponding to the virtual instruction corpus.

[0059] Similar to Figure 1 the steps S104 shown, the second virtual samples are also not the instruction corpora actually input by the user during the execution of the specified type of business and the response corpora actually obtained, but are generated by the "fabrication" of the advanced large language model.

[0060] Specifically, the server can first obtain the prompt information in the specified type of business, input the prompt information into the advanced large language model, and based on the prompt information, use the advanced large language model to generate virtual training samples. That is, taking the prompt information as a priori knowledge, the advanced large language model extracts knowledge points from the prompt information, and based on these knowledge points, the advanced large language model generates "self-questioning and self-answering" virtual training samples.

[0061] Different from step S104, the first virtual samples generated by the fine-tuning large language model only contain a virtual instruction corpus and a corresponding virtual response corpus, while the second virtual samples generated by the advanced large language model contain a virtual instruction corpus and more than two corresponding virtual response corpora. That is, in the process of generating "self-questioning and self-answering" virtual training samples according to the above knowledge points, the advanced large language model can not only infer the instruction corpus that the user may input during the execution of the above specified type of business according to these knowledge points, but also generate at least two optional response corpora required for the possible input instruction corpus according to the possible input instruction corpus and these knowledge points. Then, the above possible input instruction corpus inferred by the advanced large language model and at least two optional response corpora required for the possible input instruction corpus can be used as virtual training samples, that is, the second virtual samples.

[0062] For example, the server can directly input the product description of the mobile phone as a product on the e-commerce platform into the advanced large language model. Then, based on the trained understanding ability and reasoning ability, the advanced large language model can extract various functions such as the mobile phone can take pictures, surf the Internet, play games, etc. as knowledge points. Then, the advanced large language model can generate the instruction corpus that the user may input, "Describe the uses of the mobile phone", and four corresponding optional response corpora, "The mobile phone can make calls", "The mobile phone can take pictures", "The mobile phone can surf the Internet", "The mobile phone can play games", etc. as the second virtual samples.

[0063] S210: For each second virtual sample, select a standard response corpus from among the at least two virtual response corpora included in the second virtual sample and perform annotation.

[0064] In the embodiments of this specification, since each second virtual sample includes a virtual instruction corpus and more than two virtual response corpora, in step S210, for each second virtual sample, one or several preferred response corpora can be selected from among the at least two virtual response corpora included in the second virtual sample as the standard response corpus and annotated.

[0065] Continuing with the above example, the second virtual sample includes the virtual instruction corpus "Describe the uses of a mobile phone" and the corresponding four optional virtual response corpora "A mobile phone can make calls", "A mobile phone can take pictures", "A mobile phone can access the Internet", "A mobile phone can play games". If it is desired to fine-tune the trained large language model to be more inclined to output "A mobile phone can play games" when encountering a problem like the virtual instruction corpus "Describe the uses of a mobile phone", then the virtual response corpus "A mobile phone can play games" in the second virtual sample can be annotated as the standard response corpus.

[0066] S212: Use each of the annotated second virtual samples to perform DPO fine-tuning training on the advanced large language model as the final fine-tuning training to obtain the final large language model.

[0067] When using each of the annotated second virtual samples to perform DPO fine-tuning training on the advanced large language model, a preset reward function can be used to perform DPO fine-tuning training on the advanced large language model, where the reward function is used to determine a reward value based on the instruction corpus input to the large language model and the response output by the large language model. The reward function is set such that after the virtual instruction corpus in the second virtual sample is input into the large language model, the reward value corresponding to the response output by the large language model being the standard response corpus is greater than the reward value corresponding to the response output not being the standard response corpus. Thus, the reward value determined by the reward function can be maximized as the objective of the fine-tuning training to adjust the model parameters of the advanced large language model.

[0068] As can be seen from the above method, the DPO fine-tuning training method can simplify the complex reinforcement learning problem of "how to select the response with the maximum reward" into a relatively simple binary classification problem of "is the reward for this response high or low", thereby reducing the training complexity of the large language model in the reinforcement learning stage, simplifying the fine-tuning training of the large language model, and improving the efficiency of the fine-tuning training.

[0069] Of course, when passing through Figure 2After obtaining the final large language model by the method shown, the performance of the final large language model can also be tested. If the test result reaches the preset expected performance index, the final large language model is used to provide the specified type of service for the user, including language services that require text generation. If the test result does not reach the preset expected performance index, the final large language model can be used again as the pre-trained large language model and return to Figure 2 the step S200 shown in Figure 2 the method shown to continue fine-tuning the re-determined pre-trained large language model.

[0070] Specifically, when testing the performance of the final large language model, the test instruction corpus used for testing (this specification does not limit whether the test instruction corpus is a virtual instruction corpus) can be input into the final large language model, and the response corpus output by the final large language model based on the test instruction corpus is obtained, and the test result of testing the performance of the final large language model is determined according to the response corpus.

[0071] The above is a fine-tuning method for a large language model provided by an embodiment of this specification. Based on the same idea, this specification also provides corresponding devices, storage media, and electronic devices.

[0072] Figure 3 The following is a schematic diagram of a fine-tuning device for a large language model provided by an embodiment of this specification. The device includes:

[0073] An acquisition module 301, configured to acquire a pre-trained large language model and real training samples;

[0074] A preliminary fine-tuning module 302, configured to perform preliminary fine-tuning training on the pre-trained large language model based on the real training samples to obtain a roughly fine-tuned large language model;

[0075] A sample synthesis module 303, configured to generate virtual training samples by using the roughly fine-tuned large language model;

[0076] An advanced fine-tuning module 304, configured to perform advanced fine-tuning training on the roughly fine-tuned large language model based on the virtual training samples to obtain an advanced large language model.

[0077] Optionally, the acquisition module 301 is specifically configured to acquire the instruction corpus actually input by the user when performing the specified type of service historically, and the response corpus returned to the user for the actually input instruction corpus, as real training samples.

[0078] Optionally, the preliminary fine-tuning module 302 is specifically configured to determine the annotation of the true training samples; and preliminarily adjust the model parameters of the pre-trained large language model by using a supervised training method based on the true training samples and the annotation.

[0079] Optionally, the sample synthesis module 303 is specifically configured to input the prompt information in a specified type of business into the coarsely-tuned large language model; and generate virtual training samples based on the prompt information by using the coarsely-tuned large language model.

[0080] Optionally, the sample synthesis module 303 is specifically configured to use the prompt information as prior knowledge, generate instruction corpus that a user may input in the specified type of business by using the coarsely-tuned large language model; generate response corpus required to be returned for the possible input instruction corpus according to the possible input instruction corpus and the prior knowledge, and use the possible input instruction corpus and the required response corpus as virtual training samples.

[0081] Optionally, the apparatus further includes:

[0082] A final fine-tuning module 305, configured to perform final fine-tuning training on the advanced large language model by using a reinforcement learning method to obtain a final large language model.

[0083] Optionally, the apparatus further includes:

[0084] A testing module 306, configured to test the performance of the final large language model; if the test result reaches a preset expected performance index, use the final large language model to provide services of the specified type of business for the user; if the test result does not reach the preset expected performance index, use the final large language model as the pre-trained large language model again, and continue to perform fine-tuning on the re-determined pre-trained large language model.

[0085] This specification also provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, it can be used to execute the above Figure 1 or Figure 2 provided fine-tuning method of the large language model.

[0086] Based on Figure 1 or Figure 2 the fine-tuning method of the large language model shown, embodiments of this specification also provide Figure 4 the structural schematic diagram of an electronic device shown. As Figure 4, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 or Figure 2 fine-tuning method of the large language model described above.

[0087] The above are only examples of this specification and are not used to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A fine-tuning method for a large language model, the method comprising: Obtaining a pre-trained large language model and real training samples; Based on the real training samples, performing preliminary fine-tuning training on the pre-trained large language model to obtain a roughly-tuned large language model; Using the roughly-tuned large language model to generate virtual training samples; Based on the virtual training samples, performing advanced fine-tuning training on the roughly-tuned large language model to obtain an advanced large language model.

2. The method according to claim 1, wherein obtaining real training samples specifically comprises: Obtaining the instruction corpus actually input by users in the past when performing a specified type of business, and the response corpus returned to the users for the actually input instruction corpus, as real training samples.

3. The method according to claim 1, wherein based on the real training samples, performing preliminary fine-tuning training on the pre-trained large language model specifically comprises: Determining the annotations of the real training samples; Using the method of supervised training, and based on the real training samples and the annotations, preliminarily adjusting the model parameters of the pre-trained large language model.

4. The method according to claim 1, wherein using the roughly-tuned large language model to generate virtual training samples specifically comprises: Inputting the prompt information in a specified type of business into the roughly-tuned large language model; Based on the prompt information, using the roughly-tuned large language model to generate virtual training samples.

5. The method according to claim 4, wherein based on the prompt information, using the roughly-tuned large language model to generate virtual training samples specifically comprises: Taking the prompt information as prior knowledge, and generating, through the roughly-tuned large language model, the instruction corpus that users may input in the specified type of business; According to the possible input instruction corpus and the prior knowledge, generating the response corpus required to be returned for the possible input instruction corpus, and using the possible input instruction corpus and the required response corpus as virtual training samples.

6. The method according to claim 1, the method further comprising: Using the method of reinforcement learning to perform final fine-tuning training on the advanced large language model to obtain a final large language model.

7. The method according to claim 6, the method further comprising: Testing the performance of the final large language model; If the test result reaches the preset expected performance index, then using the final large language model to provide a specified type of business for users; If the test result does not reach the preset expected performance index, then using the final large language model as the pre-trained large language model again, and continuing to perform fine-tuning on the re-determined pre-trained large language model.

8. A fine-tuning device for a large language model, the device comprising: An acquisition module, configured to obtain a pre-trained large language model and real training samples; A preliminary fine-tuning module, configured to perform preliminary fine-tuning training on the pre-trained large language model based on the real training samples to obtain a roughly-tuned large language model; A sample synthesis module, configured to use the roughly-tuned large language model to generate virtual training samples; An advanced fine-tuning module, configured to perform advanced fine-tuning training on the coarsely-tuned large language model based on the virtual training samples, so as to obtain an advanced large language model.

9. A computer-readable storage medium storing a computer program, which when executed by a processor implements the method according to any one of claims 1-7 above.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1-7 above when executing the program.

Citation Information

Cited By

  • Marketing decision-making method and device, equipment and medium

    CN120707176A