Model fine tuning and question and answer data processing method and device and storage medium

By introducing preference models and alignment fine-tuning operations into the large language model, and using the feedback data of the target application scenario for reinforcement learning, the problem of insufficient accuracy and reliability of the execution results of large language models in specific tasks is solved, and more efficient reply generation is achieved.

CN120179759APending Publication Date: 2025-06-20HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311726503.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Large language models have poor accuracy and reliability in execution results in specific tasks, and may generate unanticipated false content and unsafe speech.

Method used

By performing reinforcement learning based on feedback data in the target application scenario, model the large language model is fine-tuned, and the preference model is used to score the Q&A pairs, and aligning fine-tuning operations are carried out to improve the accuracy and reliability of the output reply.

Benefits of technology

It effectively improves the accuracy and reliability of the output replies of large language models in specific tasks, and reduces the risk of generating false content and unsafe speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179759A_ABST
    Figure CN120179759A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model fine tuning method and device, a question and answer data processing method and device and a storage medium. In the model fine tuning method, a plurality of instruction texts in an instruction data set are sequentially input into a language model after instruction fine tuning, and a plurality of candidate reply texts of the plurality of instruction texts are obtained. The instruction texts and the candidate reply texts corresponding to the instruction texts form multiple groups of question and answer pairs. And inputting the multiple groups of question and answer pairs into the preference model to obtain respective preference scores of the multiple groups of question and answer pairs, and performing at least one round of alignment fine tuning operation on the language model according to the multiple groups of question and answer pairs and the preference scores of the multiple groups of question and answer pairs. Wherein the preference model is obtained by training the returned question and answer data in the target application scene, so that the preference model can accurately output the preference score matched with the actual preference information in the target application scene according to the input question and answer pair, thereby forming accurate and reliable guidance for the alignment fine adjustment process; and the professional performance of the language model in the professional field is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and in particular, to a method, device, and storage medium for model fine-tuning and question-and-answer data processing. Background Art

[0002] With the development of artificial intelligence, large language models (LLMs) have been gradually widely used. Such large language models are pre-trained on extremely large-scale unsupervised corpora. After pre-training, the large language models can master various knowledge in the large-scale corpora. In practical applications, in order to enable the large language models to demonstrate their capabilities in specific tasks, it is necessary to perform instruction fine-tuning on the large language models based on specific data sets in different fields so that they can complete specific tasks in different application fields. However, the training objective of the large language models is to predict the probability of the next word in the reply, so the fine-tuned large language models may generate unexpected results, such as some fabricated false content and some unsafe remarks, etc. The above unexpected generated results make the execution results of the large language models for specific tasks have poor accuracy and reliability. Therefore, there is a need to propose a new solution. Summary of the Invention

[0003] Multiple aspects of this application provide a method, device, and storage medium for model fine-tuning and question-and-answer data processing, which are used to perform reinforcement learning on a large language model based on feedback data in a target application scenario, and improve the accuracy and reliability of the replies output by the large language model.

[0004] An embodiment of this application provides a model fine-tuning method, including: sequentially inputting multiple instruction texts into a language model after instruction fine-tuning to obtain first reply results of the multiple instruction texts respectively; the first reply result of any instruction text includes: multiple candidate reply texts; the multiple instruction texts and their respective corresponding multiple candidate reply texts form multiple groups of question-and-answer pairs; inputting the multiple groups of question-and-answer pairs into a preference model to obtain preference scores of the multiple groups of question-and-answer pairs respectively; the preference score of any question-and-answer pair is used to describe the preference information of the instruction text in the question-and-answer pair for the candidate reply text; wherein, the preference model is trained according to the historical candidate reply texts generated by the language model for multiple historical instructions in a target application scenario and the preference feedback data of the historical candidate reply texts of the multiple historical instructions respectively; performing at least one round of alignment fine-tuning operation on the language model according to the multiple groups of question-and-answer pairs and the preference scores of the multiple groups of question-and-answer pairs.

[0005] Optionally, before inputting the multiple groups of question-and-answer pairs into the preference model to obtain the preference scores of the multiple groups of question-and-answer pairs, it further includes: obtaining a plurality of sample data groups; any one of the sample data groups includes: a historical instruction, a plurality of historical candidate response texts of the historical instruction, and preference feedback data of the plurality of historical candidate response texts; wherein, the historical instruction and the plurality of historical candidate response texts form multiple groups of question-and-answer pair samples; inputting the multiple groups of question-and-answer pair samples into the preference model to obtain the preference scores of the multiple groups of question-and-answer pair samples; determining the preference analysis loss of the preference model according to the preference scores of the multiple groups of question-and-answer pair samples and the preference feedback data of the plurality of historical candidate response texts; training the preference model according to the preference analysis loss, and stopping the training when the preference analysis loss of the preference model converges to a specified range.

[0006] Optionally, it further includes: during the process of sequentially inputting multiple instruction texts in the instruction dataset into the language model after instruction fine-tuning, performing at least one instruction fine-tuning operation on the language model according to the set fine-tuning conditions; the fine-tuning conditions include: the time duration from the last instruction fine-tuning operation is a specified time duration, or a specified number of instruction texts are input after the last instruction fine-tuning operation.

[0007] Optionally, before performing at least one round of alignment fine-tuning operation on the language model according to the multiple groups of question-and-answer pairs and the preference scores of the multiple groups of question-and-answer pairs, it further includes: for any one of the multiple instruction texts, inputting the instruction text into the corresponding initial language model of the language model to obtain a second response result; determining the respective fine-tuning offsets of the multiple groups of question-and-answer pairs corresponding to the instruction text according to the first response result and the second response result of the instruction text; updating the preference scores of the multiple groups of question-and-answer pairs corresponding to the instruction text according to the respective fine-tuning offsets of the multiple groups of question-and-answer pairs corresponding to the instruction text to obtain the updated preference scores of the multiple groups of question-and-answer pairs corresponding to the instruction text.

[0008] Optionally, according to the first response result and the second response result of the instruction text, determine the fine-tuning offsets of each group of question-answer pairs corresponding to the instruction text, including: obtaining the first vocabulary probability distribution corresponding to each of the multiple candidate response texts of the instruction text according to the first response result of the instruction text; the first vocabulary probability distribution of any candidate response text is used to describe the probability that each word in the candidate response text is predicted as the next output word by the language model in the vocabulary; obtaining the second vocabulary probability distribution corresponding to each of the multiple candidate response texts of the instruction text from the second response result of the instruction text; the second vocabulary probability distribution corresponding to any candidate response text is used to describe the probability that each word in the candidate response text is predicted as the next output word by the initial language model in the vocabulary; determining the alignment fine-tuning offsets of each of the multiple candidate response texts of the instruction text according to the first vocabulary probability distribution and the second vocabulary probability distribution corresponding to each of the multiple candidate response texts of the instruction text.

[0009] Optionally, according to the multiple groups of question-answer pairs and the preference scores of the multiple groups of question-answer pairs, perform at least one round of alignment fine-tuning operations on the language model, including: sampling the multiple groups of question-answer pairs to obtain a first question-answer pair; determining the preference loss of the first question-answer pair according to the preference score of the first question-answer pair; performing an alignment fine-tuning operation on the language model according to the preference loss and the instruction fine-tuning loss of the initial language model in the instruction fine-tuning stage.

[0010] Optionally, it further includes: when the instruction text is input into the language model after instruction fine-tuning, inputting the instruction text into the instruction evaluation model corresponding to the language model to obtain the first instruction value score of the instruction text; according to the multiple groups of question-answer pairs and the preference scores of the multiple groups of question-answer pairs, performing at least one round of alignment fine-tuning operations on the language model, including: using the proximal policy optimization algorithm to perform at least one round of alignment fine-tuning operations on the language model according to the multiple groups of question-answer pairs, the preference scores of the multiple groups of question-answer pairs, and the first instruction value scores of the multiple instruction texts.

[0011] Optionally, according to the multiple sets of question-answer pairs, the preference scores of the multiple sets of question-answer pairs, and the first instruction value scores of the multiple instruction texts respectively, the proximal policy optimization algorithm is used to perform at least one round of alignment fine-tuning operation on the language model, including: in any round of alignment fine-tuning operation, sampling the multiple sets of question-answer pairs to obtain a second question-answer pair; the second question-answer pair includes a target instruction text and a target candidate answer text; inputting the target instruction text into the language model and the instruction evaluation model after at least one instruction fine-tuning respectively to obtain a third answer result and a second instruction value score of the target instruction text; obtaining a third vocabulary probability distribution of the target candidate answer text according to the third answer result of the target instruction text; determining a proximal policy optimization loss according to the third vocabulary probability distribution and the first vocabulary probability distribution of the target candidate answer text; determining an instruction quality loss according to the second instruction value score of the target instruction and the preference score of the second question-answer pair; performing an alignment fine-tuning operation on the language model at least according to the instruction quality loss, the proximal policy optimization loss, and the instruction fine-tuning loss of the initial language model in the instruction alignment fine-tuning stage.

[0012] Optionally, determining a proximal policy optimization loss according to the third vocabulary probability distribution and the first vocabulary probability distribution of the target candidate answer text in the second question-answer pair includes: determining a vocabulary probability ratio difference according to the first vocabulary probability distribution and the third vocabulary probability distribution of the target candidate answer text; determining an advantage score according to the preference score of the target candidate answer text and the first instruction value score; determining the proximal policy optimization loss according to the advantage score and the probability ratio difference.

[0013] Optionally, the instruction evaluation model and the language model share a model structure and model parameters.

[0014] The embodiment of the present application further provides a method for processing question-answer data, including: obtaining a user instruction in a target application scenario; inputting the user instruction into the language model after alignment fine-tuning to obtain an answer result output by the language model; wherein, the language model is fine-tuned by using the model fine-tuning method provided by the embodiment of the present application.

[0015] The embodiment of the present application further provides a server, including: a memory and a processor; the memory is used for storing one or more computer instructions; the processor is used for executing the one or more computer instructions to: execute the steps in the method provided by the embodiment of the present application.

[0016] The embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it can implement the steps in the method provided by the embodiment of the present application.

[0017] In the embodiment of the present application, after obtaining the instruction dataset, multiple instruction texts in the instruction dataset can be sequentially input into the language model after instruction fine-tuning to obtain multiple candidate response texts for each of the multiple instruction texts. The multiple instruction texts and their respective multiple candidate response texts form multiple groups of question-and-answer pairs. The multiple groups of question-and-answer pairs are input into the preference model to obtain the preference scores for each of the multiple groups of question-and-answer pairs. Based on the multiple groups of question-and-answer pairs and the preference scores of the multiple groups of question-and-answer pairs, at least one round of alignment fine-tuning operation can be performed on the language model. In this embodiment, on the one hand, by using the preference model to score the preference of the question-and-answer pairs, the preference scores of the question-and-answer pairs can be automatically and efficiently generated, reducing the dependence on manual analysis operations. On the other hand, the preference model is trained with the question-and-answer data flowing back from the target application scenario. Therefore, the preference model can accurately output the preference scores that match the actual preference information in the target application scenario according to the input question-and-answer pairs, thus forming an accurate and reliable guidance for the alignment fine-tuning process, and further effectively improving the professional performance of the language model in the professional field. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0019] Figure 1 is a schematic flowchart of a model fine-tuning method provided by an exemplary embodiment of the present application;

[0020] Figure 2 is a schematic flowchart of a process for training a preference model provided by an exemplary embodiment of the present application;

[0021] Figure 3 is a schematic flowchart of a data generation stage provided by an exemplary embodiment of the present application;

[0022] Figure 4 is a schematic flowchart of an alignment fine-tuning process provided by an exemplary embodiment of the present application;

[0023] Figure 5 is a schematic flowchart of a question-and-answer data processing method provided by an exemplary embodiment of the present application;

[0024] Figure 6 is a schematic structural diagram of a server provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0026] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Multiple" generally includes at least two, but does not exclude the case of including at least one.

[0027] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0028] It should also be noted that the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, so that a commodity or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a commodity or system. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of another identical element in the commodity or system including the said element.

[0029] Regarding the problem that the fine-tuned large language model may generate unexpected results, and the technical problem of having poor accuracy and reliability in the execution results of specific tasks, in some embodiments of this application, a solution is provided. The following will detail the technical solutions provided by each embodiment of this application in conjunction with the drawings.

[0030] Figure 1 is a schematic flowchart of a model fine-tuning method provided by an exemplary embodiment of this application. The method may include as Figure 1 shown in the steps:

[0031] Step 101: Input multiple instruction texts into the instruction fine-tuned language model in sequence to obtain the first response results of the multiple instruction texts respectively; the first response result of any instruction text includes: multiple candidate response texts; the multiple instruction texts and their corresponding multiple candidate response texts form multiple groups of question-and-answer pairs.

[0032] Step 102: Input the multiple groups of question-answer pairs into the preference model to obtain the preference scores of each of the multiple groups of question-answer pairs; the preference score of any question-answer pair is used to describe the preference information of the instruction text in the question-answer pair for the candidate response text; wherein, the preference model is trained based on the historical candidate response texts generated by the language model for multiple historical instructions in the target application scenario and the preference feedback data of the historical candidate response texts of each of the multiple historical instructions.

[0033] Step 103: Perform at least one round of alignment fine-tuning operation on the language model according to the multiple groups of question-answer pairs and the preference scores of the multiple groups of question-answer pairs.

[0034] The model fine-tuning method provided in this embodiment is used to perform alignment fine-tuning on the language model after instruction fine-tuning. In some embodiments, the language model refers to a large language model with a parameter quantity greater than a set parameter quantity threshold. In this embodiment, the language model after instruction fine-tuning is a supervised fine-tuning model (SFT), that is, a model obtained by performing instruction fine-tuning on a pre-trained large model in combination with the application data in the target application scenario of the language model. The language model after instruction fine-tuning can have good effects on many tasks in the target application scenario. Among them, the target application scenario refers to the application scenario of the language model in a vertical industry or vertical field. To further improve the performance of the language model in the target application scenario, the method provided in this embodiment can be used to perform alignment fine-tuning on the language model after instruction fine-tuning in combination with the application data flowing back in the target application scenario. Among them, alignment fine-tuning means making the large model align with human preferences through fine-tuning.

[0035] In the data generation stage before alignment fine-tuning, a training data set required for alignment fine-tuning can be generated according to the instruction data set. Among them, the instruction data set can include multiple instruction texts obtained from the target application scenario of the language model. Optionally, user instructions can be obtained from the application data actually generated by the target model in the target application scenario as instruction texts. In the alignment fine-tuning stage, multiple instruction texts in the instruction data set can be sequentially input into the language model. After any instruction text is input into the language model, the language model can output a response result for the instruction text. In this embodiment, the response result output by the language model for any instruction text is described as the first response result of the instruction text. The first response result of any instruction text can include: multiple candidate response texts. Among them, multiple instruction texts and their respective corresponding multiple candidate response texts form multiple groups of question-answer pairs. Among them, any group of question-answer pairs includes an instruction text and a corresponding candidate response text.

[0036] Among multiple candidate response texts corresponding to any instruction text, different candidate response texts have different degrees of matching with user preferences in the target application scenario. Among them, the user preferences can include subjective preferences of users, or can also include preferences under the constraints of regulations and systems. For example, users prefer true and reliable candidate response texts, or users prefer candidate response texts that meet the needs of safe speech.

[0037] In this embodiment, the preference score of a question-and-answer pair is used to describe the degree to which the question-and-answer pair meets user preferences. Among them, the higher the preference score of any question-and-answer pair, the closer the candidate response text in the question-and-answer pair is to the user's expectation of a reasonable response to the instruction text in the target application scenario, that is, the candidate response text in the target question-and-answer pair can better meet user preferences. In this embodiment, a preference model can be used to evaluate multiple groups of question-and-answer pairs to obtain the preference scores of each group of question-and-answer pairs. Among them, the preference model is trained according to the application data returned in the target application scenario. Among them, the returned application data includes historical candidate response texts generated by the language model for multiple historical instructions and preference feedback data for the respective historical candidate response texts of the multiple historical instructions. Among them, the multiple historical instructions are actual user instructions received when the language model is used in the target application scenario, reflecting the real user needs in the target application scenario; the historical candidate response texts of the historical instructions are real response data generated by the language model according to the actual user instructions, and the preference feedback data of each historical candidate response text can be used to represent the real preference information of users in the target application scenario for the candidate response texts output by the language model. Therefore, training the preference model based on the returned application data can enable the preference model to efficiently and accurately learn the response preference knowledge in the target application scenario, so as to guide the alignment and fine-tuning process of the language model. The training process of the preference model will be further exemplarily described below.

[0038] Optionally, multiple sample data groups can be obtained. Any sample data group includes: a historical instruction, multiple historical candidate response texts of the historical instruction, and preference feedback data among the multiple historical candidate response texts. Among them, the historical instruction and the multiple historical candidate response texts form multiple groups of question-and-answer pair samples. For example, after the language model is launched and used, the language model can be used to implement automatic question answering for user questions in the target application scenario. As Figure 2 shown, for a real user question Q in the target application scenario, the language model can generate multiple candidate responses, such as Figure 2The responses A, B, C, and D shown. In some cases, the user can make preference selections for multiple of the above candidate responses. For example, the user can like certain preferred responses and dislike responses that are not as expected. The language model can generate multiple sets of preference pairs based on the user's preference selection operations. For example, if the user selects response A, the language model can automatically generate three sequences of preference pairs, namely: A > B, A > C, A > D, a total of 3 preference pairs. In other cases, after the language model generates multiple responses, the operation and maintenance personnel or technical personnel related to the language model can sort the multiple responses. For example, the sorting result can be: such as A > B > C = D. Based on the above sorting result, the language model can construct 6 preference pairs. The above preference pairs can be used as preference feedback data for multiple responses in the target application scenario.

[0039] In the target application scenario, the user question input into the language model can be used as a historical instruction; the multiple responses output by the language model to the user question can be used as multiple historical candidate response texts for the historical instruction; the preference pairs corresponding to the multiple responses can be used as preference feedback data for the multiple historical candidate response texts. Continuing with the above example, a sample data group corresponding to the user question Q can include: {(question Q, response A), (question Q, response B), (question Q, response C), (question Q, response D), A > B, A > C, A > D}.

[0040] Based on multiple historical instructions, multiple historical candidate response texts for each historical instruction, and preference feedback data for the multiple historical candidate response texts, a preference model can be trained.

[0041] Continuing with any sample data group as an example. Optionally, multiple sets of question-and-answer pairs in the sample data group can be input into the preference model to obtain the preference scores for the multiple sets of question-and-answer pair samples. Based on the preference scores for the multiple sets of question-and-answer pair samples and the preference feedback data between the multiple historical candidate response texts, the training loss of the preference model can be determined. Based on the training loss, the preference model can be trained so that the preference model learns the preference of the question text for the multiple candidate response texts under the supervision of the preference information. When the training loss of the preference model converges to a specified range, the preference model can be output.

[0042] After obtaining the preference scores for each of the multiple sets of question-and-answer pairs output by the preference model, the language model can be aligned and fine-tuned based on the multiple sets of question-and-answer pairs and their preference scores. During the alignment and fine-tuning process, the language model can, under the guidance of the preference scores, learn to output candidate response texts that can meet the preference requirements of the target application scenario.

[0043] In this embodiment, after obtaining the instruction dataset, multiple instruction texts in the instruction dataset can be sequentially input into the language model after instruction fine-tuning to obtain multiple candidate response texts for each of the multiple instruction texts. The multiple instruction texts and their respective multiple candidate response texts form multiple groups of question-and-answer pairs. The multiple groups of question-and-answer pairs are input into the preference model to obtain preference scores for each of the multiple groups of question-and-answer pairs. Based on the multiple groups of question-and-answer pairs and the preference scores of the multiple groups of question-and-answer pairs, at least one round of alignment fine-tuning operation can be performed on the language model. In this embodiment, on the one hand, by using the preference model to score the preference of the question-and-answer pairs, the preference scores of the question-and-answer pairs can be automatically and efficiently generated, reducing the dependence on manual analysis operations. On the other hand, the preference model is trained by the question-and-answer data returned from the target application scenario. Therefore, the preference model can accurately output preference scores that match the actual preference information in the target application scenario according to the input question-and-answer pairs, thus forming an accurate and reliable guidance for the alignment fine-tuning process, and further effectively improving the professional performance of the language model in the professional field.

[0044] The process of obtaining multiple groups of question-and-answer pairs based on the instruction dataset and using the preference model to generate preference scores for each of the multiple groups of question-and-answer pairs can be described as the data generation stage, and the generated data can be described as the empirical dataset required for alignment fine-tuning.

[0045] It should be noted that in the process of sequentially inputting multiple instruction texts in the instruction dataset into the language model after instruction fine-tuning, at least one instruction fine-tuning operation can be performed on the language model according to the set fine-tuning conditions. Among them, any instruction fine-tuning operation is used to update all or part of the model parameters of the language model. In some optional embodiments, the fine-tuning conditions may include: the time duration since the last instruction fine-tuning operation is a specified duration. Among them, the specified duration can be 5 minutes, 10 minutes, or any other duration, which is not limited in this embodiment. For example, before inputting the instruction text into the language model, an instruction fine-tuning operation can be performed; after 5 minutes, the next instruction fine-tuning operation is performed, and so on. In some other optional embodiments, the fine-tuning conditions may include: a specified number of instruction texts are input after the last instruction fine-tuning operation. Among them, the specified number can be 50, 80, 100, or other numbers, which is not limited in this embodiment. For example, an instruction fine-tuning operation can be performed on the language model every time 50 instruction texts are input into the language model.

[0046] Based on this embodiment, in the data generation stage, a fine-tuning operation can be performed on the language model, so that the generated training data has higher diversity to improve the reinforcement learning effect in the alignment fine-tuning stage.

[0047] In some exemplary embodiments, to reduce the impact of the model fine-tuning operation in the above data generation stage on the subsequent alignment fine-tuning process, the fine-tuning offset of the language model relative to the initial language model (i.e., the SFT model) can be obtained, and based on the fine-tuning offset, the preference scores of multiple groups of question-answer pairs can be adjusted. Among them, the initial language model refers to the model obtained after performing instruction fine-tuning on the pre-trained large language model. This initial language model has not performed other fine-tuning operations after the instruction fine-tuning is completed. Therefore, the initial language model possesses all the knowledge learned in the instruction fine-tuning stage. Among them, the language model is initialized from the initial language model after instruction fine-tuning.

[0048] The following will take any instruction text as an example to exemplarily illustrate the method for obtaining the fine-tuning offset.

[0049] Optionally, the instruction text can be input into the initial language model corresponding to the language model to obtain the second reply result corresponding to the initial language model. The second reply result of the initial language model for the instruction text is used to represent the reply of the initial language model regarding the instruction text generated based on the knowledge learned in the instruction fine-tuning stage. Based on the second reply result output by the initial language model for this instruction text and the first reply result output by the language model for this instruction text, the fine-tuning offset of each group of question-answer pairs corresponding to this instruction text can be determined.

[0050] Continuing with an example of multiple candidate reply texts corresponding to any instruction text for exemplary illustration.

[0051] Optionally, the first reply result output by the language model for any instruction text may include: a first vocabulary probability matrix. Among them, the first vocabulary probability matrix includes at least one probability vector, and each probability vector corresponds to a text position. The elements in the probability vector corresponding to any text position are used to describe the probability that each word in the vocabulary is predicted by the language model as the output word (i.e., the next output word) corresponding to this text position.

[0052] Correspondingly, the second reply result output by the initial language model for this instruction text may include: a second vocabulary probability matrix. Among them, the second vocabulary probability matrix includes at least one probability vector, and each probability vector corresponds to a text position. The elements in the probability vector corresponding to any text position are used to describe the probability that each word in the vocabulary is predicted by this initial language model as the output word (i.e., the next output word) corresponding to this text position.

[0053] In some alternative embodiments, according to the first response result of the instruction text, the first vocabulary probability distribution corresponding to each of the multiple candidate response texts of the instruction text can be obtained; the first vocabulary probability distribution of any candidate response text is used to describe the probability that each word in the candidate response text is predicted as the next output word by the language model in the vocabulary. For example, the instruction text is: "What is the shipping method?", and one of the candidate response texts output by the language model is: "It is expected to be sent by air logistics tomorrow". The first vocabulary probability distribution of this candidate response text obtained from the first vocabulary probability matrix is: tomorrow (90%), send (97%), air (95%), logistics (98%). That is, in the prediction process, when the language model predicts the first word in the candidate response text, the probability that the word "tomorrow" in the vocabulary is the first word is 90%; after predicting the first word, the probability that the word "send" in the vocabulary is the next word is 97%; after predicting the second word, the probability that the word "air" in the vocabulary is the next word is 95%; after predicting the third word, the probability that the word "logistics" in the vocabulary is the next word is 98%.

[0054] Correspondingly, from the second response result, the second vocabulary probability distribution corresponding to each of the multiple candidate response texts of the instruction text can be obtained. The second vocabulary probability distribution corresponding to any candidate response text is used to describe the probability that each word in the candidate response text is predicted as the next output word by the initial language model in the vocabulary. According to the first vocabulary probability distribution and the second vocabulary probability distribution corresponding to each of the multiple candidate response texts of the instruction text, the alignment fine-tuning offsets of each of the multiple candidate response texts of the instruction text can be determined.

[0055] Continuing with the foregoing example for illustration, in the prediction process, in the second vocabulary probability matrix output by the initial language model, the probability that the word "tomorrow" in the vocabulary is the first word is 80%; after predicting the first word, the probability that the word "send" in the vocabulary is the next word is 97%; after predicting the second word, the probability that the word "air" in the vocabulary is the next word is 85%; after predicting the third word, the probability that the word "logistics" in the vocabulary is the next word is 90%.

[0056] According to the first vocabulary probability distribution and the second vocabulary probability distribution corresponding to each of the multiple candidate response texts, calculate the fine-tuning offsets of each of the multiple question-and-answer pairs. Optionally, the fine-tuning offset of any question-and-answer pair (x, y i ) can be calculated using the Kullback-Leibler Divergence (i.e., relative entropy divergence), as shown in the following formula 1:

[0057]

[0058] Among them, x represents the instruction text, and y i represents the i-th candidate response text of the instruction text. π SFT (y i |x) represents the candidate response text y i corresponding to the second vocabulary probability distribution in the second response result output by the initial language model, that is Figure 3 the L2(x, yi) shown. represents the first vocabulary probability distribution of the candidate response text y i output by the language model in the first response result, that is Figure 3 the L1(x, yi) shown. β is a constant coefficient and can be determined according to empirical values. The KL divergence is used to measure the degree of difference between the first vocabulary probability distribution and the second vocabulary probability distribution.

[0059] After determining the respective fine-tuning offsets of multiple groups of question-and-answer pairs based on the above embodiments, the respective preference scores of the multiple groups of question-and-answer pairs can be updated according to the respective fine-tuning offsets of the multiple groups of question-and-answer pairs, and the updated preference scores of the multiple groups of question-and-answer pairs can be obtained. Optionally, for any group of question-and-answer pairs, the preference score of the question-and-answer pair can be superimposed with the fine-tuning offset of the candidate response text in the question-and-answer pair to obtain the updated preference score of the question-and-answer pair.

[0060] Continuing with any candidate response text y i of the instruction text x as an example, the question-and-answer pair composed of the instruction text x and the candidate response text y i is (x, y i ). Suppose, as Figure 2 shown, the preference score output by the preference model for the question-and-answer pair (x, y i ) is r(x, y i ), then the updated preference score R(x, y i ) of the question-and-answer pair (x, y i ) can be expressed by the following formula 2:

[0061]

[0062] Based on the respective fine-tuning offsets of multiple groups of question-and-answer pairs, during the alignment fine-tuning process, the alignment fine-tuning operation of the language model can be constrained so that the deviation of the alignment fine-tuned language model from the initial language model after instruction fine-tuning is small, and thus the alignment fine-tuned language model can retain as much knowledge learned during the instruction fine-tuning stage as possible.

[0063] After updating the preference scores of multiple groups of question-answer pairs based on the above embodiments, the updated preference scores of multiple groups of question-answer pairs can be saved in the empirical data cache pool for aligning and fine-tuning the language model according to the data in the empirical data buffer pool. The following will be described by way of example in combination with different embodiments. It should be noted that the preference scores of the question-answer pairs involved hereinafter refer to the updated preference scores and will not be described separately.

[0064] In some optional Example A1 cases, a reinforcement learning method with policy gradients can be used to perform at least one round of alignment and fine-tuning on the language model.

[0065] Based on the methods described in the foregoing embodiments, the preference scores of multiple groups of question-answer pairs corresponding to multiple instruction texts can be obtained.

[0066] Optionally, in any round of alignment and fine-tuning operation in the alignment and fine-tuning stage, multiple groups of question-answer pairs can be sampled to obtain a first question-answer pair. Herein, the first question-answer pair refers to any group of question-answer pairs obtained by sampling multiple groups of question-answer pairs. The use of "first" to limit the question-answer pairs obtained by sampling is only for convenience of description and distinction and does not constitute a limitation on the order of the question-answer pairs. After sampling the first question-answer pair, the preference loss of the first question-answer pair can be determined according to the preference score of the first question-answer pair. Among them, the preference loss can be a function of the preference score of the first question-answer pair. For example, the preference loss corresponding to any question-answer pair (x, y) can be expressed as:

[0067] L R = f[R(x, y)] Formula 3

[0068] After obtaining the preference loss based on the above embodiments, the language model can be aligned and fine-tuned according to the preference loss and the instruction fine-tuning loss of the initial language model in the instruction fine-tuning stage. Among them, the instruction fine-tuning loss of the initial language model in the instruction fine-tuning stage can be marked as: L ptx , then in the alignment and fine-tuning stage, the total training loss of the language model can be expressed as: L loss = L ptx + f[R(x, y)]. Furthermore, in the alignment and fine-tuning stage, L loss can be used as the policy guidance direction to guide the language model to predict the probability that the word in the vocabulary is the next word in the candidate response text. The above alignment and fine-tuning process can be iteratively executed until the total training loss of the language model converges to a specified range, at which point the iterative training is stopped and the language model after alignment and fine-tuning is output.

[0069] In some optional Example A2Among them, a reinforcement learning method based on the Proximal Policy Optimization (PPO) algorithm can be adopted to perform at least one round of alignment fine-tuning on the language model.

[0070] Among them, the proximal optimization policy algorithm is implemented based on the AC framework. Among them, the AC framework includes a policy model (Actor) and an instruction evaluation model (Critic). Among them, the policy model is used to output actions based on the input state. Among them, the instruction evaluation model is used to output the scores of actions based on the input state. In this embodiment, the states input to the policy model and the instruction evaluation model are specifically: the instruction text. The action output by the policy model based on the input state is specifically: the probability that the word in the vocabulary is the next word in the candidate response text. The score of the action output by the instruction evaluation model based on the input state is specifically: the instruction value score of the instruction text. The following will be specifically described in combination with the AC framework.

[0071] In this embodiment, the language model can be used as the policy model, and the instruction evaluation model corresponding to the language model can be determined. Among them, the instruction evaluation model corresponding to the language model can be initialized from the initial language model, and the model structures and model parameters of the language model and the instruction evaluation model are the same. During the data generation stage described in the foregoing embodiment, each time the language model is fine-tuned, the parameters of the instruction evaluation model are fine-tuned synchronously. In some embodiments, the instruction evaluation model shares the model structure and model parameters with the language model, thereby saving some storage resources.

[0072] Optionally, in the data generation stage, when the instruction text is input into the language model after instruction fine-tuning, the instruction text can be input into the instruction evaluation model corresponding to the language model to obtain the first instruction value score of the instruction text. Among them, the instruction value score of any instruction text is used to evaluate the instruction text in terms of application field, focus, difficulty, and quality. For the convenience of description and distinction, the instruction value score obtained in the data generation stage is described as the first instruction value score, and the first instruction value score of any instruction text x can be marked as V(x).

[0073] Among them, multiple groups of question-and-answer pairs corresponding to multiple instruction texts, the first instruction value scores of multiple instruction texts respectively, and the preference scores of multiple groups of question-and-answer pairs respectively can be saved in the experience data cache pool for aligning and fine-tuning the language model according to the experience data set in the experience data cache pool. As Figure 3 shown, in the experience data cache pool, there are stored the question-and-answer pair (x, y1) of the instruction text x and its preference score R(x, y1), the question-and-answer pair (x, y2) and its preference score R(x, y2), the question-and-answer pair (x, y3) and the positive preference score R(x, y3), and the instruction value score V(x) of the instruction text x.

[0074] When performing alignment fine-tuning on a language model, question-and-answer pairs in the empirical data cache pool can be sampled, and at least one round of alignment fine-tuning operations can be performed based on the relevant data of the sampled question-and-answer pairs. The following will take any round of alignment fine-tuning operation as an example for exemplary illustration.

[0075] Optionally, as Figure 4 shown, multiple groups of question-and-answer pairs in the empirical data cache pool can be sampled to obtain a second question-and-answer pair; where the second question-and-answer pair refers to any group of question-and-answer pairs obtained by sampling multiple groups of question-and-answer pairs. Here, the term "second" is used to limit the question-and-answer pairs obtained by sampling, only for the convenience of description and distinction, and does not impose a limit on the order of the question-and-answer pairs. Among them, the second question-and-answer pair includes a target instruction text and a target candidate response text. For the convenience of subsequent description, the second question-and-answer pair is marked as (x, y i ).

[0076] Optionally, in the alignment fine-tuning stage, the target instruction text can be input into the language model that has undergone at least one round of instruction fine-tuning in the data generation stage to obtain a third response result for the target instruction text, and the target instruction text can be input into the instruction evaluation model that has undergone at least one round of instruction fine-tuning in the data generation stage to obtain a second instruction value score V(x) for the target instruction text ′ , as Figure 4 shown.

[0077] According to the second instruction value score V(x) of the target instruction text ′ and the preference score of the second question-and-answer pair, the instruction quality loss can be determined. Optionally, the instruction quality loss can be determined according to the mean square error of the second instruction value score of the target instruction text and the preference score of the second question-and-answer pair. That is, the instruction quality loss L value can be described by the following formula 4:

[0078] L value = MSE[V(x) ′ , R(x, y i )] Formula 4

[0079] where MSE() represents the mean square error calculation function, V(x) ′ represents the second instruction value score, and R(x, y i ) represents the preference score of the second question-and-answer pair.

[0080] Optionally, the third response result output by the language model for any instruction text may include: a third vocabulary probability matrix. The third vocabulary probability matrix includes at least one probability vector, and each probability vector corresponds to a text position. Among them, the elements in the probability vector corresponding to any text position are used to describe the probability that each word in the vocabulary is predicted by the language model that has undergone at least one instruction fine-tuning in the data generation stage as the output word (i.e., the next output word) corresponding to the text position.

[0081] Among them, according to the third response result of the target instruction text, the third vocabulary probability distribution of the target candidate response text can be obtained, as Figure 4 shown in L3(x, yi). Among them, the third vocabulary probability distribution corresponding to the target candidate response text is used to describe the probability that each word in the target candidate response text is predicted by the current language model as the next output word. According to the third vocabulary probability distribution L3(x, yi) and the first vocabulary probability distribution L1(x, yi) of the target candidate response text, the proximal policy optimization loss can be determined.

[0082] Optionally, when determining the proximal optimization policy loss, the vocabulary probability ratio difference can be determined according to the first vocabulary probability distribution and the third vocabulary probability distribution of the target candidate response text, and the advantage score can be determined according to the preference score of the target candidate response text and the first instruction value score. According to the advantage score and the probability ratio difference, the proximal policy optimization loss L ppo .

[0083] Optionally, the vocabulary probability ratio difference Dr can be described by the following formula 5:

[0084]

[0085] Among them, w j represents the j-th word in the target candidate response text y i in the second Q&A pair, L1(w j ) represents the probability value of the j-th word in the first vocabulary probability distribution, L3(w j ) represents the probability value of the j-th word in the third vocabulary probability distribution, n is the total number of words in the target candidate response text y i , and n is a positive integer.

[0086] Optionally, the advantage score A can be determined according to the difference between the preference score R(x, y i ) of the second Q&A pair and the first instruction value score R(x, y i ), as shown in the following formula 6:

[0087] A = R(x, y i ) - V(x) Formula 6

[0088] Determine the proximal policy optimization loss L based on the advantage score and the probability ratio difference. ppo Among them, the proximal policy optimization loss L ppo can be as shown in the following formula 7:

[0089] L ppo = min(r(x, y i )A, clip(r(x, y i ), 1 - ∈, 1 + ∈)A) Formula 7

[0090] Among them, clip() represents a value clipping function used to limit the magnitude of the policy gradient update, min represents a function to find the minimum value, and ∈ represents a truncation parameter to constrain the gap between two updates of the language model. After determining the instruction quality loss and the proximal policy optimization loss based on the above implementation, the language model can be fine-tuned for alignment at least according to the instruction quality loss, the proximal policy optimization loss, and the instruction fine-tuning loss of the initial language model in the instruction fine-tuning stage. That is, in the alignment fine-tuning stage, the total training loss of the language model can be expressed as: L loss = L ptx + L ppo + L value . Furthermore, in the alignment fine-tuning stage, the language model can be optimized by backpropagation according to L loss . The above alignment fine-tuning operation can be executed multiple times. When the total training loss of the language model converges, stop training and output the language model after alignment fine-tuning.

[0091] Figure 5 is a schematic flowchart of a question-and-answer data processing method provided by an exemplary embodiment of the present application. The method may include steps as Figure 5 shown:

[0092] Step 501: Obtain user instructions in the target application scenario.

[0093] Step 502: Input the user instruction into the language model after alignment fine-tuning to obtain the reply result output by the language model. Among them, the alignment fine-tuning operation of the language model includes: sequentially inputting multiple instruction texts into the language model after instruction fine-tuning to obtain the first reply results of the multiple instruction texts respectively. The first reply result of any instruction text includes: multiple candidate reply texts. The multiple instruction texts and their respective corresponding multiple candidate reply texts form multiple groups of question-answer pairs. Input the multiple groups of question-answer pairs into the preference model to obtain the preference scores of the multiple groups of question-answer pairs respectively. The preference score of any question-answer pair is used to describe the preference information of the instruction text in the question-answer pair for the candidate reply text. Among them, the preference model is trained according to the historical candidate reply texts generated by the language model for multiple historical instructions in the target application scenario and the preference feedback data of the historical candidate reply texts of the multiple historical instructions respectively. Perform at least one round of alignment fine-tuning operation on the language model according to the multiple groups of question-answer pairs and the preference scores of the multiple groups of question-answer pairs.

[0094] The execution subject of this embodiment can be a server or a terminal device. If implemented as a server, the server can obtain the user instruction sent by the terminal device, use the language model after alignment fine-tuning to reply to the user instruction, and return the reply result to the terminal device. If implemented as a terminal device, the terminal device can obtain the user instruction input by the user through a peripheral input device, a voice input device, or a touch output device. The user instruction can be the user's question text or a prompt text. The language model after alignment fine-tuning can be run locally on the terminal device, and after obtaining the user instruction, the language model can be used to reply to the user instruction.

[0095] Among them, the target application scenario refers to the application scenario of the language model in a vertical industry or vertical field. For example, the target application scenario can include: a safety knowledge Q&A scenario, an online customer service Q&A scenario, a knowledge Q&A scenario in online education, etc., which will not be listed one by one. Among them, the optional implementation manner of the alignment fine-tuning operation of the language model can refer to the description of the foregoing embodiment, and will not be described herein again.

[0096] In this implementation manner, the language model for replying to the user instruction is obtained through alignment fine-tuning. In the alignment fine-tuning operation of the language model, the preference model can be used to accurately output the preference score that matches the actual preference information in the target application scenario according to the input question-answer pair, so as to form accurate and reliable guidance for the alignment fine-tuning process, and then effectively improve the professional performance of the language model in the professional field and improve the reliability of the reply result to the user instruction.

[0097] It should be noted that the execution entity of each step of the method provided in the above embodiments can be the same device, or the method can also be executed by different devices. For example, the execution entity of steps 101 to 104 can be device A; for another example, the execution entity of steps 101 and 102 can be device A, and the execution entity of step 103 can be device B; and so on.

[0098] In addition, in some of the processes described in the above embodiments and the accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this document or in parallel. The operation numbers such as 101 and 102 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0099] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0100] Figure 6 Schematically shows the structural diagram of a server provided by an exemplary embodiment of the present application, as Figure 6 shown, the server includes: a memory 601, a processor 602, and a communication component 603.

[0101] The memory 601 is used to store computer programs and can be configured to store various other data to support operations on the server. Examples of these data include instructions for any application or method operating on the server.

[0102] A processor 602, coupled to a memory 601, is configured to execute a computer program in the memory 601 for: sequentially inputting a plurality of instruction texts into a language model after instruction fine-tuning to obtain first response results for each of the plurality of instruction texts; the first response result of any one instruction text includes: a plurality of candidate response texts; the plurality of instruction texts and their respective corresponding plurality of candidate response texts form multiple groups of question-answer pairs; inputting the multiple groups of question-answer pairs into a preference model to obtain preference scores for each of the multiple groups of question-answer pairs; the preference score of any one question-answer pair is used to describe the preference information of the instruction text in the question-answer pair for the candidate response text; wherein, the preference model is trained according to historical candidate response texts generated by the language model for multiple historical instructions in a target application scenario and preference feedback data of the historical candidate response texts of the multiple historical instructions; performing at least one round of alignment fine-tuning operations on the language model according to the multiple groups of question-answer pairs and the preference scores of the multiple groups of question-answer pairs.

[0103] Optionally, before the processor 602 inputs the multiple groups of question-answer pairs into the preference model to obtain preference scores for each of the multiple groups of question-answer pairs, it is further configured to: obtain a plurality of sample data groups; any one sample data group includes: a historical instruction, a plurality of historical candidate response texts of the historical instruction, and preference feedback data of the plurality of historical candidate response texts; wherein, the historical instruction and the plurality of historical candidate response texts form multiple groups of question-answer pair samples; inputting the multiple groups of question-answer pair samples into the preference model to obtain preference scores for each of the multiple groups of question-answer pair samples; determining a preference analysis loss of the preference model according to the preference scores of the multiple groups of question-answer pair samples and the preference feedback data of the plurality of historical candidate response texts; training the preference model according to the preference analysis loss, and stopping training when the preference analysis loss of the preference model converges to a specified range.

[0104] Optionally, the processor 602 is further configured to: during the process of sequentially inputting a plurality of instruction texts in an instruction dataset into a language model after instruction fine-tuning, perform at least one instruction fine-tuning operation on the language model according to set fine-tuning conditions; the fine-tuning conditions include: the time duration from the last instruction fine-tuning operation is a specified time duration, or a specified number of instruction texts are input after the last instruction fine-tuning operation.

[0105] Optionally, before performing at least one round of alignment fine-tuning operations on the language model according to the multiple sets of Q&A pairs and the preference scores of the multiple sets of Q&A pairs, the processor 602 is further configured to: for any one of the multiple instruction texts, input the instruction text into the corresponding initial language model of the language model to obtain a second response result; determine the fine-tuning offset of each set of Q&A pairs corresponding to the instruction text according to the first response result and the second response result of the instruction text; update the preference scores of each set of Q&A pairs corresponding to the instruction text according to the fine-tuning offset of each set of Q&A pairs corresponding to the instruction text, so as to obtain the updated preference scores of each set of Q&A pairs corresponding to the instruction text.

[0106] Optionally, when the processor 602 determines the fine-tuning offset of each set of Q&A pairs corresponding to the instruction text according to the first response result and the second response result of the instruction text, it is specifically configured to: according to the first response result of the instruction text, obtain the first vocabulary probability distribution corresponding to each of the multiple candidate response texts of the instruction text; the first vocabulary probability distribution of any one candidate response text is used to describe the probability that each word in the candidate response text is predicted by the language model as the next output word in the vocabulary; obtain the second vocabulary probability distribution corresponding to each of the multiple candidate response texts of the instruction text from the second response result of the instruction text; the second vocabulary probability distribution corresponding to any one candidate response text is used to describe the probability that each word in the candidate response text is predicted by the initial language model as the next output word in the vocabulary; determine the alignment fine-tuning offset of each of the multiple candidate response texts of the instruction text according to the first vocabulary probability distribution and the second vocabulary probability distribution corresponding to each of the multiple candidate response texts of the instruction text.

[0107] Optionally, when the processor 602 performs at least one round of alignment fine-tuning operations on the language model according to the multiple sets of Q&A pairs and the preference scores of the multiple sets of Q&A pairs, it is specifically configured to: sample the multiple sets of Q&A pairs to obtain a first Q&A pair; determine the preference loss of the first Q&A pair according to the preference score of the first Q&A pair; perform alignment fine-tuning operations on the language model according to the preference loss and the instruction fine-tuning loss of the initial language model in the instruction fine-tuning stage.

[0108] Optionally, the processor 602 is further configured to: when inputting the instruction text into the instruction fine-tuned language model, input the instruction text into the instruction evaluation model corresponding to the language model to obtain a first instruction value score of the instruction text; when performing at least one round of alignment fine-tuning operations on the language model according to the multiple sets of question-and-answer pairs and the preference scores of the multiple sets of question-and-answer pairs, the processor 602 is specifically configured to: according to the multiple sets of question-and-answer pairs, the preference scores of the multiple sets of question-and-answer pairs, and the first instruction value scores of the multiple instruction texts, adopt the proximal policy optimization algorithm to perform at least one round of alignment fine-tuning operations on the language model.

[0109] Optionally, when performing at least one round of alignment fine-tuning operations on the language model according to the multiple sets of question-and-answer pairs, the preference scores of the multiple sets of question-and-answer pairs, and the first instruction value scores of the multiple instruction texts, by adopting the proximal policy optimization algorithm, the processor 602 is specifically configured to: in any round of alignment fine-tuning operations, sample the multiple sets of question-and-answer pairs to obtain a second question-and-answer pair; the second question-and-answer pair includes a target instruction text and a target candidate response text; input the target instruction text into the at least once instruction fine-tuned language model and the instruction evaluation model respectively to obtain a third response result and a second instruction value score of the target instruction text; according to the third response result of the target instruction text, obtain a third vocabulary probability distribution of the target candidate response text; according to the third vocabulary probability distribution and the first vocabulary probability distribution of the target candidate response text, determine the proximal policy optimization loss; according to the second instruction value score of the target instruction and the preference score of the second question-and-answer pair, determine the instruction quality loss; at least according to the instruction quality loss, the proximal policy optimization loss, and the instruction fine-tuning loss of the initial language model in the instruction alignment fine-tuning stage, perform alignment fine-tuning operations on the language model.

[0110] Optionally, when determining the proximal policy optimization loss according to the third vocabulary probability distribution and the first vocabulary probability distribution of the target candidate response text in the second question-and-answer pair, the processor 602 is specifically configured to: determine a vocabulary probability ratio difference according to the first vocabulary probability distribution and the third vocabulary probability distribution of the target candidate response text; determine an advantage score according to the preference score of the target candidate response text and the first instruction value score; determine the proximal policy optimization loss according to the advantage score and the probability ratio difference.

[0111] Optionally, the instruction evaluation model and the language model share a model structure and model parameters.

[0112] In some alternative embodiments, Figure 6The server shown can also be used to execute a question-and-answer data processing method. Specifically, the processor 602 is configured to: obtain a user instruction in a target application scenario through the communication component 603; input the user instruction into the language model after alignment and fine-tuning to obtain a reply result output by the language model. Among them, the alignment and fine-tuning operation of the language model includes: sequentially inputting multiple instruction texts into the language model after instruction fine-tuning to obtain the first reply results of the multiple instruction texts respectively; the first reply result of any instruction text includes: multiple candidate reply texts; the multiple instruction texts and their corresponding multiple candidate reply texts form multiple groups of question-and-answer pairs; inputting the multiple groups of question-and-answer pairs into a preference model to obtain the preference scores of the multiple groups of question-and-answer pairs respectively; the preference score of any question-and-answer pair is used to describe the preference information of the instruction text in the question-and-answer pair for the candidate reply text. Among them, the preference model is trained according to the historical candidate reply texts generated by the language model for multiple historical instructions in the target application scenario and the preference feedback data of the historical candidate reply texts of the multiple historical instructions respectively; perform at least one round of alignment and fine-tuning operation on the language model according to the multiple groups of question-and-answer pairs and the preference scores of the multiple groups of question-and-answer pairs.

[0113] Further, as Figure 6 shown, the server further includes: a power supply component 604 and other components. Figure 6 Only some components are schematically shown, which does not mean that the server only includes Figure 6 the components shown.

[0114] Among them, the memory 601 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc.

[0115] Among them, the communication component 603 is configured to facilitate communication, in a wired or wireless manner, between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as Wi-Fi (Wireless Network Communication Technology), 2G (such as Global System for Mobile Communications (GSM), etc.), 3G (such as Wideband Code Division Multiple Access (WCDMA)), 4G (such as Long Term Evolution (LTE), etc.), 4G+ (such as LTE-Advanced (LTE-A), etc.) or 5G (5th Generation Mobile Communication Technology), or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can be implemented based on technologies such as Near Field Communication (NFC) technology, Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0116] Among them, the power supply component 604 is used to supply power to various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0117] In this embodiment, after obtaining the instruction dataset, multiple instruction texts in the instruction dataset can be sequentially input into the language model fine-tuned by instructions to obtain multiple candidate response texts for each of the multiple instruction texts. The multiple instruction texts and their respective multiple candidate response texts form multiple groups of question-and-answer pairs. The multiple groups of question-and-answer pairs are input into the preference model to obtain the preference scores for each of the multiple groups of question-and-answer pairs. Based on the multiple groups of question-and-answer pairs and the preference scores of the multiple groups of question-and-answer pairs, at least one round of alignment fine-tuning operation can be performed on the language model. In this embodiment, on the one hand, by using the preference model to score the preference of the question-and-answer pairs, the preference scores of the question-and-answer pairs can be automatically and efficiently generated, reducing the dependence on manual analysis operations. On the other hand, the preference model is trained by the question-and-answer data returned from the target application scenario. Therefore, the preference model can accurately output the preference scores that match the actual preference information in the target application scenario according to the input question-and-answer pairs, thus forming an accurate and reliable guidance for the alignment fine-tuning process, and further effectively improving the professional performance of the language model in the professional field.

[0118] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement the steps executable by the server in the above method embodiment.

[0119] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM (Compact Disc Read-Only Memory), optical storage, etc.) containing computer-usable program code.

[0120] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0121] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device that implements the functions specified in one or more of the acts and / or blocks Figure 1 of one or more of the acts and / or blocks Figure 1 specified in one or more of the acts and / or blocks

[0122] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one or more of the acts and / or blocks Figure 1 of one or more of the acts and / or blocks Figure 1 specified in one or more of the acts and / or blocks

[0123] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), an input / output interface, a network interface, and memory

[0124] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory such as read only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium

[0125] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (Parallel Random Access Machine, PRAM), static random access memory (SRAM), dynamic random access memory (Dynamic Random Access Memory, DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technologies, compact disc read only memory (CD-ROM), digital versatile discs (Digital Video Disc, DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves

[0126] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0127] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for model fine-tuning, characterized in that, Including: Sequentially inputting multiple instruction texts into the instruction fine-tuned language model to obtain the first response results of the multiple instruction texts respectively; The first response result of any instruction text includes: multiple candidate response texts; the multiple instruction texts and their corresponding multiple candidate response texts form multiple groups of question-answer pairs; Inputting the multiple groups of question-answer pairs into the preference model to obtain the preference scores of the multiple groups of question-answer pairs respectively; the preference score of any question-answer pair is used to describe the preference information of the instruction text in the question-answer pair for the candidate response text; wherein, the preference model is trained according to the historical candidate response texts generated by the language model for multiple historical instructions in the target application scenario and the preference feedback data of the historical candidate response texts of the multiple historical instructions; Performing at least one round of alignment fine-tuning operation on the language model according to the multiple groups of question-answer pairs and the preference scores of the multiple groups of question-answer pairs.

2. The method according to claim 1, characterized in that, Before inputting the multiple groups of question-answer pairs into the preference model to obtain the preference scores of the multiple groups of question-answer pairs respectively, it further includes: Obtaining multiple sample data groups; any sample data group includes: a historical instruction, multiple historical candidate response texts of the historical instruction, and preference feedback data of the multiple historical candidate response texts; wherein, the historical instruction and the multiple historical candidate response texts form multiple groups of question-answer pair samples; Inputting the multiple groups of question-answer pair samples into the preference model to obtain the preference scores of the multiple groups of question-answer pair samples; Determining the preference analysis loss of the preference model according to the preference scores of the multiple groups of question-answer pair samples and the preference feedback data of the multiple historical candidate response texts; Training the preference model according to the preference analysis loss and stopping training when the preference analysis loss of the preference model converges to a specified range.

3. The method according to claim 1, characterized in that, It further includes: During the process of sequentially inputting multiple instruction texts in the instruction dataset into the instruction fine-tuned language model, performing at least one instruction fine-tuning operation on the language model according to the set fine-tuning conditions; the fine-tuning conditions include: the time duration from the previous instruction fine-tuning operation is a specified time duration, or a specified number of instruction texts are input after the previous instruction fine-tuning operation.

4. The method according to claim 3, characterized in that, Before performing at least one round of alignment fine-tuning operation on the language model according to the multiple groups of question-answer pairs and the preference scores of the multiple groups of question-answer pairs, it further includes: For any instruction text among the multiple instruction texts, inputting the instruction text into the initial language model corresponding to the language model to obtain a second response result; the initial language model is obtained by performing instruction fine-tuning on a pre-trained large language model; Determining the fine-tuning offsets of the multiple groups of question-answer pairs corresponding to the instruction text according to the first response result and the second response result of the instruction text; Updating the preference scores of the multiple groups of question-answer pairs corresponding to the instruction text according to the fine-tuning offsets of the multiple groups of question-answer pairs corresponding to the instruction text to obtain the updated preference scores of the multiple groups of question-answer pairs corresponding to the instruction text.

5. The method according to claim 4, characterized in that, Determine the fine-tuning offsets of each of the multiple pairs of question and answer corresponding to the instruction text according to the first reply result and the second reply result of the instruction text, including: Obtain the first vocabulary probability distribution corresponding to each of the multiple candidate reply texts of the instruction text according to the first reply result of the instruction text; the first vocabulary probability distribution of any candidate reply text is used to describe the probability that each word in the candidate reply text is predicted as the next output word by the language model in the vocabulary. Obtain the second vocabulary probability distribution corresponding to each of the multiple candidate reply texts of the instruction text from the second reply result of the instruction text; the second vocabulary probability distribution corresponding to any candidate reply text is used to describe the probability that each word in the candidate reply text is predicted as the next output word by the initial language model in the vocabulary. Determine the alignment fine-tuning offsets of each of the multiple candidate reply texts of the instruction text according to the first vocabulary probability distribution and the second vocabulary probability distribution corresponding to each of the multiple candidate reply texts of the instruction text.

6. The method according to claim 4, characterized in that, Perform at least one round of alignment fine-tuning operations on the language model according to the multiple pairs of question and answer and the preference scores of the multiple pairs of question and answer, including: Sample the multiple pairs of question and answer to obtain a first pair of question and answer. Determine the preference loss of the first pair of question and answer according to the preference score of the first pair of question and answer. Perform an alignment fine-tuning operation on the language model according to the preference loss and the instruction fine-tuning loss of the initial language model in the instruction fine-tuning stage.

7. The method according to claim 4, characterized in that, Further include: When the instruction text is input into the language model after instruction fine-tuning, input the instruction text into the instruction evaluation model corresponding to the language model to obtain the first instruction value score of the instruction text. Perform at least one round of alignment fine-tuning operations on the language model according to the multiple pairs of question and answer and the preference scores of the multiple pairs of question and answer, including: Perform at least one round of alignment fine-tuning operations on the language model using the proximal policy optimization algorithm according to the multiple pairs of question and answer, the preference scores of the multiple pairs of question and answer, and the first instruction value scores of the multiple instruction texts.

8. The method according to claim 7, characterized in that, Perform at least one round of alignment fine-tuning operations on the language model using the proximal policy optimization algorithm according to the multiple pairs of question and answer, the preference scores of the multiple pairs of question and answer, and the first instruction value scores of the multiple instruction texts, including: During any round of alignment fine-tuning operation, sample the multiple pairs of question and answer to obtain a second pair of question and answer; the second pair of question and answer includes a target instruction text and a target candidate reply text. Input the target instruction text into the language model after at least one instruction fine-tuning and the instruction evaluation model respectively to obtain the third reply result and the second instruction value score of the target instruction text. Obtain the third vocabulary probability distribution of the target candidate reply text according to the third reply result of the target instruction text. Determine the proximal policy optimization loss according to the third vocabulary probability distribution of the target candidate reply text and the first vocabulary probability distribution. Determine an instruction quality loss according to the second instruction value score of the target instruction and the preference score of the second question-and-answer pair; Perform alignment fine-tuning operations on the language model at least according to the instruction quality loss, the proximal policy optimization loss, and the instruction fine-tuning loss of the language model in the instruction alignment fine-tuning stage.

9. The method according to claim 8, characterized in that, Determine the proximal policy optimization loss according to the third vocabulary probability distribution and the first vocabulary probability distribution of the target candidate response text in the second question-and-answer pair, including: Determine a vocabulary probability ratio difference according to the first vocabulary probability distribution and the third vocabulary probability distribution of the target candidate response text; Determine a superiority score according to the preference score of the target candidate response text and the first instruction value score; Determine the proximal policy optimization loss according to the superiority score and the probability ratio difference.

10. The method according to any one of claims 7-9, characterized in that, The instruction evaluation model and the language model share a model structure and model parameters.

11. A method for processing question-and-answer data, characterized in that, Including: Obtain a user instruction in a target application scenario; Input the user instruction into the language model after alignment fine-tuning to obtain a response result output by the language model; wherein, the language model is fine-tuned by using the method according to any one of claims 1-10.

12. A server, characterized in that, Including: A memory and a processor; The memory is used to store one or more computer instructions; The processor is used to execute the one or more computer instructions to: execute the steps in the method according to any one of claims 1-11.

13. A computer-readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it can implement the steps in the method according to any one of claims 1-11.

Citation Information

Cited By

  • Industrial scene dangerous behavior identification method and device based on multi-modal information

    CN121365274A