Preference alignment training method of large language model LLM, electronic equipment and storage medium

Through multiple rounds of self-iteration DPO training methods, the inefficiency and low accuracy problems caused by relying on manual annotation in the existing technology are solved, and a more efficient and stable LLM training effect is achieved.

CN120069082APending Publication Date: 2025-05-30ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510222707.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art relies on manual annotation when training large language model LLM to align human preferences, resulting in low training efficiency and low accuracy and stability of model output.

Method used

Multiple rounds of self-iteration direct preference optimization DPO training method is adopted, sample questions are randomly selected through the preset question bank, sample answers are automatically scored using the scoring model, and training data is constructed based on the scoring results, and the model is optimized round by round.

Benefits of technology

It significantly improves the training efficiency of LLM, reduces labor costs, improves the accuracy and stability of model output, avoids overfitting of the model, and enhances the ability to generalize different types of problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069082A_ABST
    Figure CN120069082A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a preference alignment training method of a large language model LLM, electronic equipment and a storage medium. The training method comprises the following steps: carrying out multiple rounds of self-iteration direct preference optimization (DPO) training on LLM to be trained, and stopping training when a stopping condition is met; wherein for a positive integer i, performing the ith round of training on the i-1-level LLM obtained by the (i-1) th round of training, which comprises the following steps: randomly selecting a sample question from a preset question library, inputting the sample question into the (i-1)-level LLM to obtain a sample answer generated by the model, and scoring the alignment degree of the sample answer and human preference by using a preset scoring model; determining available sample questions from the sample questions according to scoring results of the sample answers, and constructing training data based on the available sample questions and the corresponding available sample answers; and training the (i-1)-level LLM by using the training data to obtain an i-level LLM.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of artificial intelligence, and in particular, to a method for training preference alignment of a large language model (LLM), an electronic device, and a storage medium. Background Art

[0002] In recent years, AI (Artificial Intelligence) technology has developed rapidly. Among them, the LLM (Large Language Model) based on the Transformer architecture has been widely used in dialogue systems in various industries due to its powerful natural language understanding and generation capabilities. In a dialogue system, the LLM can be the object of the user's conversation, that is, the LLM can reason based on the question input by the user and output the corresponding text to answer the question.

[0003] In order to make the text generated (i.e., output) by the LLM fully meet human language and text preferences, related technologies usually adopt reward mechanism-based solutions such as RLHF (Reinforcement Learning from Human Feedback) or Vanilla DPO (Vanilla Direct Preference Optimization) to train the LLM to align with human preferences. However, such solutions usually require manual annotation (such as designing questions and answers before training starts and manually scoring the answers), resulting in relatively low training efficiency of such solutions, as well as the accuracy and stability of the output results of the trained models, which urgently need to be improved. Summary of the Invention

[0004] In view of this, one or more embodiments of this specification provide the following technical solutions:

[0005] According to the first aspect of one or more embodiments of this specification, a method for training preference alignment of a large language model (LLM) is proposed, including:

[0006] Performing multiple rounds of self-iterative direct preference optimization (DPO) training on the LLM to be trained, and stopping the training when the stopping condition is satisfied; wherein, for a positive integer i, performing the i-th round of training on the (i - 1)-th level LLM obtained from the (i - 1)-th round of training includes:

[0007] Randomly selecting a sample question from a preset question bank, inputting the sample question into the (i - 1)-th level LLM to obtain the sample answer generated by the model, and using a preset scoring model to score the degree of alignment between the sample answer and human preferences;

[0008] Determine available sample questions from the sample questions according to the scoring results of the sample answers, and construct training data based on the available sample questions and their corresponding available sample answers;

[0009] Use the training data to train the (i - 1)-level LLM to obtain the i-level LLM.

[0010] According to the second aspect of one or more embodiments of this specification, an electronic device is provided, including: a processor; a memory for storing processor-executable instructions; wherein, the processor realizes the steps of the method as described in the first aspect by running the executable instructions.

[0011] According to the third aspect of one or more embodiments of this specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in the first aspect are realized.

[0012] According to the fourth aspect of one or more embodiments of this specification, a computer program product is provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method as described in the first aspect are realized.

[0013] As can be seen from the above embodiments, this specification proposes an improved DPO training scheme, that is, a multi-round DPO training is performed on the LLM to be trained in a self-iterative manner to obtain the final LLM. Among them, when performing the i-th round of training on the (i - 1)-level LLM obtained from the (i - 1)-th round of training, the sample questions are input into the (i - 1)-level LLM to obtain the sample answers generated by the model, and a preset scoring model is used to score the degree of alignment between the sample answers and human preferences; then available sample questions are determined from the sample questions according to the scoring results of the sample answers, and training data is constructed based on the available sample questions and their corresponding available sample answers; then the training data is used to train the (i - 1)-level LLM to obtain the i-level LLM. It can be understood that in the first round of training (i.e., i = 1), the (i - 1)-level LLM is the LLM to be trained that has not yet started self-iterative training. The so-called self-iterative training means using the sample answers generated by the model completed in the previous round of training (i.e., the (i - 1)-level LLM) to construct training data, and using this training data to perform the next round of training on this model (i.e., performing the i-th round of training on the (i - 1)-level LLM to obtain the i-level LLM).

[0014] It can be seen that this solution not only requires multiple rounds of self-iterative DPO training for the LLM to be trained, but also the specific process of each round of DPO training is different from that of traditional DPO (i.e., the aforementioned naive DPO): during the i-th round of self-iterative DPO training in this solution, a preset scoring model is used to automatically score the sample answers, rather than manually scoring in advance before training. Then, appropriate available sample questions and available sample answers are selected according to the scoring results to construct training data to complete this round of training. Obviously, the self-iterative DPO training process of this solution does not rely on manual annotation, significantly improving the training efficiency and reducing the labor cost during the preference alignment training of the LLM; in addition, high-quality training data can be continuously introduced through multiple rounds of self-iterative DPO training, making the model output have high accuracy (i.e., being able to fully align with human preferences) and stability; moreover, it can avoid the model from overfitting specific patterns in the training data, enabling it to learn more comprehensive distribution characteristics, thereby improving the generalization ability of the model to different types of questions, and the output results can comprehensively handle various types of practical problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a schematic diagram of the architecture of a model management system provided by an exemplary embodiment.

[0016] Figure 2 is a flowchart of a method for preference alignment training of a large language model LLM provided by an exemplary embodiment.

[0017] Figure 3 is a flowchart of another method for preference alignment training of a large language model LLM provided by an exemplary embodiment.

[0018] Figure 4 is a schematic diagram of the structure of a device provided by an exemplary embodiment.

[0019] Figure 5 is a block diagram of a device for preference alignment training of a large language model LLM provided by an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To solve the aforementioned technical problems in the related art, this specification proposes a training solution for the LLM, aiming to reduce the dependence on manual annotation in the training process by automatically scoring the model output through a scoring model, thereby improving the training efficiency of the model and the inference accuracy and stability of the trained model. The following describes this solution in combination with the accompanying drawings and related embodiments.

[0021] Figure 1 is a schematic diagram of the architecture of a model management system provided by an exemplary embodiment. As Figure 1As shown, the system may include several servers, such as server 11 and server 12; and several electronic devices, such as PC (Personal Computer) 13, PC 14, and PC 15, etc. It should be noted that Figure 1 This is to illustrate as comprehensively as possible the possible implementation manners of this solution. Therefore, the system drawn includes the above-mentioned multiple devices.

[0022] Any one of server 11 and server 12 can be a physical server containing an independent host, or it can also be a virtual server hosted by a host cluster. During operation, the server can run the server-side program of a certain application to implement the related functions of the application. For example, when server 11 runs the program of the model training service, it can be implemented as the server side of the model management service, such as cooperating with the client of the model training service and invoking relevant resources to train the LLM; and when server 12 runs the program of the model invocation service, it can be implemented as the server side of the model invocation service, such as being used to deploy and run the final LLM trained by server 11, and providing the model invocation service to the client of the model invocation service (such as opening an API interface for the client to call the final LLM), which will not be elaborated here.

[0023] PC is only one type of electronic device that users can use. In fact, users can obviously also use electronic devices of the following types: mobile phones, tablet devices, laptop computers, personal digital assistants (PDAs), wearable devices (such as smart glasses, smart watches, etc.). One or more embodiments of this specification do not limit this. During operation, the electronic device can run the client-side program of a certain application to implement the related functions of the application. For example, when PC 13 runs the program of the model training service, it can be implemented as the client of this service. At this time, the user (such as a technical person participating in model training) can interact with the server side of the model training service running in server 11 through this client to control the server to complete the training process of the to-be-trained LL through initiating relevant instructions. The final LLM after training can be deployed on the aforementioned server 12 to participate in building a dialogue system after running. Another example is that when PC 14 or PC 15 runs the program of the model invocation service, it can be implemented as the client of this service. At this time, the user (such as an ordinary user who has a dialogue with the final LLM through the dialogue system) can interact with the server side of the model invocation service running in server 12 through this client to call the final LLM to complete the dialogue by initiating relevant instructions.

[0024] Among them, any of the above-mentioned client-side programs can be started and run on the corresponding PC. The client-side program can be a native application installed on the PC, or the client-side program can be a MiniProgram, a Quick APP, or other similar forms. Of course, when using web technologies such as HTML5 or similar, relevant functions can be implemented through the page displayed by the browser. Here, the browser can be an independent browser application or a browser module embedded in some applications.

[0025] It should be noted that the preference alignment training method for the LLM described in the embodiments of this specification can be applied to a model training system (i.e., the system serves as the execution entity of this method). From a software perspective, the system can include the server side of the aforementioned model training service; from a hardware perspective, the system can include any hardware device that runs the server side (such as the aforementioned server 11, etc.). Of course, during the model training process, the relevant resources (such as network, storage, computing, etc.) called by the system can belong to the aforementioned hardware device itself, or can also belong to other devices associated with the hardware device (of course, corresponding usage permissions may need to be opened to the hardware device), which will not be elaborated here.

[0026] The network 10 for the interaction between electronic devices such as PC13 and servers such as server 11 can be specifically selected to implement communication using wired or wireless networks based on the communication methods supported by the corresponding electronic devices. This specification does not limit this.

[0027] Figure 2 is a flowchart of a preference alignment training method for an LLM provided by an exemplary embodiment. As Figure 2 shown, the method can include the following step 200.

[0028] Step 200, perform multiple rounds of self-iterative direct preference optimization (DPO) training on the LLM to be trained, and stop training when the stop condition is met.

[0029] The LLM to be trained is the training object of the training solution described in this specification. Therefore, the model needs to be obtained first before the solution is executed.

[0030] In one embodiment, the basic LLM that has been pre-trained and fine-tuned using large-scale text data (such as books, web pages, papers, etc.) can be directly used as the LLM to be trained. Such basic models usually already possess basic language modeling capabilities, that is, they can perform simple language understanding and text generation; in addition, the model can also master some general knowledge (such as grammar, vocabulary, factual information, etc.) and can perform basic context understanding (that is, it can understand the logical relationships and semantics in the text). Among them, there are a large number of basic LLMs available in the industry. Directly using the basic LLM as the LLM to be trained can effectively simplify the model acquisition and preprocessing process.

[0031] In another embodiment, the basic LLM can also be cold-start trained using basic preference data to obtain the LLM to be trained. Among them, the basic preference data can be obtained through manual annotation, or it can also be generated by preset rules or other reliable LLMs and screened through scoring. In short, the basic preference data should have reliable data quality, that is, the data should be able to reflect human real language and text preferences as much as possible. In addition, the basic preference data should also have a large amount of data and cover as many task types, technical fields, language styles and other scenarios as possible to avoid preference biases that may be caused by a single scenario. Using the LLM trained with the above basic preference data as the LLM to be trained helps to ensure that the model can be initially aligned with real human preferences, so that it has a high convergence speed, helps to reduce the number of iterations of subsequent self-iterative training, and shortens the time-consuming of self-iterative training.

[0032] In addition, the cold-start training of the basic LLM using basic preference data can refer to the Vanilla DPO training method in related technologies, which will not be elaborated here.

[0033] After obtaining the LLM to be trained, the model training system will perform multiple rounds of self-iterative DPO training (or self-iterative optimization, or self-training Self-training) on the model, and monitor whether the corresponding stop conditions are met during the training process (such as judging whether the current model meets the stop conditions after each round of training), and stop the training when the conditions are met, and use the current model as the result of self-iterative training. Among them, in any round of self-iterative training process, the high-quality output results of the model obtained in the previous round of self-iterative training need to be used to construct the training data required for this round, so as to effectively alleviate the distribution shift problem between the training data and the LLM through multiple rounds of optimization. The specific manner of the stop condition can be seen in the following embodiments, which will not be elaborated here.

[0034] Without loss of generality, assume that the LLM to be trained stops training after N rounds of self-iterative training (e.g., the model obtained in the Nth round of training satisfies the stop condition). In the following embodiments, the model after the i-th round of training is referred to as the i-level LLM (where i is a positive integer). Specifically, when i = 1, the model before the start of the first round of training (i.e., the LLM to be trained) is referred to as the 0-level LLM, and the model obtained after the completion of the first round of training (performed on the 0-level LLM) is referred to as the 1-level LLM; when i = N, the model obtained after the completion of the Nth round of training (performed on the (N - 1)-level LLM) is referred to as the N-level LLM, and this model is regarded as the final LLM obtained after the completion of N rounds of self-iterative training.

[0035] Among them, for the i-th round of training, the self-iterative DPO (i.e., Iterative DPO) scheme proposed in the embodiments of this specification can be used for training. Specifically, the i-th round of training can be performed on the (i - 1)-level LLM obtained in the (i - 1)-th round of training according to the following steps 202 to 206 to obtain the i-level LLM.

[0036] Step 202: Randomly select a sample question from a preset question library, input the sample question into the (i - 1)-level LLM to obtain the sample answer generated by the model, and use a preset scoring model to score the degree of alignment between the sample answer and human preferences.

[0037] During the self-iterative training process, the LLM being trained is used as the dialogue object for interaction, that is, the sample question is input into the model, and the sample answer generated by the model after reasoning for this question is received. It can be understood that since the training object of this scheme is the LLM, both the sample question and the sample answer are text-format data and do not include multimedia data such as images, audio, and video.

[0038] A question library (or question pool) containing multiple questions can be created in advance. Based on this, the model training system can randomly select a sample question from the question library (i.e., use the question selected from the question library as the sample question) and input the selected sample question into the (i - 1)-level LLM. Among them, the number of sample questions can be multiple. At this time, each sample question can be input into the (i - 1)-level LLM respectively, and the sample answers generated by the model for each sample question can be obtained. It can be understood that by randomly selecting sample questions from the database, it can be ensured as much as possible that the selected sample questions are evenly distributed in the database, so as to cover more comprehensive scenarios. In addition, a large number of questions covering more scenarios should be added to the question library so that the sample questions selected in the subsequent training process can cover more scenarios to avoid distribution shift during the training process.

[0039] It should be noted that since the self-iterative training process uses the sample answers generated by the model after reasoning on the sample questions to construct the training data for this round, the applicable scenarios of the final model can be guided by the sample questions. For example, in the multi-round self-iterative DPO training process, questions in a specific scenario can be selected as sample questions to ensure that the final LLM obtained after training stops can be applicable to that scenario. Exemplarily, if the sample questions in the online shopping scenario are used in N rounds of training, the final LLM can be applicable to the online shopping scenario; if the sample questions in the education and training and frontier science scenarios are used in N rounds of training, the final LLM can be applicable to the education and training and frontier science scenarios, etc. The specific scenario can also be any one or more of medical and health, fashion design, history and culture, news media, and the embodiments of this specification do not impose any restrictions.

[0040] The purpose of training the LLM in this solution is to make the final LLM after training more accurately align with human preferences. Therefore, in the i-th round of training, it is necessary to score the alignment degree between the sample answers and human preferences, and select samples based on this score to construct the training data for this round, so as to take the alignment of the output text with human preferences as the training direction of the model. Among them, the human preferences described in this specification refer to the language and text preferences of humans. For example, Chinese characters need to conform to Chinese grammar and basic logical relationships, and should be as concise and accurate as possible, etc., which will not be elaborated here. The human preferences can be reflected through various forms of scoring, as detailed in the following embodiments.

[0041] In one embodiment, the sample questions can be input into the (i - 1)-level LLM to obtain multiple sample answers (i.e., the generated text content) output by the model for each sample question, and a preset scoring model can be used to score each sample answer respectively. Among them, the input interface of the (i - 1)-level LLM can be called to input each sample question one by one or input all sample questions in batches. For each input sample question, the (i - 1)-level LLM can output multiple sample answers for the model. All the sample answers corresponding to all the sample questions can be regarded as candidate answers - the available sample answers for constructing the training data for this round are selected from these answers according to the scoring results; correspondingly, all the sample questions for this round can be regarded as candidate questions - the available sample questions for constructing the training data for this round are selected from these questions according to the scoring results. Suppose m questions are selected in this round, and the (i - 1)-level LLM outputs k sample answers for each of these questions, then a total of k * m sample answers can be obtained as candidate answers; correspondingly, the m sample questions become candidate questions.

[0042] For each sample answer generated by the i-1 level LLM, a preset scoring model can be used to score each sample answer respectively, so as to select appropriate available sample answers according to the scoring results of each scheme. Among them, the scoring result of any sample answer can be used to reflect the quality of the result. For example, the quality level can be reflected by the size of the score (that is, the degree of alignment with human preferences). In this way, it can be ensured that high-quality sample answers are selected as available sample answers for constructing training data in the subsequent process according to the scoring results, thereby improving the training efficiency of this round. Among them, the quality of the sample answer can be evaluated from at least one dimension (that is, there may be one or more evaluation indicators for the answer quality), and the quality of different dimensions can be represented by different scores, and different scores can be calculated in different ways.

[0043] In one embodiment, from the perspective of the scoring method, a preset scoring rule can be used to score the degree of alignment between the sample answer and human preferences; and / or, the sample answer can also be input into multiple alignment evaluation models respectively, so as to use each alignment evaluation model to score the degree of alignment between the sample answer and human preferences respectively. Among them, the multiple alignment evaluation models can be other LLMs different from the LLM to be trained, and each alignment evaluation model is different from each other.

[0044] In another embodiment, from the perspective of the manifestation form of the scoring result, the scoring result of the sample answer corresponding to each sample question includes the rationality score of each sample answer calculated comprehensively from multiple evaluation dimensions; it can also include, when any sample answer contains the formatted text corresponding to the standard format, the format accuracy score of the sample answer used to characterize the degree to which the formatted text conforms to the standard format; and / or, it can also include the answer correct rate of the sample answer corresponding to the sample question, where the answer correct rate is the proportion of the number of correct sample answers among all sample answers corresponding to the sample question, etc., as described in the following embodiments.

[0045] In one embodiment, scoring can be performed using a scoring rule. For example, a preset scoring rule can be used to calculate the accuracy score of the sample answer, and this score can be used to characterize the accuracy degree of the sample answer (corresponding to the accuracy dimension). Exemplarily, the accuracy score can be positively correlated with the accuracy degree of the sample answer, that is, the higher the accuracy score of any sample answer, the more accurate the answer; conversely, the lower the accuracy score of any sample answer, the less accurate the answer. Through the accuracy score, it is possible to more accurately evaluate whether the sample answer is accurate.

[0046] Among them, the accuracy dimension can also be divided into sub-dimensions such as correctness (or content accuracy) and format accuracy. As mentioned above, the sample answer is data in text format. Content accuracy refers to the accuracy of the text content in the sample answer, which is used to evaluate whether the text generated by the current model (i.e., the i-1 level LLM, the same below) can correctly answer the corresponding sample question; while format accuracy refers to the degree of compliance (or following degree) of the text organization form in the sample answer with a specific data format, which is used to evaluate whether the model can output text that conforms to a specific data format.

[0047] Exemplarily, the answering correct rate of the sample answer can be calculated using a preset first scoring rule. The answering correct rate of the sample answer corresponding to any sample question can be the proportion of the number of correct sample answers among all sample answers corresponding to that sample question (for example, if any sample question corresponds to 8 sample answers, among which 6 are correct answers and 2 are wrong answers, then the answering correct rate of the sample question corresponding to this sample question is 6 / 8 = 75%). Obviously, the answering correct rate can be used to characterize the content accuracy of all sample answers corresponding to that sample question (that is, the higher the answering correct rate, the higher the content accuracy). For example, the true answer to the sample question can be determined by querying an external trusted knowledge base, and whether each sample answer is a correct answer or a wrong answer can be determined by comparing the semantics of each sample answer with the true answer. Among them, the semantic similarity between any sample answer and the true answer can be calculated (such as cosine similarity or semantic similarity, etc.) to determine whether their semantics are consistent. If the semantic similarity is greater than 80%, it is considered that their semantics are consistent, and then it is determined that the sample answer is a correct answer; otherwise, it is a wrong answer.

[0048] In the case where any sample answer contains formatted text corresponding to the standard format, the second scoring rule associated with the standard format can be used to calculate the format accuracy score of the sample answer, which is used to characterize the degree to which the formatted text conforms to the standard format. Among them, the second scoring rule may include formulas for calculating the format accuracy score based on various factors such as the grammatical correctness of the text in the sample answer, whether the text length meets the requirements, and whether it contains necessary keywords or information points. In addition, the standard format can be any format such as JSON (JavaScript Object Notation), XML (Extensible Markup Language), Excel (spreadsheet), etc. Taking the JSON format as an example, the second scoring rule corresponding to this format can record the formula for calculating the score when the text format (such as the data structure of the text, key-value pairs in the text, key names, value types, string formats, etc.) does not conform to the JSON format. For example, the more the number of texts that do not conform to the JSON format, the smaller the calculated format accuracy score. Through the above method, the appropriate scoring rule can be selected according to the actual situation of the sample answer to calculate the accuracy score of the answer, which is used to evaluate the accuracy of the i-1 level LLM's answer to the sample question.

[0049] It should be noted that the accuracy score of the sample answer can only include the above answer correct rate, or only include the above format accuracy score, or can also include both the above answer correct rate and format accuracy score, or can also be calculated based on the above answer correct rate and format accuracy score. For example, for any sample question, the weighted average of the answer correct rate of all sample answers corresponding to this question and the format accuracy scores of each sample answer can be calculated as the accuracy score of all sample answers corresponding to this question - at this time, this score can comprehensively characterize the accuracy of all sample answers in terms of content and format, and the evaluation of the sample answer is more comprehensive.

[0050] In another embodiment, scoring can also be performed using an alignment evaluation model, that is, inputting the sample answers into multiple alignment evaluation models respectively, and calculating the rationality score (corresponding to the rationality dimension) of the sample answers by using the scores generated by each alignment evaluation model respectively. For example, calculate the weighted average of the scores generated by each alignment evaluation model for the same sample answer, and use it as the rationality score of the sample answer. Among them, the multiple alignment evaluation models can be other LLMs different from the to-be-trained LLM, and each alignment evaluation model is different from each other, that is, the rationality of the sample answers output by the i-1 level LLM is evaluated by multiple different other LLMs respectively (that is, using the multiple alignment evaluation models as "judges" to evaluate the output rationality of the i-1 level LLM). Of course, each of the above alignment evaluation models can be pre-trained to output a rationality score in scalar form (such as a floating-point number) according to the input. Usually, different LLMs often have different understandings and emphases on text semantics and human preferences. Therefore, it is very difficult for the rationality scores obtained by the multiple alignment evaluation models for the same sample answer to be exactly the same. Therefore, using the multiple alignment evaluation models to evaluate the rationality scores of the sample answers respectively can avoid the understanding and evaluation biases that may be brought by a single LLM and avoid reward hacking (that is, the reward model is hijacked by irrelevant or weakly relevant features in the sample, resulting in the reward score no longer being able to correctly model the sample quality), and comprehensively and elaborately evaluate the rationality of the current model output results as much as possible.

[0051] In addition, the scores generated by at least some of the multiple alignment evaluation models are used to characterize the rational degree of the sample answers under multiple evaluation dimensions. For example, each model among the multiple alignment evaluation models can comprehensively evaluate the sample answers from multiple evaluation dimensions such as usefulness (whether it can effectively answer the sample answers, whether it answers off-topic), correctness (whether the text semantics is correct), coherence (whether the text and punctuation marks are semantically coherent), complexity (whether the semantic logic is easy to understand), and redundancy (whether the text content is concise, whether it contains repeated content), and give the corresponding rationality scores. In this way, the rationality scores generated by the alignment evaluation model can be used to comprehensively characterize the rational degree of the sample answers output by the i-1 level LLM under multiple evaluation dimensions. Of course, the alignment evaluation model can also directly output the rationality score components of the sample answers under the above multiple evaluation dimensions. At this time, the model training system can calculate the rationality score of the sample answers according to the above components, such as using the weighted average of each component as the rationality score, etc. In this way, the rationality evaluation of the sample answers can be split into more fine-grained multiple dimensions, so that the rationality score of the sample answers can more comprehensively reflect the rational degree of the answers, thus being more in line with human real preferences and helping to improve the accuracy of the model.

[0052] The foregoing embodiments are described by taking the rationality scores of the sample answers output by each alignment evaluation model respectively as an example. In fact, in order to simplify subsequent operations, it is also possible to calculate the average value of the scores generated by each alignment evaluation model respectively (such as weighted average or arithmetic mean, etc.), and use it as the rationality score of the sample answer. In this way, each sample answer has only one rationality score, which helps to reduce the computational complexity of subsequent processing.

[0053] In addition, for any sample question and its corresponding sample answers, it is possible to calculate the weighted average of the accuracy scores of these sample questions and the rationality scores of each sample answer (this value can be regarded as the comprehensive score of the sample answer), and use it as the scoring result of the sample answer corresponding to this sample question.

[0054] Step 204, determine available sample questions from the sample questions according to the scoring result of the sample answer, and construct training data based on the available sample questions and their corresponding available sample answers.

[0055] After scoring the alignment degree between the sample answer and human preference using a preset scoring model, available sample questions can be determined from the sample questions according to the corresponding scoring result. As mentioned above, after inputting a sample question into the i-1 level LLM, the model will output multiple sample answers; therefore, if any sample question is determined to be an available sample question, its corresponding available sample answers can be determined from all its corresponding sample answers; correspondingly, if any sample answer is determined to be an available sample answer, its corresponding sample question becomes an available sample question.

[0056] In one embodiment, the higher the answer correct rate of the sample answers corresponding to any sample question indicates that these answers are generally more accurate; on the contrary, the lower the answer correct rate indicates that these answers are generally less accurate. Based on this, when the scoring result includes the answer correct rate, available sample answers can be determined from the sample questions according to the answer correct rate of the sample answer. For example, available sample questions with the answer correct rate of the corresponding sample answers not lower than the first threshold and not higher than the second threshold can be determined from the sample questions. Exemplarily, if the value range of the answer correct rate is [0, 100%], sample questions with the answer correct rate in the interval of [30%, 70%] (at this time, the first threshold is 30% and the second threshold is 70%) can be screened out from all sample questions according to the size of the answer correct rate of the corresponding sample answers, and these questions are determined to be available sample questions.

[0057] It is understandable that since the correct answer rate of the sample answers corresponding to any sample question is positively correlated with the accuracy of these answers, if the correct answer rate of the sample answers corresponding to a certain sample question is lower than the first threshold, it indicates that the current model (i.e., the i-1 level LLM, the same below) is still unable to stably and correctly answer this question (that is, the difficulty of this question significantly exceeds the dialogue ability of the current model, and this question is too difficult for the model); and if the correct answer rate of the sample answers corresponding to a certain sample question is higher than the second threshold, it indicates that the current model can stably answer this question correctly from multiple perspectives (that is, the difficulty of this question is significantly lower than the dialogue ability of the current model, and this question is too easy for the model). Therefore, the difficulty of these two types of questions does not match the dialogue ability of the current model. And the sample questions with a correct answer rate not lower than the first threshold and not higher than the second threshold are exactly the questions with moderate difficulty (not too difficult and not too easy). Selecting these questions as the available sample questions for this round (i.e., the i-th round) of training can further improve the dialogue ability of the current model, thereby continuously strengthening the dialogue performance of the model and making it gradually approach the optimal preference alignment effect.

[0058] Alternatively, it is also possible to first screen out high-quality questions with higher rationality according to the rationality score (for example, when the rationality score is positively correlated with the reasonable degree of the sample answer, determining the sample questions corresponding to the sample answers with a rationality score not lower than the preset threshold as high-quality questions), and then further screen out questions with moderate difficulty from these high-quality questions as the available sample questions. By this method, it can be ensured that the finally determined available sample questions have high rationality and moderate difficulty, which helps to improve the convergence speed and final performance of the model.

[0059] In one embodiment, the accuracy score of the aforementioned sample answer can be positively correlated with its accuracy (that is, the larger the accuracy score of the sample answer, the more accurate it is), and the rationality score can be positively correlated with its reasonable degree (that is, the higher the rationality score of the sample answer, the more reasonable it is). Based on this, when constructing training data based on the available sample questions and their corresponding available sample answers, one piece of training data can be created for each available sample question.

[0060] For example, given that the correct answer rate of the sample answers corresponding to any sample question is positively correlated with the accuracy of these answers, when the aforementioned first threshold is not zero and the second threshold is not 100%, each available sample question determined in the aforementioned manner must correspond to at least one correct answer and at least one wrong answer. Based on this, when the scoring result includes the rationality score and the correct answer rate, and the rationality score is positively correlated with the rationality of the sample answer, for each available sample question and its corresponding sample answers: positive sample answers with a rationality score not lower than the third threshold can be determined from the correct answers included in the sample answers, and negative sample answers with a rationality score not higher than the fourth threshold can be determined from the wrong answers included in the sample answers, where the third threshold is greater than the fourth threshold; then a training data can be constructed with this available sample question (prompt), the positive sample answer (chosen), and the negative sample answer (rejected). From the size relationship between the above scores and thresholds, it can be seen that the positive sample answer is an answer that is correctly and reasonably answered (that is, the answer is a correct answer and has a high degree of rationality), and is used as the positive sample for this round of training; while the negative sample answer is an answer that is wrongly and unreasonably answered (that is, the answer is a wrong answer and has a low degree of rationality), and is used as the negative sample for this round of training. When constructing the training data, a new prompt can be generated from the available sample question, the positive sample answer, and the negative sample answer according to a preset template, as the input data for the i-1 level LLM. In this way, positive sample answers and negative sample answers can be determined from all sample answers according to the scoring results of each sample answer, and together with the corresponding available sample questions, they form training data for this round of training.

[0061] It should be noted that the previous embodiments are described by taking the accuracy score being positively correlated with the accuracy of the sample answer and / or the rationality score being positively correlated with the rationality of the sample answer as examples. In fact, when implementing the solution, it is also possible to control the accuracy score to be negatively correlated with the accuracy of the sample answer and / or the rationality score to be negatively correlated with the rationality of the sample answer by setting the scoring rules and pre-training the alignment evaluation model. The embodiments of this specification do not limit this. In addition, each of the aforementioned thresholds can be flexibly selected according to actual situations such as training parameters, resource status, and index requirements, and will not be elaborated here.

[0062] Step 206, training the i-1 level LLM with the training data to obtain the i level LLM.

[0063] After constructing the training data for this round in the aforementioned manner, the current model can be trained with this training data.

[0064] Since this solution does not require manual annotation during the intermediate process of self-iterative training, the model training process is relatively controllable, which also helps to reasonably schedule training resources such as the CPU (Central Processing Unit), GPU (Graphics Processing Unit), network bandwidth, database / memory, etc., thereby optimizing resource allocation.

[0065] In one embodiment, compared with the resources occupied by the process of performing the i-th round of training using the training data in step 206, the resources occupied by the processes of scoring the sample answers, selecting available sample questions and available sample answers according to the scoring results, and constructing the training data of this round in the foregoing steps 202 to 204 are usually less. However, the execution of steps 202 to 204 often takes a certain amount of time. If all the resources required for the i-th round of training wait for the construction of the target data to be completed during this period, the overall utilization efficiency of the resources of the model training system may be reduced. In response to this, the expected time consumption for constructing the training data can be predicted at least based on the data volumes of the sample questions and the sample answers, and the training resources can be allocated for other tasks (different from the i-th round of training) to use before the start time of the i-th round of training corresponding to the expected time consumption; and the training resources can be allocated (or scheduled) to execute the i-th round of training again during the start time or the training preparation duration before this time (this duration is less than the execution duration of steps 202 to 204). As mentioned above, every time a self-iterative training is performed, steps 202 to 206 need to be executed once. By this means, during the process of each self-iterative training, steps 202 - 204 can be executed using fewer resources, and the remaining training resources can be allocated to other tasks to use; only all available training resources are occupied when executing step 206, thereby achieving seamless connection and reasonable arrangement of software and hardware resources through the above scheduling, which helps to improve the overall utilization efficiency of training resources.

[0066] During the process of self-iterative training, the model training system can detect in real time or periodically whether the stop condition is met, and stop training when it is met. For example, a preset number of iterations can be adopted (such as presetting the number of iterations to the aforementioned N), the loss function converges (such as calculating the loss function of the current model after each round of training, and considering the stop condition to be met when the loss Loss converges to an acceptable range), the performance of the validation model reaches the expectation (such as using the validation set to validate the model after each round of training, and stopping training when the validation results show that the model performance parameters reach the preset metrics), the resources are exhausted (stopping training when the available training resources are exhausted), etc. Taking the loss function as an example, the Loss can be equal to the deviation between the scoring result of the sample answer in the i-th round of training and the scoring result of the sample answer in the (i - 1)-th round (or the (i - k)-th round, where k is an integer greater than 1) of training. At this time, when this deviation is less than the threshold, it can be determined that the stop condition is met and training is stopped. Of course, the above threshold or k value, etc. can be reasonably set during the implementation of the solution, and the embodiments of this specification do not limit this.

[0067] As can be seen from the above embodiments, this specification proposes an improved DPO training scheme, that is, taking the LLM to be trained as the 0-level LLM and performing multiple rounds of self-iterative DPO training on it to obtain the final LLM. Among them, when performing the i-th round of training on the (i - 1)-level LLM, the sample question is input into the (i - 1)-level LLM to obtain the sample answer generated by the model, and a preset scoring model is used to score the alignment degree between the sample answer and the human preference; then, according to the scoring result of the sample answer, the available sample questions are determined from the sample questions, and training data is constructed based on the available sample questions and their corresponding available sample answers; then, the (i - 1)-level LLM is trained using the training data to obtain the i-level LLM. It can be understood that in the first round of training (i.e., i = 1), the (i - 1)-level LLM is the LLM to be trained that has not yet started self-iterative training. And the so-called self-iterative training means using the sample answers generated by the model completed in the previous round of training (i.e., the (i - 1)-level LLM) to construct training data, and using this training data to perform the next round of training on this model (i.e., performing the i-th round of training on the (i - 1)-level LLM to obtain the i-level LLM).

[0068] It can be seen that this solution not only requires multiple rounds of self-iterative DPO training for the LLM to be trained, but also the specific process of each round of DPO training is different from that of traditional DPO (i.e., the aforementioned naive DPO): during the i-th round of self-iterative DPO training in this solution, a preset scoring model is used to automatically score the sample answers, rather than manually scoring in advance before training, and then appropriate available sample questions and available sample answers are selected according to the scoring results to construct training data to complete this round of training. Obviously, the self-iterative DPO training process of this solution does not need to rely on manual annotation, significantly improving the training efficiency and reducing the labor cost in the process of preference alignment training for the LLM; in addition, high-quality training data can be continuously introduced through multiple rounds of self-iterative DPO training, making the model output have high accuracy (i.e., being able to fully align with human preferences) and stability; moreover, it can avoid the model overfitting specific patterns in the training data, enabling it to learn more comprehensive distribution characteristics and thus improving the generalization ability of the model to different types of questions, and the output results can comprehensively handle various types of actual problems.

[0069] Figure 3 is a flowchart of another method for preference alignment training of a large language model (LLM) provided by an exemplary embodiment. As Figure 3 shown, the method may include the following steps 302-304:

[0070] Step 302, cold-start training the basic LLM using basic preference data to obtain the LLM to be trained.

[0071] Step 304, perform multiple rounds of self-iterative DPO training on the LLM to be trained (multiple overlapping rectangles in the figure represent multiple rounds of self-iterative DPO training), and stop training when the stop condition is met.

[0072] Among them, in step 304, when performing the i-th round of training on the (i-1)-th level LLM, it can be carried out according to the following steps 3042-3046:

[0073] Step 3042: Randomly select sample questions and input them into the (i-1)-th level LLM, and obtain multiple sample answers generated by the model for each sample question respectively;

[0074] Step 3044: Use a preset scoring model to score each sample answer respectively, such as calculating the accuracy score and / or rationality score of each sample answer, etc.; and, select available sample answers from each sample answer according to the scoring results, and determine the corresponding sample questions as available sample questions;

[0075] Step 3046: Construct the training data for this round according to the available sample questions and available sample answers, and use the training data to train the (i-1)-th level LLM to obtain the i-th level LLM.

[0076] For the specific implementation methods of the above steps, reference can be made to the foregoing embodiments, which will not be elaborated herein.

[0077] Figure 4 It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 4 , at the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, a memory 408, and a non-volatile memory 410. Of course, it may also include other hardware required for other functions. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into the memory 408 and then runs it. Of course, in addition to the software implementation method, one or more embodiments of this specification do not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logical device.

[0078] Please refer to Figure 5 , the preference alignment training device of the large language model LLM can be applied to the device as shown in Figure 4 to implement the technical solutions of this specification. Among them, the training device may include:

[0079] An iterative training unit 501, configured to perform multiple rounds of self-iterative direct preference optimization (DPO) training on the LLM to be trained, and stop training when the stop condition is met; where, for a positive integer i, the iterative training unit 501 includes:

[0080] An answer scoring unit 5011, configured to randomly select sample questions from a preset question bank, input the sample questions into the (i - 1)-level LLM to obtain the sample answers generated by the model, and use a preset scoring model to score the alignment degree between the sample answers and human preferences;

[0081] A data construction unit 5012, configured to determine available sample questions from the sample questions according to the scoring results of the sample answers, and construct training data based on the available sample questions and their corresponding available sample answers;

[0082] A model training unit 5013, configured to train the (i - 1)-level LLM using the training data to obtain the i-level LLM.

[0083] Optionally, the answer scoring unit 5011 is specifically configured to:

[0084] Calculate the accuracy score of the sample answer using a preset scoring rule; and / or,

[0085] Input the sample answers into multiple alignment evaluation models respectively, and calculate the rationality score of the sample answers by using the scores generated by each alignment evaluation model respectively.

[0086] Optionally, the answer scoring unit 5011 is specifically configured to:

[0087] Perform cold start training on the basic LLM using the basic preference data to obtain the to-be-trained LLM.

[0088] Optionally, the answer scoring unit 5011 is specifically configured to:

[0089] Input the sample questions into the i-1 level LLM to obtain multiple sample answers generated by the model for each sample question respectively, and score each sample answer using a preset scoring model.

[0090] Optionally, the answer scoring unit 5011 is specifically configured to:

[0091] Score the alignment degree between the sample answer and human preferences using a preset scoring rule; and / or,

[0092] Input the sample answers into multiple alignment evaluation models respectively, so as to score the alignment degree between each sample answer and human preferences by using each alignment evaluation model respectively.

[0093] Optionally, the multiple alignment evaluation models are other LLMs different from the to-be-trained LLM, and each alignment evaluation model is different from each other.

[0094] Optionally, the scoring result of the sample answer corresponding to each sample question includes at least one of the following:

[0095] The rationality score of each sample answer calculated comprehensively from multiple evaluation dimensions;

[0096] When any sample answer contains the formatted text corresponding to the standard format, the format accuracy score of the sample answer used to characterize the degree to which the formatted text conforms to the standard format;

[0097] The answer correct rate of the sample answer corresponding to the sample question, where the answer correct rate is the proportion of the number of correct sample answers in all sample answers corresponding to the sample question.

[0098] Optionally, when the scoring result includes the answer correct rate, the data construction unit 5012 is specifically configured to:

[0099] Determine available sample questions from the sample questions, where the answer correct rate of the corresponding sample answers is not lower than the first threshold and not higher than the second threshold.

[0100] Optionally, when the scoring result includes the rationality score and the answer correct rate, and the rationality score is positively correlated with the rationality degree of the sample answer, the data construction unit 5012 is specifically configured to:

[0101] For each available sample question and its corresponding sample answers:

[0102] Determine positive sample answers with a rationality score not lower than a third threshold from the correct answers included in the sample answers, and determine negative sample answers with a rationality score not higher than a fourth threshold from the wrong answers included in the sample answers, where the third threshold is greater than the fourth threshold;

[0103] Construct the available sample question, the positive sample answer, and the negative sample answer into a piece of training data.

[0104] Optionally, it further includes:

[0105] A resource scheduling unit 502, configured to predict the expected time consumption for constructing the training data at least according to the data volume of the sample question and the sample answer, and allocate training resources to other tasks for use before the start time of the i-th round of training corresponding to the expected time consumption.

[0106] Optionally, when the stop condition is satisfied, it includes:

[0107] The deviation between the scoring result of the sample answer in the i-th round of training and the scoring result of the sample answer in the (i - 1)-th round of training is less than a threshold.

[0108] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor runs the executable instructions to implement the steps of the method as described in any one of the above embodiments.

[0109] Based on the same concept as the above method, this specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any one of the above embodiments are implemented.

[0110] Based on the same concept as the above method, this specification also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method as described in any one of the above embodiments are implemented.

Claims

1. A preference alignment training method for a large language model LLM, comprising: The LLM to be trained is trained for multiple rounds of self-iterative direct preference optimization (DPO) and the training is stopped when the stopping condition is met; wherein, for a positive integer i, the i-1-level LLM obtained by the i-1-th round of training is trained for the i-th round, including: Randomly select sample questions from a preset question bank, input the sample questions into the i-1 level LLM to obtain sample answers generated by the model, and use a preset scoring model to score the degree of alignment between the sample answers and human preferences; Determining available sample questions from the sample questions according to the scoring results of the sample answers, and constructing training data based on the available sample questions and their corresponding available sample answers; The i-1 level LLM is trained using the training data to obtain the i level LLM.

2. The method according to claim 1, obtaining the LLM to be trained, comprising: The basic LLM is cold-start trained using the basic preference data to obtain the LLM to be trained.

3. The method according to claim 1, wherein the inputting the sample question into the i-1 level LLM to obtain the sample answer generated by the model, and using a preset scoring model to score the degree of alignment between the sample answer and human preferences, comprises: The sample questions are input into the i-1 level LLM to obtain a plurality of sample answers generated by the model for each sample question, and each sample answer is scored using a preset scoring model.

4. The method according to claim 1, wherein scoring the degree of alignment between the sample answers and human preferences using a preset scoring model comprises: Scoring the degree of alignment between the sample answers and human preferences using a preset scoring rule; and / or, The sample answers are respectively input into a plurality of alignment evaluation models, so as to use each alignment evaluation model to score the degree of alignment between the sample answers and human preferences.

5. According to the method of claim 4, the multiple alignment evaluation models are other LLMs different from the LLM to be trained, and the alignment evaluation models are different from each other.

6. According to any one of the methods of claims 1-5, the scoring result of the sample answer corresponding to each sample question includes at least one of the following: The rationality score of each sample answer calculated comprehensively from multiple evaluation dimensions; In the case where any sample answer contains formatted text corresponding to the standard format, a format accuracy score of the sample answer used to characterize the degree to which the formatted text conforms to the standard format; The correct answer rate of the sample answers corresponding to the sample question, wherein the correct answer rate is the proportion of the correct sample answers among the sample answers corresponding to the sample question.

7. The method according to claim 6, wherein when the scoring result includes the correctness rate of the answer, determining the available sample questions from the sample questions according to the scoring result of the sample answers comprises: Determine available sample questions from the sample questions, the answer accuracy rate of the corresponding sample answers being not lower than a first threshold and not higher than a second threshold.

8. The method according to claim 6, wherein when the scoring result includes the rationality score and the answer accuracy, and the rationality score is positively correlated with the rationality of the sample answer, constructing training data based on the available sample questions and their corresponding available sample answers comprises: For each available sample question and its corresponding individual sample answer: Determine a positive sample answer whose rationality score is not less than a third threshold from the correct answers included in the sample answers, and determine a negative sample answer whose rationality score is not higher than a fourth threshold from the wrong answers included in the sample answers, wherein the third threshold is greater than the fourth threshold; The available sample question, the positive sample answer and the negative sample answer are constructed as a piece of training data.

9. The method according to claim 1, further comprising: The expected time required to construct the training data is predicted based on at least the data volume of the sample questions and the sample answers, and the training resources are allocated to other tasks before the start time of the i-th round of training corresponding to the expected time.

10. The method according to claim 1, wherein the stopping condition is satisfied, comprising: The deviation between the scoring result of the sample answer of the i-th round of training and the scoring result of the sample answer of the i-1-th round of training is less than the threshold.

11. An electronic device, characterized in that: include: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 10 by executing the executable instructions.

12. A computer-readable storage medium, characterized in that: Computer instructions are stored thereon, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

13. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Task processing model training method, role playing model training method and task processing method

    CN120724109A

  • Task processing model training method, role playing model training method, and task processing method

    CN120724109B

  • Unmanned aerial vehicle body cognition alignment method based on man-machine cooperation

    CN120803003A

  • Data preprocessing and optimizing method and device for generative model in question and answer scene

    CN121303362A

  • Psychological counseling multi-round dialogue method, device and equipment, storage medium and product

    CN121543712A