Mark reinforcement learning-based multi-dimensional preference alignment method and system for large language model
By using multiple reward models for markup enhancement and dataset reconstruction, direct preference optimization with weights combined with confidence, and calibrating the language model through Pratt scaling, the problem of aligning human multidimensional preferences in large language models is solved, and dialogue quality and generalization performance are improved.
Patent Information
- Application Number
- CN202510669015.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The prior art is difficult to accurately align human multidimensional preferences in large language models, resulting in poor generalization performance of generated content and complex and unstable alignment process.
Multiple different reward models are used to score conversation sample data, marking enhancement and data set reconstruction, direct preference optimization with weights is performed in combination with confidence, and language model is calibrated through Pratt scaling.
It improves the alignment effect of language models on multidimensional human preferences, enhances dialogue quality, simplifies the alignment process, and improves the generalization performance of the model.
Smart Images

Figure CN120196748A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence natural language processing, and relates to a multi-dimensional preference alignment method and system for large language models based on token-enhanced learning. Background Art
[0002] With the rapid development of artificial intelligence technology, large language models have made remarkable progress in the field of natural language processing. These models are trained on a vast amount of text data through unsupervised learning and can generate high-quality text content. However, due to the completely unsupervised nature of their training process, these models often have difficulty precisely aligning with human multi-dimensional preferences when generating content.
[0003] To improve the answer quality of language models, researchers have proposed large language model alignment methods. The alignment methods use question-and-answer pairs that reflect human preferences to help the model avoid generating answers that deviate from human preferences. Currently popular alignment methods mainly include supervised fine-tuning, reinforcement learning optimization based on human feedback, and direct preference optimization, etc. These methods can help the model align with human preferences to a certain extent, but there are still some problems.
[0004] Firstly, supervised fine-tuning has high requirements for data quality. The existing preference datasets are marked in a relatively single way and are difficult to help the model avoid generating answers that deviate from human preferences, with poor generalization performance. Secondly, the process of the reinforcement learning method based on human feedback is very complex, requiring a large amount of computing resources, and the optimization process is unstable. Although direct preference optimization simplifies the alignment process, this method is prone to overfitting and can only learn fixed preferences, with poor generalization performance. Finally, human preferences have multiple characteristics, and the above methods all rely on question-and-answer datasets with absolute preferences for alignment, which can only reflect relatively single human preferences and have great limitations. Summary of the Invention
[0005] Object of the Invention: Aiming at the problem of single human preferences reflected by the alignment object in the prior art, the present invention provides a multi-dimensional preference alignment method and system for large language models based on token-enhanced learning, making the language model closer to real human preferences and improving the dialogue quality of the large model.
[0006] Technical Solution: To achieve the above object of the invention, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a multi-dimensional preference alignment method for large language models based on token-enhanced learning, including the following steps: Use multiple different reward models to score the dialogue sample data, obtain the sample preference confidence, perform token enhancement, and reconstruct the preference dataset; After selecting dialogue samples from the reconstructed dataset for supervised training of the large language model, perform weighted direct preference optimization based on confidence for the large language model; Combine Platt scaling for large language model calibration and iteratively update the large language model parameters and calibration parameters.
[0007] Furthermore, for the same question in the preference dataset x , multiple reward models are used to score the dialogue samples ( x,y w ) and ( x,y l ) to obtain sample preference confidence: ; Among them, n is the number of reward models, y w is the "selected" answer, y l is the "rejected" answer, is the indicator function, r i ( x,y w ) and r i ( x,y l ) respectively represent the scores of the i th reward model for the dialogue samples ( x,y w ) and ( x,y l ).
[0008] Furthermore, the reconstruction of the preference dataset includes: reselecting and recording the selected answer for the same question according to the confidence. If the confidence of the original "selected" answer calculated by multiple reward models exceeds the set threshold, it remains unchanged; otherwise, the "selected" answer and the "rejected" answer are exchanged; add a confidence feature to the dataset.
[0009] Furthermore, the supervised training is to reselect the "selected" answer from the training set divided from the reconstructed dataset for supervised training of the large language model. The alignment goal is to maximize the probability that the large language model generates the "selected" answer given the question.
[0010] Furthermore, use a direct preference optimization framework based on confidence weighting. Calculate the preference confidence through multi-reward model integration, model human preferences using the probability ratio method based on the Bradley-Terry model, and construct an objective function that fuses multi-dimensional preference confidence to achieve alignment.
[0011] Furthermore, the objective function of the confidence-based weighted direct preference optimization is as follows: ; where is the policy model to be optimized, , are respectively the probabilities that the model generates the answer x for the input y w , y l . is the reference policy model, , are respectively the probabilities that the model generates the answer x for the input y w , y l . is the training set divided from the reconstructed preference dataset, containing triples( x,y w ,y l ). is the temperature coefficient, controlling the amplitude of policy update, is the Sigmoid activation function, denotes taking the expectation.
[0012] Furthermore, in the calibration of the large language model by combining Platt scaling, the calibration set divided from the reconstructed dataset is used to fit the calibration parameters, minimizing the negative log-likelihood loss on the calibration set: , , ; where is the preference distribution output after model calibration, is the probability that the model generates the answer x for the input y w is greater than the probability of generating the answer y l , A and B are the calibration parameters to be fitted; initialize the calibration parameter A = 1, B = 0, and use the gradient descent algorithm to alternately optimize the calibration parameters and the large language model parameters.
[0013] Second aspect, the present invention provides a multi-dimensional preference alignment system for large language models based on labeled reinforcement learning, including: A dataset reconstruction module, configured to score dialogue sample data using multiple different reward models to obtain sample preference confidence levels, perform labeled enhancement, and reconstruct the preference dataset; A model alignment module, configured to perform supervised training on a large language model by selecting dialogue samples from the reconstructed dataset, and then perform confidence-based weighted direct preference optimization on the large language model; A model calibration module, configured to calibrate the large language model in combination with Platt scaling, and iteratively update the parameters of the large language model and the calibration parameters.
[0014] Third aspect, the present invention provides a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the method for multi-dimensional preference alignment of large language models based on labeled reinforcement learning.
[0015] Fourth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by the processor, it implements the steps of the method for multi-dimensional preference alignment of large language models based on labeled reinforcement learning.
[0016] Advantageous effects: Compared with the prior art, the present invention has the following advantages: (1) The method of the present invention generates dialogue sample confidence levels by invoking multiple reward models as human agents, combines labeled reinforcement learning to solve the shortcoming of the lack of multi-dimensional human preference characteristics in the mainstream alignment dataset, and improves the direct preference optimization method, enhancing the comprehensiveness, effectiveness, and efficiency of the alignment process. This method has broad application prospects in practical applications; (2) The method of the present invention calibrates the language model after multi-dimensional preference alignment through Platt scaling, and by calibrating the answer output probability of the policy model, makes the language model closer to the true human preference distribution, thereby obtaining a high-quality alignment model. Description of the Drawings
[0017] Figure 1 It is a flowchart of the steps of an embodiment of the present invention.
[0018] Figure 2 It is a model architecture diagram of an embodiment of the present invention. Detailed Embodiments
[0019] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.
[0020] An embodiment of the present invention discloses a multi-dimensional preference alignment method for large language models based on labeled reinforcement learning, which mainly includes: using multiple different reward models to score dialogue sample data to obtain sample preference confidence, performing label enhancement, and reconstructing the preference dataset; after selecting dialogue samples from the reconstructed dataset to perform supervised training on the large language model, performing weighted direct preference optimization based on confidence on the large language model; combining Platt scaling for large language model calibration, and iteratively updating the large language model parameters and calibration parameters.
[0021] Mainstream large model alignment algorithms usually rely on a single preference dataset. However, human preferences are often multi-dimensional, such as security, relevance, and effectiveness. Such a single preference dataset will cause problems with poor model alignment. In this embodiment, multiple reward models focusing on different preferences are selected as human agents to score the samples multi-dimensionally, obtain the preference confidence of the samples, and then perform the alignment process. Finally, Platt scaling is introduced to further calibrate the language model to improve the quality of model answers. As Figure 1 and Figure 2 shown, this embodiment specifically includes the following steps: S1: Dataset reconstruction stage: Input the dialogue samples in the preference dataset, score them according to multiple reward models focusing on different preferences, perform labeled reinforcement learning, and reconstruct the preference dataset after obtaining the dialogue sample confidence.
[0022] Specifically, in the open-source human preference dataset, for the same question x , there are two answers reflecting human preferences: the "selected" answer y w and the "rejected" answer y l . Considering the multi-dimensionality of human preferences, such as security, effectiveness, relevance, etc., multiple reward models focusing on different preferences are used to score the dialogue samples ( x,y w ) and ( x,y l ), and then the confidence is obtained: ; Among them, is the sample preference confidence, that is, the true preference distribution of humans, is the indicator function, r i ( x,y w ) and r i ( x,y l ) respectively represent the iThe scoring of a reward model for dialogue samples ( x,y w ) and ( x,y l ). A total of n reward models are selected to label and enhance the original preference dataset, adding the confidence features of the samples. Thus, the original dataset is reconstructed: First, reselect and record the "selected" answer according to the confidence. If , the "selected" answer remains y w , otherwise, exchange its content with y l ; Second, add the sample preference confidence to the dataset to supplement the supervision information of human multi-dimensional preferences.
[0023] S2: Model alignment stage: Divide the data reconstructed in step S1 into a training set and a calibration set. Use the reselected high-quality dialogue samples in the training dataset to perform supervised training on the language model for preliminary alignment; then, combined with the sample confidence, use the training set to perform weighted direct preference optimization on the model. Specifically, it includes: S21: Divide the data reconstructed in step S1 into a training set and a calibration set. Select about 5% of the samples from the original training set as the calibration set , and the remaining samples form the new training set , ensuring the independence of the calibration set.
[0024] S22: Use the reselected "selected" answers y w in the training set to perform supervised training on the model. Use in the training set as the input of the language model, where x i is the question part in the dialogue sample, and y w_i is the "selected" answer in the dialogue sample. A total of d dialogue samples are used for supervised instruction fine-tuning for preliminary alignment. The alignment goal of this stage is to maximize the probability that the language model generates the "selected" answer given the question: ; Among them, is the policy model to be optimized, m is the length of the answer y w , is the probability that the policy model generates the specified answer x for the input y kThe probability. Supervised training can align the model with the expected behavior norms or criteria to obtain a preliminarily aligned model , as a reference policy model.
[0025] S23: Use the reward model as a human agent to simulate human preferences. The intuitive alignment goal of the model is:[[]] ; Among them, is a hyperparameter between the balanced divergence value and the expected value of the reward model, r ( x,y ) is the score of the reward model for the question-answer pair ( x,y ), is the KL divergence function, represents the policy model for the input x to generate the answer y The probability of represents the model for the input x to generate the answer y The probability of, combined with the Bradley-Terry probability model: ; Among them, , respectively represent the policy model for the input x to generate the answer y 1, y 2 probability, , respectively represent the model for the input x to generate the answer y 1, y 2 probability, is the temperature coefficient that controls the update frequency.
[0026] Combining the model's intuitive alignment goal and the Bradley-Terry probability model, aligning the language model based on the human preference dataset, the objective function is: ; Among them, is the objective loss function, is the reference policy model obtained in step S22, x is the question part of the sample, y w and y l are the "selected" and "rejected" answers in the sample respectively, , They are the models for the input x to generate answers y w , y l the probability of , They are the models for the input x to generate answers y w , y l the probability of is the Sigmoid activation function.
[0027] S24: Humans rarely have a "black or white" preference. To align with multi-dimensional human preferences, the sample preference confidence obtained by scoring according to multiple reward models in step S1 is used as the weight for direct preference optimization, then: ; Minimize the objective function to obtain a further aligned language model .
[0028] S3: Model calibration stage: According to the aligned model in step S2, in order to make the model more conform to multi-dimensional human preferences, model calibration is performed by combining Platt scaling, and the model parameters and calibration parameters are iteratively updated.
[0029] Specifically, introduce Platt scaling into weighted direct preference optimization, and use the calibration set obtained by partitioning in step S21 , to fit the calibration parameters: For each sample in the calibration dataset, use the sample preference confidence as a proxy for the preference probability annotated by humans, that is, the preference label distribution in the calibration dataset: , ; where is the preference distribution output after model calibration, A and B are the calibration parameters to be fitted, is the model optimized in step S24 for the input x to generate a "selected" answer y w with a greater probability than generating a "rejected" answer y l The probability. Minimize the negative log-likelihood loss on the calibration set: ; Among them, the initialization calibration parameter A = 1, B = 0, and the gradient descent algorithm is used to alternately optimize the calibration parameter and the model parameter.
[0030] S4: Q&A stage: Use the large language model calibrated in step S3 to have a conversation with the question examples in the test set to obtain high-quality answers aligned with multi-dimensional human preferences.
[0031] In summary, the present invention discloses a method for aligning multi-dimensional preferences of a large language model based on token-based reinforcement learning, proposes to use multiple reward models to reconstruct the preference dataset, and perform weighted direct preference optimization on the language model according to the sample confidence for comprehensive alignment. In addition, Pratt scaling is proposed to perform model calibration to make the language model closer to the real human preference distribution, and finally achieve high-quality alignment with multi-dimensional human preferences.
[0032] Next, take the experiment of the present invention for multi-dimensional preference alignment of the Qwen2.5-7B and Llama3-8B models on the public dataset HuggingFaceH4 / ultrafeedback_binarized to illustrate the advantages of the present invention over the prior art.
[0033] We used four open-source reward models to score the preference dataset HuggingFaceH4 / ultrafeedback_binarized, perform token enhancement to obtain sample preference confidence features, reconstruct the dataset, and then use the alignment method disclosed in the present invention to perform alignment fine-tuning on the Llama3-8B and Qwen2.5-7B models respectively. Under the settings of this experiment, we selected 4 open-source reward models (RM) focusing on different preferences as human agents, and calculated the model output win rate on each reward model. In this experiment, we used all samples in the test set to calculate the win rate. The experimental results are shown in Table 1. Among them, the 4 open-source reward models RM1, RM2, RM3, and RM4 are: Skywork / Skywork-Reward-Llama-3.1-8B-v0.2, LxzGordon / URM-LLaMa-3.1-8B, Skywork / Skywork-Reward-Llama-3.1-8B, and RLHFlow / ArmoRM-Llama3-8B-v0.1. We compared the current mainstream alignment methods, including: instruction fine-tuning method (SFT), reinforcement learning method based on human feedback (PPO), and direct preference optimization method (DPO).
[0034] Table 1 Winning Rates of the Invention Method Based on Each RM on the UltraFeedback Test Set
[0035] In Table 1, the baseline corresponds to the base models of pre-trained Llama 3 - 8B or Qwen 2.5 - 7B. The experimental data represent the proportion of cases where the policy model optimized using the method of the present invention answers better than the policy models optimized by other methods for all questions in the test set, and the winning rate of the policy model optimized by the method of the present invention relative to other policy models. The formula for calculating the winning rate of policy model A relative to policy model B is as follows: ; where and represent policy model A and policy model B respectively, M represents the number of test samples, represents the answer given by policy model A to the test question x m , represents the reward score of this question - answer pair, represents the answer given by policy model B to the test question x m , represents the reward score of this question - answer pair, is the indicator function.
[0036] In addition, we also calculated the normalized reward model scores and used the large AI model Claude 3.5 - haiku for model evaluation. We randomly selected 500 test samples to calculate the model winning rate, and the experimental results are shown in Table 2.
[0037] Table 2 Normalized Reward Scores and AI - Evaluated Winning Rates of the Experimental Method on the UltraFeedback Dataset
[0038] It can be seen that the experimental method in the present invention has achieved the best performance in all indicators, proving the feasibility of the invention motivation and the invention method.
[0039] Based on the same inventive concept, an embodiment of the present invention also discloses a multi-dimensional preference alignment system for large language models based on labeled reinforcement learning, including: a dataset reconstruction module, configured to score dialogue sample data using multiple different reward models to obtain sample preference confidence, perform label enhancement, and reconstruct the preference dataset; a model alignment module, configured to select dialogue samples from the reconstructed dataset to perform supervised training on the large language model, and then perform weighted direct preference optimization based on confidence on the large language model; a model calibration module, configured to perform large language model calibration in combination with Platt scaling, and iteratively update the large language model parameters and calibration parameters.
[0040] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the above-mentioned method for multi-dimensional preference alignment of large language models based on labeled reinforcement learning are implemented.
[0041] An embodiment of the present invention also discloses a computer program product, including a computer program. When the computer program is executed by the processor, the steps of the above-mentioned method for multi-dimensional preference alignment of large language models based on labeled reinforcement learning are implemented.
[0042] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, so that when the program codes are executed by the processor or controller, the steps of the method of the present invention are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed as an independent software package partially on the machine and partially on a remote machine, or executed entirely on a remote machine or server. Those parts not detailed in the present invention are all well-known techniques in the art.
[0043] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. A multi-dimensional preference alignment method for large language models based on labeled reinforcement learning, characterized in that, It includes the following steps: Use multiple different reward models to score the dialogue sample data, obtain the sample preference confidence, perform label enhancement, and reconstruct the preference dataset; After selecting dialogue samples from the reconstructed dataset to perform supervised training on the large language model, perform confidence-based weighted direct preference optimization on the large language model; Combine Platt scaling to calibrate the large language model, and iteratively update the large language model parameters and calibration parameters.
2. The multi-dimensional preference alignment method for large language models based on labeled reinforcement learning according to claim 1, wherein For the same question in the preference dataset x , multiple reward models are used to score the dialogue samples ( x,y w ) and ( x,y l ), and the sample preference confidence is obtained: ; Among them, n is the number of reward models, y w is the "selected" answer, y l is the "rejected" answer, is the indicator function, r i ( x, y w ) and r i ( x,y l ) respectively represent the scores of the i th reward model for the dialogue samples ( x,y w ) and ( x,y l ).
3. A multi-dimensional preference alignment method for large language models based on marked reinforcement learning according to claim 1, characterized in that The reconstruction of the preference dataset includes: reselecting and recording the selected answers to the same question according to the confidence. If the confidence of the original "selected" answer calculated according to multiple reward models exceeds the set threshold, it remains unchanged; otherwise, exchange the "selected" answer and the "rejected" answer; add confidence features to the dataset.
4. A multi-dimensional preference alignment method for large language models based on marked reinforcement learning according to claim 1, characterized in that, The supervised training is to reselect the "selected" answer from the training set divided from the reconstructed dataset to perform supervised training on the large language model. The alignment goal is to maximize the probability that the large language model generates the "selected" answer given the question.
5. A multi-dimensional preference alignment method for large language models based on labeled reinforcement learning according to claim 1, characterized in that Use a direct preference optimization framework based on confidence weighting. Calculate the preference confidence through multi-reward model integration, model human preferences using the probability ratio method based on the Bradley-Terry model, and construct an objective function that fuses multi-dimensional preference confidence to achieve alignment.
6. A multi-dimensional preference alignment method for large language models based on labeled reinforcement learning according to claim 5, characterized in that, The objective function of the confidence-based weighted direct preference optimization is as follows: ; Among them, is the policy model to be optimized, and are the probabilities that the model generates answers x for the input y w and y l respectively. is the reference policy model, and are the probabilities that the model generates answers x for the input y w and y l respectively. is the training set divided from the reconstructed preference dataset, which contains triples ([[]] x,y w ,y l ). is the sample preference confidence, is the temperature coefficient, which controls the amplitude of policy update, is the Sigmoid activation function, represents the expectation.
7. A multi-dimensional preference alignment method for large language models based on labeled reinforcement learning according to claim 1, characterized in that In the calibration of the large language model by combining Platt scaling, use the calibration set divided from the reconstructed dataset to fit the calibration parameters, and minimize the negative log-likelihood loss on the calibration set: , , ; Among them, is the sample preference confidence, is the preference distribution output after model calibration, is the policy model to be optimized, 、 are the models for the input x to generate the answer y w 、 y l probabilities, is the reference policy model, 、 are the models for the input x to generate the answer y w 、 y l probabilities, is the probability that the model for the input x to generate the answer y w is greater than the probability of generating the answer y l ; A and B are the calibration parameters to be fitted; Initialize the calibration parameters A = 1, B = 0, and use the gradient descent algorithm to alternately optimize the calibration parameters and the large language model parameters.
8. A multi-dimensional preference alignment system for large language models based on marker-enhanced learning, characterized in that, It includes: A dataset reconstruction module for using multiple different reward models to score the dialogue sample data, obtaining the sample preference confidence, performing label enhancement, and reconstructing the preference dataset; A model alignment module for selecting dialogue samples from the reconstructed dataset to perform supervised training on the large language model, and then performing confidence-based weighted direct preference optimization on the large language model; A model calibration module for calibrating the large language model by combining Platt scaling, and iteratively updating the large language model parameters and calibration parameters.
9. A computer system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of a multi-dimensional preference alignment method for a large language model based on label enhancement learning according to any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of a multi-dimensional preference alignment method for a large language model based on label enhancement learning according to any one of claims 1-7.
Citation Information
Patent Citations
Decision-making method and model for offline reinforcement learning and continuous online fine tuning
CN119249360A
Large model security alignment method based on dynamic constraint reinforcement learning
CN119539057A
Calibrating confidence scores for machine learning models trained as natural language interfaces
CN119790413A
User preference analysis method based on memory enhanced multi-modal large model
CN119884981A
Deep learning-based sales prediction method and system
CN120013590A
Cited By
Intelligent content evaluation and optimization method and system based on multi-standard preference learning
CN120494074A
Large language model evaluation method and system based on SimPO
CN121524565A