Preference alignment optimization method based on reward-driven selective punishment

By using a reward-driven selective penalty mechanism to dynamically classify and weight samples, this method solves the problems of insufficient sample quality differentiation and gradient imbalance in existing preference optimization methods, achieving a more efficient and stable preference alignment effect and improving the generation quality and consistency of large language models.

CN121301935APending Publication Date: 2026-01-09NORTHEASTERN UNIV CHINA
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511718682.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing preference optimization methods have shortcomings in sample selection and weight allocation, and cannot effectively distinguish between differences in sample quality and complexity. This leads to overfitting of the model on noisy or fuzzy preference data, unbalanced gradient propagation, and a lack of real-time perception and response to sample reward signals, making it difficult to maintain stable alignment performance in complex human feedback scenarios.

Method used

By introducing a reward-driven selective penalty mechanism, the implicit reward signal is used to dynamically classify the samples into four categories: high-quality alignment, low-quality alignment, incorrect preference, and fuzzy preference. An RSPO loss function is constructed to achieve differentiated weighted optimization, dynamically adjust the learning intensity, and suppress noise interference.

Benefits of technology

It significantly improves the model's alignment performance and generalization ability under complex human feedback data, enhances the stability and consistency of generated content, reduces the risk of overfitting, and strengthens the model's stability and efficiency in preference alignment tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301935A_ABST
    Figure CN121301935A_ABST
Patent Text Reader

Abstract

The invention provides a preference alignment optimization method based on reward-driven selective punishment, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining intelligent question and answer training data, and constructing an intelligent question and answer training sample set which comprises a plurality of intelligent question and answer training samples; taking a to-be-optimized large language model as a strategy model and a reference model; and training the strategy model by using the intelligent question and answer training sample set, and measuring the offset amplitude of the strategy model before and after training by using the reference model to obtain an optimized large language model. According to the method, implicit reward signals in the model are introduced, preference data are divided into multiple categories according to implicit reward distribution generated by the model, a dynamic weight function is designed, differential weighted optimization is carried out on different categories of samples, reinforcement learning of high-quality samples and suppression of low-quality samples are achieved, and the method has the advantages of being high in robustness and high in robustness. Weight of low-quality or conflict samples is reduced while high-quality sample learning is enhanced, noise interference is suppressed, and overfitting is prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, and in particular relates to a preference alignment optimization method based on reward-driven selective punishment. Background Technology

[0002] In recent years, the rapid development of large-scale language models (LLMs) in the field of natural language processing (NLP) has driven breakthroughs in various applications such as human-computer interaction, knowledge generation, and intelligent decision-making. However, while these models possess powerful language generation capabilities, they often suffer from inconsistencies with human intent and outputs that deviate from user expectations.

[0003] To address the gap between model output and human preferences, researchers have proposed "Preference Alignment Optimization" methods, enabling models to adjust their generation behavior based on human feedback. Among these, Reinforcement Learning from Human Feedback (RLHF) was the earliest and most widely adopted framework. It evaluates generation quality by training a reward model and then uses reinforcement learning algorithms (such as PPO) to optimize the policy model, ensuring it generates content that aligns with human preferences. While RLHF has achieved some success in practice, its training process is complex, computationally expensive, and lacks stability. Furthermore, the reward model often suffers from misleading or overfitting issues. Therefore, Direct Preference Optimization (DPO) was proposed as an alignment method that does not require an explicit reward model. It directly models human preference signals by optimizing the probability ratio between chosen and rejected responses.

[0004] While Direct Preference Analysis (DPO) simplifies the training process and improves efficiency, it still has several limitations: First, DPO treats all data samples equally, failing to distinguish between differences in sample quality and complexity, leading to overfitting on noisy or ambiguous preference data. Second, DPO penalizes non-preference responses too heavily during gradient propagation, while rewarding preference responses insufficiently, resulting in an imbalance in optimization direction. Furthermore, DPO is susceptible to preference labeling noise; when ambiguous or conflicting preferences exist in the data, the model's alignment performance significantly decreases. Finally, DPO lacks adaptability to samples of varying difficulty, making it difficult to maintain stable generalization performance in complex human feedback scenarios. To overcome these problems, researchers have proposed various improvement methods, such as introducing dynamic weighting, sample resampling, or curriculum learning strategies to adjust the optimization intensity based on sample characteristics. However, most of these methods rely on predefined rules or static parameters, lacking real-time perception and response capabilities to sample reward signals, and failing to fully leverage the implicit preference structure features between samples. In summary, existing preference optimization frameworks still have significant shortcomings in sample selection and weight allocation. There is an urgent need for an innovative method that can dynamically adjust the optimization strategy based on the implicit rewards within the model, and suppress noise sample interference while strengthening high-quality preference learning, so as to achieve more efficient and stable human-machine preference alignment. Summary of the Invention

[0005] To address the shortcomings of existing technologies, in a first aspect, this invention provides a preference alignment optimization method based on reward-driven selective penalty, comprising the following steps:

[0006] Acquire intelligent question answering training data, process the intelligent question answering training data, and construct an intelligent question answering training sample set, including several intelligent question answering training samples;

[0007] The large language model to be optimized is used as the strategy model. and reference model ;

[0008] Use the intelligent question answering training sample set to train the policy model Training is performed using a reference model. Measuring the policy model before and after training The offset magnitude is used to obtain the optimized large language model, including:

[0009] Input the intelligent question answering training samples into the policy model respectively. and reference model Perform forward reasoning and compute the strategy model. The probability distribution of the policy logarithm generated from the training samples of the intelligent question answering system is used to calculate the reference model. The reference logarithm of the training samples for intelligent question answering generates a probability distribution;

[0010] Calculate the implicit reward value of the response;

[0011] Establish a sample classification mechanism to divide the intelligent question answering training samples into multiple categories based on the implicit reward value of the response;

[0012] Construct the RSPO loss function and use the RSPO loss function to evaluate the policy model. Training is performed to obtain a trained policy model. That is, the optimized large language model.

[0013] Furthermore, the intelligent question-answering training data includes: several input prompts. Preference response Non-preference response ;

[0014] Processing the training data for intelligent question answering includes: processing each input suggestion Formatting and normalization processes are performed, and each input suggestion is constructed using a word segmenter. input token sequence Generate suggestions for each input. Corresponding intelligent question answering training samples .

[0015] Furthermore, the policy logarithmic generation probability distribution includes the policy preference logarithmic generation probability distribution. The probability distribution of non-preference log generation of strategies ;

[0016] Computational strategy model The probability distribution of the policy logarithm for the training samples of intelligent question answering is generated using the following method:

[0017] training samples for intelligent question answering Input policy model Perform forward reasoning to obtain the policy model. Given input prompts Under the condition of preference response Non-preference response Token-level preference generation probability distribution and token-level non-preference generation probability distribution ;

[0018] Probability distribution generated based on token-level preferences and token-level non-preference generation probability distribution Computational strategy model Given input prompts Under the condition of preference response Non-preference response The log-generated probability distribution of strategy preferences The probability distribution of non-preference log generation of strategies .

[0019] Furthermore, the reference logarithmic generation probability distribution includes the reference preference logarithmic generation probability. and the probability of generating the reference non-preference logarithm ;

[0020] Computational reference model The reference logarithm of the training samples for intelligent question answering generates a probability distribution, specifically using the following method:

[0021] training samples for intelligent question answering Input reference model Perform forward reasoning to obtain the reference model. Given input prompts Under the condition of preference response Non-preference response Reference preference log generation probability and the probability of generating the reference non-preference logarithm .

[0022] Furthermore, the implicit reward value of the response is calculated using the following method:

[0023] Calculate the implicit reward value of the preference response based on the implicit reward formula in the DPO theory. Implicit reward value of non-preference response As shown in the formula below:

[0024]

[0025] in, This is the implicit reward formula in the DPO theory. To control the temperature coefficient of the reward scale in the preference response, To control the temperature coefficient of the reward scale for non-preference responses.

[0026] Furthermore, a sample classification mechanism is established to divide the intelligent question-answering training samples into multiple categories based on the implicit reward value of the response. The specific method is as follows:

[0027] Implicit reward value based on preference response Implicit reward value of non-preference response The training samples for intelligent question answering are divided into multiple categories, including: high-quality alignment R1, low-quality alignment R2, incorrect preference R3, and ambiguous preference R4.

[0028] when At that time, the training samples for intelligent question answering are assigned to the high-quality aligned category R1;

[0029] when At that time, the training samples for intelligent question answering were assigned to the low-quality aligned category R2;

[0030] when At that time, the training samples for intelligent question answering are classified into the error preference category R3;

[0031] when At that time, the training samples for intelligent question answering are classified into the fuzzy preference category R4;

[0032] in, and All are learnable or preset thresholds.

[0033] Furthermore, the RSPO loss function is constructed as follows:

[0034] Constructing a nonlinear dynamic weighting function As shown in the formula below:

[0035]

[0036] in, For the sigmoid function, This is the minimum weight bias term;

[0037] Implicit reward value in preference response Implicit reward value of non-preference response Based on this, a nonlinear dynamic weighting function is introduced. Construct the RSPO loss function As shown in the formula below:

[0038]

[0039] in, This is the training sample set for intelligent question answering.

[0040] Secondly, this application proposes an electronic device comprising: one or more processors, and a memory for storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the preference alignment optimization method based on reward-driven selective penalty.

[0041] Thirdly, this application proposes a computer-readable storage medium storing executable instructions that, when executed, cause a processor to perform the aforementioned preference alignment optimization method based on reward-driven selective penalty.

[0042] Fourthly, this application proposes a computer program product, including a computer program or instructions that, when executed by a processor, implement the aforementioned preference alignment optimization method based on reward-driven selective punishment.

[0043] The beneficial effects of adopting the above technical solution are as follows: This invention provides a preference alignment optimization method based on reward-driven selective penalty to address the shortcomings of existing Direct Preference Optimization (DPO) in terms of sample weighting, noise handling, and gradient imbalance. This invention introduces implicit reward signals within the model to dynamically classify and weight preference data, thereby achieving reinforcement learning of high-quality samples and suppression of low-quality samples. Specifically, this invention introduces fine-grained data quality awareness into the DPO framework, classifying preference data into multiple categories based on the implicit reward distribution generated by the model. A sigmoid-coupled dynamic weight function is designed to achieve differentiated weighting optimization for samples of different categories, thereby automatically reducing the training weights of low-quality or conflicting samples while reinforcing the learning of high-quality samples, effectively suppressing noise interference and preventing overfitting. Simultaneously, an asymmetric scaling factor is used to balance the learning dynamics of the model in generation and suppression tasks, constructing a more stable, robust, and efficient preference optimization objective, significantly improving the alignment performance and generalization ability of large language models under complex human feedback data. The RSPO of this invention can effectively identify and strengthen high-confidence preference samples, suppress the negative impact of fuzzy or noisy samples, significantly improve the stability, generalization and training efficiency of the model in preference alignment tasks, and ultimately achieve higher quality and more robust human preference consistency optimization. Attached Figure Description

[0044] Figure 1 A schematic diagram of the preference alignment optimization method based on reward-driven selective penalty provided in Embodiment 1 of the present invention;

[0045] Figure 2 The distribution diagram of four types of preference response pairs provided in Embodiment 1 of the present invention, wherein (a) is a diagram showing the change of the proportion of samples trained on Mistral-SFT with the number of training steps, and (b) is a diagram showing the change of the proportion of samples trained on Llama-3-Instruct with the number of training steps.

[0046] Figure 3 A comparison chart of the gradient change trends of RSPO and DPO on Mistral-SFT provided in Embodiment 1 of the present invention;

[0047] Figure 4The comparison chart of the length control win rate (LC) of RSPO and DPO on Mistral-SFT provided in Embodiment 1 of the present invention. Detailed Implementation

[0048] The specific implementation methods of this application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0049] Example 1:

[0050] While existing preference optimization methods have improved the human alignment performance of large language models to some extent, they still have significant shortcomings. First, DPO and its improved algorithms generally use fixed optimization weights, treating all preference samples equally and ignoring significant differences in quality, difficulty, and consistency among samples. This results in the model's inability to effectively distinguish high-quality preference signals from low-quality or noisy samples. Second, existing methods lack in-depth modeling of implicit reward signals, failing to dynamically identify which samples contribute to stable alignment and which may be misleading from the model's own probability distribution, thus limiting the model's adaptability to complex human preference scenarios. Furthermore, DPO exhibits asymmetry in gradient propagation: the model tends to more easily reduce the generation probability of rejection samples while struggling to effectively increase the generation probability of preferred samples, leading to gradient imbalance during optimization and affecting the model's learning of the true preference direction. Finally, because DPO updates gradients for both noisy and ambiguous samples during training, this can cause the model to overfit on complex or inconsistent preference-labeled data, resulting in unstable generated content and decreased generalization ability. Some studies have attempted to alleviate this problem by introducing sample reweighting or curriculum learning mechanisms, but these methods are mostly static rules that cannot respond in real time to the dynamic changes in implicit rewards during training, and they also lack fine-grained control over different sample categories. Finally, in training with large-scale preference data, models are easily disturbed by anomalous samples when faced with diverse preference signals, making it difficult to maintain a stable optimization trajectory and effective gradient updates. In summary, current preference optimization methods still lack a mechanism that can dynamically evaluate sample quality based on the model's internal implicit rewards, reinforce high-quality samples, and penalize low-quality samples, thus failing to achieve stable, efficient, and robust alignment optimization in complex preference data scenarios.

[0051] To address the aforementioned problems, this embodiment provides a preference alignment optimization method based on reward-driven selective penalty, such as... Figure 1 As shown, training a large model for intelligent question answering in an intelligent question answering scenario includes the following steps:

[0052] Acquire intelligent question answering training data, process the intelligent question answering training data, and construct an intelligent question answering training sample set, including several intelligent question answering training samples;

[0053] The training data for intelligent question answering includes several input suggestions. Preference response Non-preference response Processing the training data for intelligent question answering includes: processing each input prompt... Formatting and normalization processes are performed, and each input suggestion is constructed using a word segmenter. input token sequence Generate suggestions for each input. Corresponding intelligent question answering training samples ;

[0054] The large language model to be optimized is used as the strategy model. and reference model ;

[0055] Use the intelligent question answering training sample set to train the policy model Training is performed using a reference model. Measuring the policy model before and after training The offset magnitude is used to obtain the optimized large language model, including:

[0056] Use the intelligent question answering training sample set to train the policy model Training is performed using a reference model. Measuring the policy model before and after training The offset magnitude is used to obtain the optimized large language model, including:

[0057] Input the intelligent question answering training samples into the policy model respectively. and reference model Perform forward reasoning and compute the strategy model. The probability distribution of the policy logarithm generated from the training samples of the intelligent question answering system is used to calculate the reference model. The reference logarithm of the training samples for intelligent question answering generates a probability distribution, specifically using the following method:

[0058] The log-probability distribution of policy generation includes the log-probability distribution of policy preference and the log-probability distribution of policy non-preference;

[0059] training samples for intelligent question answering Input policy model Perform forward reasoning to obtain the policy model. Given input prompts Under the condition of preference response Non-preference response Token-level preference generation probability distribution and token-level non-preference generation probability distribution ;

[0060] Probability distribution generated based on token-level preferences and token-level non-preference generation probability distribution Computational strategy model Given input prompts Under the condition of preference response Non-preference response The log-generated probability distribution of strategy preferences The probability distribution of non-preference log generation of strategies Used to represent strategy models The relative degree of preference for different candidate answers;

[0061] The reference log generation probability distribution includes the reference preferred log generation probability and the reference non-preferential log generation probability;

[0062] training samples for intelligent question answering Input reference model Perform forward reasoning to obtain the reference model. Given input prompts Under the condition of preference response Non-preference response Reference preference log generation probability and the probability of generating the reference non-preference logarithm ;

[0063] The method for calculating the implicit reward value of the response is as follows:

[0064] Reference Model As a stable behavioral baseline, the reference model Output reference preference log generation probability distribution and reference non-preference log generation probability distribution Respectively with the strategy model Output policy preference log generation probability distribution The probability distribution of non-preference log generation of strategies Compare to measure the strategy model Relative to the reference model The offset range is determined by the following method:

[0065] Calculate the implicit reward value of the preference response based on the implicit reward formula in the DPO theory. Implicit reward value of non-preference response The implicit reward formula in the DPO theory is shown below:

[0066]

[0067] in, The temperature coefficient is used to control the reward scale. Input prompts The responses include preferred and unpreferred responses;

[0068] Preference response implicit reward value Implicit reward value of non-preference response As shown in the formula below:

[0069]

[0070] in, To control the temperature coefficient of the reward scale in the preference response, To control the temperature coefficient of the reward scale for non-preference responses;

[0071] A sample classification mechanism is established to divide the intelligent question answering training samples into multiple categories based on the implicit reward value of the response. The specific method is as follows:

[0072] Implicit reward value based on preference response Implicit reward value of non-preference response The training samples for intelligent question answering are divided into multiple categories, including: high-quality alignment R1, low-quality alignment R2, incorrect preference R3, and ambiguous preference R4.

[0073] when At that time, the training samples for intelligent question answering are assigned to the high-quality aligned category R1;

[0074] when At that time, the training samples for intelligent question answering were assigned to the low-quality aligned category R2;

[0075] when At that time, the training samples for intelligent question answering are classified into the error preference category R3;

[0076] when At that time, the training samples for intelligent question answering are classified into the fuzzy preference category R4;

[0077] in, and All of these are learnable or preset thresholds used to distinguish the degree of alignment;

[0078] The sample classification mechanism enables fine-grained partitioning of the training data, providing a basis for subsequent differentiated optimization. Specifically, in a sample classification system based on implicit rewards, each input cue corresponds to a preference pair. Both respond to implicit reward values ​​through preference. Implicit reward value of non-preference response This approach reveals the policy model's internal preference for preferred and non-preferred responses. This reward is not merely the logarithm of the probability ratio, but a direct reflection of the policy model's current cognitive state. Therefore, this embodiment utilizes the implicit reward value of preferred responses. Implicit reward value of non-preference response The alignment and reliability of the training samples for intelligent question answering are inferred from the sign and relative size of the samples. Based on this, the four categories R1, R2, R3, and R4 constitute a structured hierarchical structure of the training samples for intelligent question answering, with each layer reflecting different learning values ​​and risks.

[0079] High-quality aligned training samples for intelligent question answering in category R1 are strong signals. When intelligent question answering training samples are classified into category R1, the reward difference is significant. Not only are they positive, but they also typically have a clear directionality. The policy model has demonstrated implicit judgments consistent with human preferences, indicating that it can clearly distinguish the quality of the two types of responses. The intelligent question-answering training samples in category R1 provide the most stable and consistent gradient signals during optimization, ensuring the loss function continuously progresses in the correct direction during updates. Therefore, the intelligent question-answering training samples in category R1 are high-confidence preference samples, characterized by strong signals and low gradient noise, making them crucial data for enhancing the alignment capabilities of the policy model.

[0080] The intelligent question-answering training samples in low-quality aligned category R2 are weak signals. When intelligent question-answering training samples are assigned to category R2, the policy model will favor the response. Non-preference response If all are considered undesirable, it means that the strategy model cannot extract information from preference pairs. A clear distinction is achieved at this point. The reward difference... The absolute value of these values ​​is often small, and their direction is even unstable, making them prone to gradient oscillations. Therefore, the intelligent question-answering samples in low-quality aligned category R2 carry weak signals. Their main characteristic is that the policy model cannot distinguish between good and bad answers and passively assigns low rewards to both answers simultaneously. Using intelligent question-answering samples in category R2 indiscriminately increases training noise and makes the learning process unstable. Therefore, the role of intelligent question-answering samples in low-quality aligned category R2 is more auxiliary and supplementary than to dominate policy model updates. It is necessary to reduce the weight of intelligent question-answering samples in low-quality aligned category R2 or re-evaluate them in conjunction with an external reward model to avoid interfering with the policy model's preference direction.

[0081] The intelligent question-answering training samples in error preference category R3 are erroneous signals. The policy model's preference direction for these samples is opposite to that of the manually labeled samples, resulting in a reward difference. Typically significantly negative, resulting in a reward difference. The gradient direction is completely inverted. If the intelligent question-answering samples in the erroneous preference category R3 are used with the same weight as intelligent question-answering samples in other categories, it is equivalent to continuously reinforcing the erroneous preferences of the policy model. Training may even experience backward convergence under the pull of such samples. Therefore, the value of identifying intelligent question-answering samples in the erroneous preference category R3 lies not in ignoring them, but in using them as correction signals. By adding a correction penalty term, the policy model actively corrects its erroneous judgments on these samples, extending preference optimization from learning correct samples to correcting erroneous samples, thereby more effectively improving the accuracy of the policy model's preference structure.

[0082] Training samples for intelligent question answering in the fuzzy preference R4 category are fuzzy signals. These samples typically arise when the task itself has multiple solutions, the annotations are ambiguous, and the gradient direction is unclear. Using these samples to train a policy model may weaken preference differences, causing the policy model to misinterpret preference responses during the learning process. Non-preference response The ambiguous region of "good" oscillates repeatedly. Therefore, it is necessary to identify and weaken the intelligent question-answering samples in the ambiguous preference category R4, and then use a stronger reward model in subsequent processes to further determine whether the preferred response is indeed better than the non-preferred response, so as to prevent the ambiguous preference from eroding the alignment capability of the strategy model.

[0083] In summary, this embodiment uses implicit reward values ​​based on preference responses. Implicit reward value of non-preference response The training samples for intelligent question answering are divided into four categories, R1–R4, transforming them from coarse-grained preference pairs into multi-level reward signals. This enables the policy model to distinguish between strong, weak, erroneous, and ambiguous signals. This sample classification mechanism not only clarifies the training objective but also provides a basis for subsequent differential optimization. It transforms preference learning from passively fitting homogeneous data to actively and selectively absorbing valuable data, thereby significantly reducing gradient noise, improving alignment efficiency, and enhancing the final quality of the generated model.

[0084] Construct the RSPO loss function and use the RSPO loss function to evaluate the policy model. Training is performed to obtain a trained policy model. That is, the optimized large language model.

[0085] Constructing a nonlinear dynamic weighting function As shown in the formula below:

[0086]

[0087] in, For the sigmoid function, This is the minimum weight bias term;

[0088] Minimum weight bias term This is used to ensure that each intelligent question-answering training sample is always involved in training and to prevent gradient vanishing; Used to measure the ability of a policy model to promote a preference response. A larger value indicates that learning should be reinforced more; Used to measure the difficulty of a policy model in suppressing undesirable responses; when suppression of undesirable responses fails, the implicit reward value of the undesirable response is... It typically oscillates around 0, neither gradually rising nor gradually falling. This leads to a decrease in overall weights. To address the difference in learning speed between the policy model and other tasks involving reducing the probability of generating non-preferred responses and increasing the probability of generating preferred responses, this embodiment employs an asymmetric scaling factor to balance the two types of learning dynamics. To reduce sensitivity to the ability to generate preference responses, set Used to enhance sensitivity to the ability to suppress unfavorable responses;

[0089] Implicit reward value in preference response Implicit reward value of non-preference response Based on this, a nonlinear dynamic weighting function is introduced. Construct the RSPO loss function As shown in the formula below:

[0090]

[0091] in, This is a training sample set for intelligent question answering.

[0092] The loss function used is RSPO (Reward-Driven Selective Penalization for Preference Alignment Optimization). For strategy model Training is performed to obtain a trained policy model. That is, the optimized large language model.

[0093] Compared to existing methods, this embodiment introduces a reward-driven selective penalty mechanism into the Direct Preference Optimization (DPO) framework, achieving dynamic weighted optimization of samples with different quality preferences. This demonstrates significant advantages in both theoretical analysis and experimental results. Theoretically, this embodiment incorporates implicit reward classification and dynamic penalty weights into the loss function to construct the RSPO loss function. This eliminates the reliance on a uniform fixed ratio constraint during optimization, allowing the learning intensity to automatically adjust based on the sample's reward distribution. This strengthens the contribution of high-quality samples and suppresses the influence of low-quality or noisy samples during the gradient update phase. This mechanism effectively alleviates the gradient imbalance and over-optimization problems inherent in traditional DPO, enabling the model to maintain training stability and directional consistency even under complex preference data distributions.

[0094] This embodiment performs DPO training on Mistral-SFT and Llama-3-Instruct. The distribution of the four types of preference response pairs is as follows: Figure 2 As shown, (a) is a graph showing the change in the proportion of samples trained on Mistral-SFT as a function of the number of training steps, and (b) is a graph showing the change in the proportion of samples trained on Llama-3-Instruct as a function of the number of training steps.

[0095] like Figure 2 As shown, the distribution of preference response pairs indicates that the policy model struggles to consistently capture preference signals on most training data. However, DPO treats all preference response pairs equally, regardless of their quality or complexity. This can lead to the model's inability to effectively distinguish high-quality preference signals from low-quality or noisy samples, thus affecting the final training results, especially when noisy or ambiguous preference response pairs dominate the training dataset.

[0096] Comparison of the gradient trends of RSPO and DPO on the Mistral-SFT: Figure 3 As shown, compared to DPO, the RSPO loss function produces smoother and more stable gradients. This indicates that RSPO not only alleviates drastic fluctuations in gradient magnitude but also promotes a more stable and controllable optimization trajectory. This smoothness of gradient dynamics reduces the risk of gradient explosion or vanishing gradients, thereby significantly improving training stability.

[0097] The provided trends of length-controlled win rates (LC) of RSPO and DPO on Mistral-SFT as a function of training are compared. Figure 4As shown. The Win Rate (WR) represents the proportion of times the model's generated answer outperforms the reference model (GPT-4) in preference evaluation, measuring the overall answer quality. The Length Control Win Rate (LC) is the model's win rate under the condition that the generated answer length is close to the reference answer under additional constraints, aiming to evaluate whether the model still has a real quality improvement without length bias. Compared with DPO, the RSPO loss function achieves a higher Length Control Win Rate (LC) faster, indicating that this mechanism effectively alleviates the gradient imbalance and over-optimization problems of traditional DPO, resulting in better training performance for the model under complex preference data distributions.

[0098] This embodiment uses different methods to conduct performance comparison experiments on multiple benchmark subsets. The experimental results are based on the Mistral-Base (7B) model and report the strict match accuracy on GSM8K, ARC, TruthfulQA (TQA), MMLU, and IFEval tasks. The "average" is the average performance of all tasks. The experimental results are shown in Table 1.

[0099] Table 1. Performance comparison of different methods on multiple benchmark subsets

[0100]

[0101] As shown in Table 1, the RSPO loss function achieves significant performance improvements on various mainstream open-ended dialogue and instruction-following tasks. In mathematical and logical reasoning tasks, specifically the GSM8K and MMLU subsets, RSPO effectively enhances the learning signals in the high-quality alignment R1 category and reduces the interference of noisy samples on the training direction. Table 1 also shows that the RSPO method based on the Mistral-base model achieves strict matching accuracy improvements of 4.3 and 1.1 percentage points compared to DPO on the GSM8K and MMLU subsets, respectively, significantly outperforming standard DPO and other RLHF methods. These results demonstrate that RSPO, by introducing a reward-driven selective penalty mechanism, enables the model to maintain higher logical consistency and computational stability in multi-step reasoning and knowledge integration tasks, thereby improving the overall robustness of reasoning tasks.

[0102] In intelligent question-answering scenarios, specifically the AlpacaEval 2 and IFEval subsets, RSPO also performs exceptionally well on the Llama-3-Instruct model. Specifically, on the AlpacaEval 2 benchmark, the LengthControl Win Rate (LC) and Win Rate (WR) metrics reach 45.0% and 42.5%, respectively, representing an improvement of approximately 4.5% compared to the standard DPO. These results demonstrate that RSPO effectively balances the relationship between semantic consistency, factual accuracy, and the naturalness of the answer style, resulting in generated answers that better reflect user intent and reducing interference from ambiguous or conflicting preference samples.

[0103] Based on the aforementioned performance improvements, this embodiment deploys the RSPO-optimized model in an educational tutoring question-and-answer system, conducting practical tests on generative question-and-answer tasks for elementary school mathematics and Chinese language questions. In actual interactive testing, the system's answer accuracy rate improved by approximately 12% compared to the original model, and user satisfaction increased by 18%, significantly improving the model's usability and reliability in the education field. This demonstrates that RSPO not only improves the model's performance on standard benchmark tasks but also enhances the model's learning efficiency and consistency with human preferences in real-world applications.

[0104] In this embodiment, different training methods were used to train the Mistral-base (7B) and Llama-3-Instruct (8B) models on the AlpacaEval 2 benchmark. The performance of the Mistral-base (7B) and Llama-3-Instruct (8B) models after training is shown in Table 2.

[0105] Table 2 Performance comparison of different training methods

[0106]

[0107] In the AlpacaEval 2 benchmark, RSPO achieves an overall win rate improvement of approximately 4% to 5% compared to DPO, significantly outperforming the control DPO and other comparative methods. Furthermore, in multiple downstream tasks (such as MMLU, ARC, GSM8K, TruthfulQA, and IFEval), RSPO generally surpasses DPO in average scores, exhibiting stronger robustness and generalization ability, particularly in logical reasoning and factual consistency tasks. Notably, this invention adds almost no additional computational burden, only a slight increase in GPU memory and time overhead (approximately 3GB of GPU memory and 5 minutes of training time) to achieve significant performance improvements. In summary, RSPO, by constructing a reward-driven choice penalty mechanism, enables the model to dynamically adapt its learning focus when facing real human preference data, significantly improving alignment, generalization ability, and training stability. It possesses significant theoretical value and practical application potential, and can be widely applied in large language model-related fields such as intelligent assistants, question-answering systems, text generation, and decision reasoning.

[0108] Example 2:

[0109] This embodiment proposes an electronic device, including: one or more processors, and a memory for storing instructions, which, when executed by the one or more processors, cause the one or more processors to perform the preference alignment optimization method based on reward-driven selective penalty.

[0110] The electronic device may be a mobile phone, computer, or tablet computer, etc., and includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the preference alignment optimization method based on reward-driven selective penalty as described in the embodiments. It is understood that the electronic device may also include an input / output (I / O) interface and communication components.

[0111] The processor is used to execute all or part of the steps in the preference alignment optimization method based on reward-driven selective penalty as described in the above embodiments. The memory is used to store various types of data, which may include, for example, instructions for any application or method in the electronic device, as well as application-related data.

[0112] The processor may be implemented as an Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), controller, microcontroller, microprocessor, or other electronic components, and is used to execute the preference alignment optimization method based on reward-driven selective penalty described in the above embodiments.

[0113] Example 3:

[0114] This embodiment proposes a computer-readable storage medium that stores executable instructions. When these instructions are executed, if they are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0115] The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the preference alignment optimization method based on reward-driven selective punishment described in the various embodiments of this application.

[0116] The aforementioned storage media include: flash memory, hard disks, multimedia cards, card-type memory (e.g., SD (Secure Digital Memory Card) or DX (Memory Data Register, MDR) memory), random access memory (RAM), static random-access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, disks, optical discs, servers, APP (Application) app stores, and other media capable of storing program verification codes. These media store computer programs, which, when executed by a processor, can implement the various steps of the reward-driven selective penalty-based preference alignment optimization method described above.

[0117] Example 4:

[0118] This embodiment proposes a computer program product, including a computer program or instructions, which, when executed by a processor, implements the aforementioned preference alignment optimization method based on reward-driven selective punishment.

[0119] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a computer program product.

[0120] The various embodiments in this application are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0121] The scope of protection of this application is not limited to the embodiments described above. Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from the scope and spirit of this disclosure. If such modifications and variations fall within the scope of this disclosure and its equivalents, then the intent of this disclosure also includes these modifications and variations.

Claims

1. A preference alignment optimization method based on reward-driven selective penalty, characterized in that, Includes the following steps: Acquire intelligent question answering training data, process the intelligent question answering training data, and construct an intelligent question answering training sample set, including several intelligent question answering training samples; The large language model to be optimized is used as the strategy model. and reference model ; Use the intelligent question answering training sample set to train the policy model Training is performed using a reference model. Measuring the policy model before and after training The offset magnitude is used to obtain the optimized large language model, including: Input the intelligent question answering training samples into the policy model respectively. and reference model Perform forward reasoning and compute the policy model. The probability distribution of the policy logarithm generated from the training samples of the intelligent question answering system is used to calculate the reference model. The reference logarithm of the training samples for intelligent question answering generates a probability distribution; Calculate the implicit reward value of the response; Establish a sample classification mechanism to divide the intelligent question answering training samples into multiple categories based on the implicit reward value of the response; Construct the RSPO loss function and use the RSPO loss function to evaluate the policy model. Training is performed to obtain a trained policy model. That is, the optimized large language model.

2. The preference alignment optimization method based on reward-driven selective penalty according to claim 1, characterized in that, The intelligent question-answering training data includes: several input prompts. Preference response Non-preference response ; Processing the training data for intelligent question answering includes: processing each input suggestion Formatting and normalization processes are performed, and each input suggestion is constructed using a word segmenter. input token sequence Generate suggestions for each input. Corresponding intelligent question answering training samples .

3. The preference alignment optimization method based on reward-driven selective penalty according to claim 1, characterized in that, The policy logarithmic generation probability distribution includes the policy preference logarithmic generation probability distribution. The probability distribution of non-preference log generation of strategies ; Computational strategy model The probability distribution of the policy logarithm for the training samples of intelligent question answering is generated using the following method: training samples for intelligent question answering Input policy model Perform forward reasoning to obtain the policy model. Given input prompts Under the condition of preference response Non-preference response Token-level preference generation probability distribution and token-level non-preference generation probability distribution ; Probability distribution generated based on token-level preferences and token-level non-preference generation probability distribution Computational strategy model Given input prompts Under the condition of preference response Non-preference response The log-generated probability distribution of strategy preferences The probability distribution of non-preference log generation of strategies .

4. The preference alignment optimization method based on reward-driven selective penalty according to claim 3, characterized in that, The reference logarithmic generation probability distribution includes the policy preference logarithmic generation probability. and the reference non-preference log generation probability ; Computational reference model The reference logarithm of the training samples for intelligent question answering generates a probability distribution, specifically using the following method: training samples for intelligent question answering Input reference model Perform forward reasoning to obtain the reference model. Given input prompts Under the condition of preference response Non-preference response Reference preference log generation probability and the reference non-preference log generation probability .

5. The preference alignment optimization method based on reward-driven selective penalty according to claim 4, characterized in that, The method for calculating the implicit reward value of the response is as follows: Calculate the implicit reward value of the preference response based on the implicit reward formula in the DPO theory. Implicit reward value of non-preference response As shown in the formula below: in, This is the implicit reward formula in the DPO theory. To control the temperature coefficient of the reward scale in the preference response, To control the temperature coefficient of the reward scale for non-preference responses.

6. The preference alignment optimization method based on reward-driven selective penalty according to claim 5, characterized in that, A sample classification mechanism is established to divide the intelligent question answering training samples into multiple categories based on the implicit reward value of the response. The specific method is as follows: Implicit reward value based on preference response Implicit reward value of non-preference response The training samples for intelligent question answering are divided into multiple categories, including: high-quality alignment R1, low-quality alignment R2, incorrect preference R3, and ambiguous preference R4. when At that time, the training samples for intelligent question answering are assigned to the high-quality aligned category R1; when At that time, the training samples for intelligent question answering were assigned to the low-quality aligned category R2; when At that time, the training samples for intelligent question answering are classified into the error preference category R3; when At that time, the training samples for intelligent question answering are classified into the fuzzy preference category R4; in, and All are learnable or preset thresholds.

7. The preference alignment optimization method based on reward-driven selective penalty according to claim 6, characterized in that, The RSPO loss function is constructed as follows: Constructing a nonlinear dynamic weighting function As shown in the formula below: in, For the sigmoid function, This is the minimum weight bias term; Implicit reward value in preference response Implicit reward value of non-preference response Based on this, a nonlinear dynamic weighting function is introduced. Construct the RSPO loss function As shown in the formula below: in, This is the training sample set for intelligent question answering.

8. An electronic device, characterized in that, include: One or more processors, and a memory for storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the preference alignment optimization method based on reward-driven selective penalty as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed, cause the processor to perform the preference alignment optimization method based on reward-driven selective penalty as described in any one of claims 1-7.

10. A computer program product, characterized in that, Includes a computer program or instructions that, when executed by a processor, implement the preference alignment optimization method based on reward-driven selective punishment as described in any one of claims 1-7.

Citation Information

Cited By

  • Generative virtual tutoring teacher model training method based on preference classification

    CN121936502A

  • Question and answer method and device based on large language model, equipment and medium

    CN121936631A

  • AI-enabled intelligent aggregation method and system for virtual power plant resources

    CN122418891A

  • AI-enabled intelligent aggregation method and system for virtual power plant resources

    CN122418891B