Safety preference modeling method based on data distillation

The method addresses the challenge of noisy and biased reward model training data by refining it through data distillation and optimization, enhancing safety and generalization, and reducing reliance on manual annotations.

CN120317099APending Publication Date: 2025-07-15HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510356804.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing reward model training data is difficult to meet the requirements of high signal-to-noise ratio, low deviation and wide coverage. The existence of noise and deviation affects the training effect, resulting in unsafe model behavior.

Method used

Multi-round data distillation and optimization methods are used to generate a referee model through the initial training set, extract inconsistent samples, generate evaluation labels and evaluation texts, perform direct preference optimization, and finally generate a high-quality distillation training set for reward modeling.

Benefits of technology

It significantly improves the quality of the reward model training data, reduces noise and deviation, enhances the model's safe generalization ability and the security adaptability of diversified instruction responses, reduces dependence on manual labeled data, and improves training efficiency and cost-effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120317099A_ABST
    Figure CN120317099A_ABST
Patent Text Reader

Abstract

The invention relates to a safety preference modeling method based on data distillation, belongs to the technical field of artificial intelligence safety, and aims to solve the problems that current reward model training data is difficult to meet the requirements of high signal-to-noise ratio, low deviation and wide coverage, and noise and deviation influence the training effect and cause unsafe model behaviors. According to the method, the base large model and the initial training set are obtained, multi-round distillation and optimization are carried out, the steps of building a prompt template, extracting inconsistent samples, optimizing the model and the like are included, and finally reward modeling is carried out on the base large model based on the high-quality training set. According to the method, the influence of data noise and deviation on security preference modeling is effectively reduced, the dependence on manual annotation is reduced, the generalization ability of the security preference model is enhanced, the accurate fitting of human security preferences is improved, the output quality and application value of a large model in a security sensitive scene are remarkably optimized, and the method is suitable for popularization and application. Therefore, the safety and the reliability of the large model in the open environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a security preference modeling method based on data distillation, belonging to the technical field of artificial intelligence security. Background Art

[0002] Reward modeling is an important part in the training of large models to ensure the security and reliability of the output. Its core goal is to train a reward model to evaluate the security of a given instruction and response, and output a scalar reward value reflecting the security quality of the response. This reward value can be used to guide the fine-tuning of the large model, so that the generated text is more in line with human security preferences and expectations. The training of the reward model usually requires a large amount of labeled data, which includes instructions, responses, and manually labeled security preference tags, for the model to learn the security advantages and disadvantages of different responses. A high-quality reward model can not only improve the security performance of the large model, but also help to avoid unsafe, unethical or even harmful content in the model output. Therefore, reward modeling is an important way to achieve the alignment of the security of large models with human intentions in an open environment.

[0003] Ideal training data for the reward model should have a high signal-to-noise ratio, low bias, and be able to cover a wide range of instruction and response types. However, current dataset construction methods often struggle to meet these requirements simultaneously. Noise in the dataset, such as incorrect annotations or inconsistent scoring, will reduce the training effect of the reward model and affect its security generalization ability. In addition, the bias existing in the dataset, such as the preference for specific types of instructions or responses, may cause the model to learn unsafe behaviors. Therefore, improving the security and reliability of the reward model by enhancing the quality of the reward model training data is a key problem that urgently needs to be solved currently. Summary of the Invention

[0004] The present invention proposes a security preference modeling method based on data distillation to solve the problems that the current reward model training data is difficult to meet the requirements of high signal-to-noise ratio, low bias and wide coverage, and the existence of noise and bias affects the training effect and leads to unsafe model behaviors.

[0005] A security preference modeling method based on data distillation includes the following steps:

[0006] S1, obtaining a base large model and an initial training set including instruction, response pairs, and evaluation labels;

[0007] S2, performing the first round of distillation based on the initial training set to generate a first referee model and a first distilled training set;

[0008] S3, extracting inconsistent samples from the initial training set and the first distilled training set;

[0009] S4. Use the base large model to decode the inconsistent samples to generate evaluation labels and evaluation texts;

[0010] S5. Based on the evaluation labels and evaluation texts, directly optimize the preference of the base large model to obtain a second referee model;

[0011] S6. Use the second referee model to distill the initial training set again to generate a second distilled training set;

[0012] S7. Perform reward modeling on the base large model based on the second distilled training set.

[0013] Further, S2 includes the following steps:

[0014] S21. Construct a first prompt template containing instruction-response pairs, and the first prompt template is used to clarify the evaluation rules, and the evaluation rules include the accuracy of instruction execution, the integrity of response content, and objectivity requirements;

[0015] S22. Use the first prompt template, the instructions and response pairs in the initial training set to construct inputs, and use the evaluation labels in the initial training set to construct outputs, and construct input-output training data;

[0016] S23. Use the input-output training data to train the base large model to obtain a first referee model;

[0017] S24. Use the first referee model to re-decode based on the inputs in S22 to generate new evaluation labels, and combine the input instructions and responses to construct a first distilled training set.

[0018] Further, the general training loss function in S23 is:

[0019]

[0020] where k represents the token index in the output sequence, n represents the number of tokens in the output sequence, y k represents the kth token in the output sequence, y <k represents all tokens before the kth token in the output sequence, P(y k |y <k ) represents the probability that the model predicts the current token y k based on the previous k - 1 tokens.

[0021] Further, in S3, the extraction of the inconsistent samples includes the following steps:

[0022] S31. Compare the evaluation labels in the initial training set with the new evaluation labels generated for the corresponding instructions in the first distilled training set;

[0023] S32. Screen out the samples with inconsistent evaluation labels and determine them as inconsistent samples.

[0024] Further, S4 includes the following steps:

[0025] S41. Construct a second prompt template containing instruction-response pairs to clarify the evaluation rules and prompt the large model to generate evaluation texts;

[0026] S42. Use the second prompt template and inconsistent samples to construct input data;

[0027] S43. Use the base large model to decode the input data to generate output data, and the output data includes two parts: evaluation labels and evaluation texts.

[0028] Further, in S5, the process of direct preference optimization includes:

[0029] S51. Construct a third prompt template containing instruction-response pairs to clarify the evaluation rules and prompt the large model to reason before outputting evaluation labels;

[0030] S52. Use the third prompt template and the evaluation labels and evaluation texts of inconsistent samples to construct preference comparison data;

[0031] S53. Use the preference comparison data to perform direct preference optimization on the base large model to obtain a second referee model.

[0032] Further, the training loss function of direct preference optimization in S53 is:

[0033]

[0034] Among them, x represents the input data, y w and y l represent the output data corresponding to inconsistent samples, including evaluation texts and evaluation labels, π θ represents the probability distribution of the current training model, π ref represents the probability distribution of the base large model, β is a parameter controlling the deviation from the base large model, and σ is the standard Sigmoid function.

[0035] Further, S6 includes the following steps:

[0036] S61. Use the third prompt template and the instruction-response pairs of the initial training set to construct input data;

[0037] S62. Use the second referee model to re-decode based on the input in S61, generate new evaluation labels, and combine the input instructions and responses to construct a second distillation training set.

[0038] Further, in S7, the training loss function of reward modeling is specifically:

[0039]

[0040] where σ is the standard Sigmoid function, and r ψ (x, y) represents the score of the reward model for the instruction x and the response y, and y c represents the better response of the second distillation training set, and y r represents the worse response of the second distillation training set. The r ψ (x, y) is obtained by connecting a reward head to the large model. The reward head is a parameterized neural network module used to predict the reward score according to the input instruction and response.

[0041] A computer device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the above-mentioned method for secure preference modeling based on data distillation.

[0042] Advantages of the present invention: The method for secure preference modeling based on data distillation of the present invention, through multiple rounds of data distillation and optimization, not only significantly improves the quality of the training data of the reward model, reduces noise and bias, but also enhances the secure generalization ability of the model while enhancing its secure adaptability to diverse instruction and response types. The method of the present invention effectively reduces the dependence on a large amount of manually labeled data, optimizes the data processing process, thereby improving the training efficiency and cost-effectiveness. In addition, the high-quality reward model trained by the present invention can more accurately capture and reflect human secure preferences, making the large model more compliant with security requirements when generating text, further reducing the risk of output of insecure, unethical, and harmful content. Overall, the present invention provides strong support for the security optimization and application of large models in open environments. Description of the Drawings

[0043] Figure 1 is the flowchart of the method for secure preference modeling based on data distillation of the present invention;

[0044] Figure 2 is the flowchart of S2 in the method for secure preference modeling based on data distillation of the present invention;

[0045] Figure 3 is the flowchart of S3 in the method for secure preference modeling based on data distillation of the present invention;

[0046] Figure 4 is the flowchart of S4 in the method for secure preference modeling based on data distillation of the present invention;

[0047] Figure 5 It is the flowchart of S5 in a security preference modeling method based on data distillation of the present invention;

[0048] Figure 6 It is the flowchart of S6 in a security preference modeling method based on data distillation of the present invention. Specific implementation manner

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] Refer to Figure 1 As shown, a security preference modeling method based on data distillation includes the following steps:

[0051] S1, obtaining a base large model and an initial training set including instruction, response pair, and evaluation label;

[0052] S2, performing the first round of distillation based on the initial training set to generate a first referee model and a first distilled training set;

[0053] S3, extracting inconsistent samples from the initial training set and the first distilled training set;

[0054] S4, using the base large model to decode the inconsistent samples to generate evaluation labels and evaluation texts;

[0055] S5, directly optimizing the preference of the base large model based on the evaluation labels and evaluation texts to obtain a second referee model;

[0056] S6, using the second referee model to distill the initial training set again to generate a second distilled training set;

[0057] S7, performing reward modeling on the base large model based on the second distilled training set.

[0058] Specifically, the core idea of the present invention is to improve the quality of the training data of the reward model through the distillation and optimization processes, thereby enhancing the security of the reward model. First, the base large model is preliminarily trained using the initial training set to obtain the first referee model. Then, by comparing the evaluation labels generated by the initial training set and the first referee model, inconsistent samples are identified. These samples may contain errors or noises in the original data. Next, the base large model is used to re-evaluate these inconsistent samples to generate more accurate evaluation labels and explanatory evaluation texts. Subsequently, direct preference optimization is performed on the base large model using this high-quality evaluation data to obtain a more powerful second referee model. Finally, the second referee model is used to distill the initial training set again to obtain a second distilled training set with higher quality, and the reward modeling of the base large model is carried out using it. Through the present invention, the training data can be continuously refined and optimized, and finally a reward model with excellent performance can be trained.

[0059] It should be noted that in S1, the base large model should have strong generalization ability and certain basic performance so as to effectively absorb and improve the data quality in the subsequent distillation process.

[0060] The base large model can select models that have been pre-trained on a large-scale text corpus as the base large model, such as LLaMA, Qwen, etc. These models already have certain language understanding and generation capabilities, which helps to improve the effect of the entire reward modeling process.

[0061] It should be noted that in S5, the goal of direct preference optimization is to enable the model to better understand and simulate human preferences. Specifically, the deviation between the model and the base large model can be balanced by adjusting the hyperparameters (such as the β value) in the optimization process to ensure that the model does not deviate too much from the original knowledge system while improving. In addition, various optimization algorithms, such as Adam, SGD, etc., can also be adopted to find the most suitable optimization strategy for the current task.

[0062] Refer to Figure 2 As shown, further, S2 includes the following steps:

[0063] S21, construct a first prompt template containing instruction-response pairs, and the first prompt template is used to clarify the evaluation rules, and the evaluation rules include the accuracy of instruction execution, the integrity and objectivity requirements of the response content;

[0064] S22, construct the input using the first prompt template, the instructions and response pairs of the initial training set, and construct the output using the evaluation labels of the initial training set to construct input-output training data;

[0065] S23, use the input-output training data to train the base large model to obtain the first referee model;

[0066] S24. Re - decode using the first referee model based on the input of S22, generate new evaluation labels, and combine the input instructions and responses to construct the first distillation training set.

[0067] Specifically, the first prompt template constructed in S21 should not only clarify the evaluation rules but also consider how to make the model's output as close as possible to human judgment criteria. Specifically, more guiding information can be added to the prompt template, such as requirements for the logic, coherence, and factuality of the response, which can ensure that the model considers more dimensions of information when generating evaluation labels and improve the comprehensiveness and accuracy of the evaluation.

[0068] Furthermore, the general training loss function in S23 is:

[0069]

[0070] where k represents the token index in the output sequence, n represents the number of tokens in the output sequence, y k represents the k - th token in the output sequence, y <k represents all tokens before the k - th token in the output sequence, and P(y k |y <k ) represents the probability that the model predicts the current token y k based on the previous k - 1 tokens.

[0071] Referring to Figure 3 as shown, furthermore, in S3, the extraction of the inconsistent samples includes the following steps:

[0072] S31. Compare the evaluation labels in the initial training set with the new evaluation labels generated for the corresponding instructions in the first distillation training set;

[0073] S32. Screen out the samples with inconsistent evaluation labels and determine them as inconsistent samples.

[0074] Specifically, when comparing the evaluation labels in S31, since there are only two cases for the evaluation labels (the first response is good or the second response is good), the labels of the corresponding samples in the initial training set and the first distillation training set can be directly compared. When the two labels are inconsistent, that is, one label indicates that the first response is good while the other indicates that the second response is good, the sample can be determined as an inconsistent sample.

[0075] Referring to Figure 4 as shown, furthermore, S4 includes the following steps:

[0076] S41. Construct a second prompt template containing instruction - response pairs to clarify the evaluation rules and prompt the large model to generate evaluation text;

[0077] S42. Construct input data using the second prompt template and inconsistent samples;

[0078] S43. Use the base large model to decode the input data to generate output data, where the output data includes two parts: an evaluation label and an evaluation text.

[0079] Specifically, the main purpose of S4 is to let the base large model re-evaluate inconsistent samples, not only generating an evaluation label but also generating an explanatory evaluation text. This can help the model better understand why a certain response is better, so as to learn more accurate evaluation criteria in subsequent training.

[0080] It should be noted that when constructing input data in S42, due to the disagreement between the two evaluation labels of inconsistent samples, the model can only normally decode one of the evaluation labels. Therefore, for the other evaluation label, it needs to be added to the input data in a manually spliced manner to ensure that the model can generate corresponding evaluation texts based on complete evaluation information. This processing method can ensure that the model takes into account both the content of the instructions and responses when generating evaluation texts and can fully understand the judgment basis of different evaluation labels.

[0081] Refer to Figure 5 As shown, further, in S5, the process of direct preference optimization includes:

[0082] S51. Construct a third prompt template containing instruction-response pairs to clarify the evaluation rules and prompt the large model to reason before outputting the evaluation label;

[0083] S52. Use the third prompt template and the evaluation labels and evaluation texts of inconsistent samples to construct preference comparison data;

[0084] S53. Use the preference comparison data to perform direct preference optimization on the base large model to obtain a second referee model.

[0085] Specifically, the third prompt template constructed in S51 needs to pay more attention to the reasoning process to help the model better understand the relationship between the evaluation label and the evaluation text.

[0086] Add guidance on the reasoning process to the third prompt template, such as asking the model to explain why a certain response is better, which can improve the model's reasoning ability and the accuracy of evaluation.

[0087] Further, the direct preference optimization training loss function in S53 is:

[0088]

[0089] where x represents the input data, y w and y lDenote the output data corresponding to inconsistent samples, including the evaluation text and the evaluation label, π θ Denote the probability distribution of the current training model, π ref Denote the probability distribution of the base large model, where β is a parameter controlling the deviation from the base large model, and σ is the standard Sigmoid function.

[0090] Refer to Figure 6 As shown, further, in S6, the following steps are included:

[0091] S61, use the third prompt template and the instructions and responses of the initial training set to construct the input data;

[0092] S62, use the second referee model to re-decode based on the input in S61, generate new evaluation labels, and combine the input instructions and responses to construct the second distillation training set.

[0093] Specifically, the main goal of S6 is to use the second referee model optimized by direct preference to re-distill the initial training set to obtain higher-quality training data. By constructing the input data using the third prompt template, it can ensure that the model fully considers the inference process when generating new evaluation labels, thereby improving the accuracy and interpretability of the evaluation.

[0094] Further, in S7, the training loss function of reward modeling is specifically:

[0095]

[0096] Among them, σ is the standard Sigmoid function, r ψ (x, y) represents the score of the reward model for the instruction x and the response y, y c Represents the better response of the second distillation training set, y r Represents the worse response of the second distillation training set, where the r ψ (x, y) is obtained by connecting a reward head to the large model. The reward head is a parameterized neural network module used to predict the reward score according to the input instruction and response.

[0097] A computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the above-mentioned method for secure preference modeling based on data distillation.

[0098] Specifically, the present invention realizes a security preference modeling method based on data distillation through a specific computer device (including a memory, a processor, and corresponding computer programs), which significantly improves the quality of the training data of the reward model, reduces the noise and bias in the data, and at the same time enhances the generalization ability of the model in the security alignment task and the security adaptability to diverse instruction and response types; effectively reduces the dependence on a large amount of manually labeled data, optimizes the data processing process, thereby improving the training efficiency and cost-effectiveness; the trained high-quality reward model can accurately capture and reflect human security preferences, making the text content generated by the large model more in line with security requirements, significantly reducing the output risk of unsafe, unethical or harmful content, and providing strong technical support for the security optimization and application of the large model in an open environment, which plays an important role in promoting the development of the artificial intelligence technology field.

[0099] Through multiple rounds of data distillation and optimization, the present invention not only significantly improves the quality of the training data of the reward model, reduces noise and bias, but also enhances the security adaptability of the model to diverse instruction and response types while improving the security generalization ability of the model. This method effectively reduces the dependence on a large amount of manually labeled data, optimizes the data processing process, thereby improving the training efficiency and cost-effectiveness. In addition, the trained high-quality reward model can more accurately capture and reflect human security preferences, making the large model more in line with security requirements when generating text, and significantly reducing the output risk of unsafe, unethical and harmful content. Overall, the present invention provides strong technical support for the security optimization and application of the large model in an open environment.

[0100] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. In addition, in this article, "front", "rear", "left", "right", "up" and "down" are all referenced with respect to the placement state shown in the drawings.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A security preference modeling method based on data distillation, characterized in that, It includes the following steps: S1. Obtain a base large model and an initial training set including instructions, response pairs, and evaluation labels; S2. Perform the first round of distillation based on the initial training set to generate a first referee model and a first distilled training set; S3. Extract inconsistent samples from the initial training set and the first distilled training set; S4. Use the base large model to decode the inconsistent samples to generate evaluation labels and evaluation texts; S5. Based on the evaluation labels and evaluation texts, perform direct preference optimization on the base large model to obtain a second referee model; S6. Use the second referee model to distill the initial training set again to generate a second distilled training set; S7. Perform reward modeling on the base large model based on the second distilled training set.

2. The security preference modeling method based on data distillation according to claim 1, wherein S2 includes the following steps: S21. Construct a first prompt template including instruction-response pairs, which is used to clarify the evaluation rules. The evaluation rules include the accuracy of instruction execution, the integrity of response content, and objectivity requirements; S22. Use the first prompt template, the instructions and response pairs of the initial training set to construct the input, and use the evaluation labels of the initial training set to construct the output, and construct input-output training data; S23. Use the input-output training data to train the base large model to obtain a first referee model; S24. Use the first referee model to re-decode based on the input in S22 to generate new evaluation labels, and combine the input instructions and responses to construct a first distilled training set.

3. The security preference modeling method based on data distillation according to claim 2, characterized in that The general training loss function in S23 is: where k represents the token index in the output sequence, n represents the number of tokens in the output sequence, y k represents the k-th token in the output sequence, y <k represents all tokens before the k-th token in the output sequence, P(y k |y <k ) represents the probability that the model predicts the current token y k based on the previous k - 1 tokens.

4. A security preference modeling method based on data distillation according to claim 1, characterized in that In S3, the extraction of the inconsistent samples includes the following steps: S31. Compare the evaluation labels in the initial training set with the new evaluation labels generated for the corresponding instructions in the first distilled training set; S32. Screen out the samples with inconsistent evaluation labels and determine them as inconsistent samples.

5. A security preference modeling method based on data distillation according to claim 1, characterized in that S4 includes the following steps: S41. Construct a second prompt template including instruction-response pairs, which is used to clarify the evaluation rules and prompt the large model to generate evaluation texts; S42. Use the second prompt template and the inconsistent samples to construct input data; S43. Use the base large model to decode the input data to generate output data, and the output data includes two parts: evaluation labels and evaluation texts.

6. The security preference modeling method based on data distillation according to claim 1, characterized in that In S5, the process of direct preference optimization includes: S51. Construct a third prompt template including instruction-response pairs, which is used to clarify the evaluation rules and prompt the large model to perform reasoning before outputting evaluation labels; S52. Use the third prompt template and the evaluation labels and evaluation texts of the inconsistent samples to construct preference comparison data; S53. Use the preference comparison data to perform direct preference optimization on the base large model to obtain a second referee model.

7. The security preference modeling method based on data distillation according to claim 6, characterized in that The direct preference optimization training loss function in S53 is: Among them, x represents the input data, and y w and y l represent the output data corresponding to inconsistent samples, including the evaluation text and evaluation labels, and π θ represents the probability distribution of the current training model, and π ref represents the probability distribution of the base large model. β is a parameter that controls the deviation from the base large model, and σ is the standard Sigmoid function.

8. The security preference modeling method based on data distillation according to claim 1, characterized in that In S6, it includes the following steps: S61. Use the third prompt template and the instructions and response pairs of the initial training set to construct input data; S62. Use the second referee model to re-decode based on the input in S61 to generate new evaluation labels, and combine the input instructions and responses to construct a second distilled training set.

9. The security preference modeling method based on data distillation according to claim 1, wherein In S7, the training loss function of reward modeling is specifically: where σ is the standard Sigmoid function, r ψ (x, y) represents the score of the reward model for the instruction x and the response y, and y c represents the better response of the second distillation training set, and y r represents the worse response of the second distillation training set, and the r ψ (x, y) is obtained by connecting a reward head to the large model. The reward head is a parameterized neural network module for predicting the reward score according to the input instruction and response.

10. A computer device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement a method for secure preference modeling based on data distillation according to any one of claims 1-9.