A Method for Privacy Protection of Content Generated by Large Language Models Based on RLHF
The RLHF-based method integrates privacy protection into large language model training using a dual-effort model with multi-dimensional scoring and regret theory, effectively balancing privacy and quality, addressing the challenge of privacy protection in large language models.
Patent Information
- Application Number
- CN202510300832.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The existing large language model has the risk of privacy leakage in terms of privacy protection, and existing methods are difficult to find a balance between privacy protection and generation quality.
Using an RLHF-based method, a multi-dimensional privacy evaluation mechanism and a gated network are designed by decoupling the benefit model and cost model, and the training process is optimized by combining the regret theory, and the privacy protection strategy of generated content is dynamically adjusted.
It realizes adaptive and dynamic privacy protection, reduces the conflict between privacy protection and model performance, and improves the quality of generated content and privacy protection efficiency.
Smart Images

Figure CN119830350B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer artificial intelligence and large model security, and specifically to a method for protecting the privacy of content generated by large language models based on RLHF. Background Art
[0002] In recent years, large language models based on deep learning (such as GPT series, BERT, LLaMA, etc.) have made remarkable progress in the field of natural language processing. These models can generate fluent and coherent text through training with massive amounts of data and are widely used in tasks such as intelligent question answering, content generation, translation, and text summarization. However, as the application scope of large language models continues to expand, their potential privacy protection issues have gradually attracted widespread attention.
[0003] During the training process of large language models, they usually rely on a large amount of publicly available data that has not been strictly screened, which may contain sensitive information (such as personal identity information, medical records, etc.). During the inference stage, the model may inadvertently generate content containing such sensitive information, leading to the risk of privacy leakage. In addition, attackers may also infer private content in the input or training data based on the model output through reverse inference, etc., which exacerbates the threat of privacy leakage.
[0004] Currently, the privacy protection methods for large language models mainly include data screening and desensitization, differential privacy, federated learning, and output filtering. Data screening and desensitization reduce the risk of privacy leakage by removing or processing sensitive information in the training data, but due to the diversity of sensitive information and the high processing cost, privacy issues cannot be completely avoided. Differential privacy provides a certain degree of privacy protection by adding noise to the model parameters or gradients during the training stage, but it may lead to performance degradation when generating high-quality text. Federated learning, as a distributed training method, avoids direct data leakage by only sharing parameter updates, but it faces problems such as high communication costs, parameter leakage risks, and uneven data distribution. Output filtering quickly reduces the risk of privacy leakage by post-processing the generated content, but this method usually relies on predefined rules, lacks flexibility, and may affect the coherence and integrity of the text. Therefore, it is still difficult to find an ideal balance between privacy protection and model performance among existing methods. Against this background, the application of Reinforcement Learning from Human Feedback (RLHF) technology shows great potential. By combining the reinforcement learning framework, RLHF can dynamically adjust the model behavior to make the generated content more in line with specific preferences. However, currently RLHF is mainly used to improve the generation quality, and its application to privacy protection issues is still in the exploratory stage. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a method for protecting the privacy of content generated by large language models based on RLHF. By introducing a reinforcement learning framework, the privacy protection objective is integrated into the model training process to achieve dynamic and adaptive protection of the privacy of the generated content and to achieve a better balance between privacy protection and generation quality.
[0006] To achieve the above object, the specific technical solution adopted by the present invention is as follows:
[0007] A method for protecting the privacy of content generated by a large language model based on RLHF, comprising the following steps:
[0008] S1. Training a base model based on supervised instruction fine-tuning: Based on a dataset of instructions for fine-tuning annotated for specific downstream tasks of generative type, training an open-source base model as an instruction fine-tuning model;
[0009] S2. Decoupling usefulness and harmlessness: Decomposing the original reward model into a benefit model and a cost model;
[0010] S3. Expanding the preference understanding of the cost model: Expanding the linear mapping layer of the cost model so that the cost model has a multi-dimensional refined scoring head structure, and adding a gating network to dynamically analyze the weight of each dimension for privacy protection for the input content;
[0011] S4. Training a multi-dimensional scoring cost model: Training a multi-dimensional scoring cost model based on an annotated dataset containing multi-dimensional scores and multi-dimensional weights;
[0012] S5. Optimizing the traditional preference probability calculation formula: Improving the traditional preference probability calculation formula based on the weighted score and the Bradley-Terry preference model, where the weighted score is the weighted score Score after obtaining the dynamic analysis weight of the input content based on the multi-dimensional scoring cost model obtained in step S4 weight ;
[0013] S6. Constructing a new method for calculating human privacy preference probability: Constructing a new method for calculating human privacy preference probability by combining multi-dimensional scoring and regret theory;
[0014] S7. Training the benefit model and the cost model: Training the benefit model and the cost model using the preference pair dataset;
[0015] S8. Training the instruction fine-tuning model by reinforcement learning: Based on the benefit model and the cost model and the new method for calculating human privacy preference probability, training the instruction fine-tuning model by proximal policy optimization.
[0016] Furthermore, in step S2, the last unembedded layer of the language model pre-trained based on Transformer is removed, and an additional linear layer is added to the last Transformer layer to construct the benefit model and the cost model; given any text, the benefit model and the cost model will assign a scalar reward value to the last token, which are used as indicators to evaluate the usefulness and privacy of the answer respectively.
[0017] Furthermore, in step S3, an additional linear layer of the cost model obtained in step S2 is expanded into five separate linear layers to perform fine-grained scoring on compliance, sensitivity, correctness, objectivity, and credibility respectively; and a gating network designed based on the idea of the mixture of experts model is added at an additional linear layer of the cost model obtained in step S2 to dynamically analyze the impact of each dimension on privacy protection for the input content and obtain different weights.
[0018] Furthermore, in step S4, first calculate the mean square error between the score of each dimension output by the multi-dimensional scoring cost model and the labeled score, and the mean square error between the weight of each dimension generated by the multi-dimensional scoring cost model and the labeled weight, and use these as the main optimization objectives to obtain the following loss function with a regularization term:
[0019]
[0020] In the formula, represents the scoring error loss of the i-th dimension, represents the weight error loss of the i-th dimension; d is the number of dimensions of the score, and λ is the regularization coefficient, which controls the influence degree of the regularization term on the total loss; The regularization term of the score, The regularization term of the weight.
[0021] Then, the loss function is used to penalize the situation where the variance of the score prediction and weight prediction of the model on a single sample is too large, avoiding the generation of unreasonable extreme values in a single sample, so as to output a smoother and more stable score and weight distribution.
[0022] Furthermore, in step S5, first, after obtaining the dynamic analysis weight of the input content based on the cost model obtained in step S4, the weighted score Score weight can be obtained, and then the weighted score Score weight is used to replace the score of the single evaluation, thereby increasing the probability that the model selects answers that conform to human privacy preferences. The new preference probability calculation formula is:
[0023]
[0024] In the formula, P(ywin > y loss |x) represents the given input x, and the preference model judges y win is better than y loss probability; Score weight (y win |x) is the numerical value of the weighted score for the input x and the output y win ; Score weight (y loss |x) is the numerical value of the weighted score for the input x and the output y loss ; exp is used to convert the score into a probability, forming a softmax function to ensure that the sum of probabilities is 1.
[0025] Furthermore, in the step S6, regret is defined as the deviation between each trajectory segment and the optimal decision, which is used to measure the loss in the expected return; regret can be decomposed into partial return and state value change. The former is the difference between the reward of each state-action pair and the optimal reward, and the latter is the value difference between the start and end states of the trajectory segment. In the preference model based on the Boltzmann rational distribution, the discounted partial return of the trajectory segment is calculated from the reward model score, the discount factor, and the total time reward of the trajectory.
[0026] Furthermore, in the step S6, in the multi-dimensional scoring cost model obtained in step S4, the score of the state representation model in each dimension is represented by the score, and the action is represented by the weight output by the gating network, which is used to adjust the score in each dimension; first, calculate the goodness (negative regret, advantage) of the score adjusted by a certain action weight relative to the baseline (unadjusted score) through the following formula: where r i (σ + ) is the score of a better answer in the i-th dimension, and the baseline return is the average score of all samples in the i-th dimension;
[0027] Then, weight the advantage of each dimension by the gating weight where γ i is the weight of a better answer in the i-th dimension. Combining the discount factor ω, calculate the future weighted total advantage: where t is the time step and T is the sequence length (time step size);
[0028] Then, for each pair of samples in each batch, use the Boltzmann distribution to calculate the preference (i.e., regret probability) between the "better" and "worse" answers, and exponentiate the sum of the weighted advantages of each dimension:
[0029]
[0030] In the formula, P(σ + > σ- ) represents the probability that the model judges that a better answer is superior to a worse answer for a given sample pair; represents the cumulative weighted advantage value of the better answer, represents the cumulative weighted advantage value of the worse answer; exp is used to convert the score into a probability, forming a softmax function to ensure that the sum of probabilities is 1.
[0031] Finally, calculate the preference probability between the two weighted answers.
[0032] Furthermore, in step S7, the benefit model is used to evaluate the usefulness of the generated content. By maximizing the probability that y win is better than y loss , the parameters φ of the benefit model are adjusted to construct a loss function for training the benefit model: In the formula, represents the expected value, D R represents the training dataset of the benefit model, (x, y win , y loss ) is a triple sampled from the dataset, where x is the input data, y win and y loss are better and worse answers to the input data respectively; R φ (y win , x) is the probability function of the benefit model for the input x and the output better answer y win , R φ (y loss , x) is the probability function of the benefit model for the input x and the output worse answer y loss ; logσ is the logarithmic form of the logistic sigmoid function, used to convert the probability to logits.
[0033] Furthermore, in step S7, the cost model is used to evaluate the privacy protection performance of the generated content. By introducing the design of regret loss, the gap between the generated content and the ideal privacy protection goal is quantified, and the regret loss is expressed as: To better utilize this information to train the cost model, the original pairwise comparison loss is corrected by introducing a classification term to construct a loss function for training the cost model:
[0034]
[0035] In the formula, represents the expected value, D C represents the training dataset of the cost model (x, y win , y loss ) and (x, y win , y loss , s win , sloss ) are triples and quintuples sampled from the dataset, where x is the input data and y win and y loss are the answers with more privacy protection and less privacy protection for the input data respectively, and s win and s loss are the scaling factors for the classification losses of the two answers respectively; C ψ (y win , x) is the cost model for the probability function of the input x and the output y with more privacy protection, and C win ψ (y win , x) is the cost model for the probability function of the input x and the output y with less privacy protection; logσ is the logarithmic form of the logistic sigmoid function, which is used to convert the probability to logits; loss is the regret loss introduced above. is the regret loss introduced above.
[0036] Furthermore, in the step S8 described above, the goal of the reinforcement learning training is to maximize the usefulness reward of the generated content while ensuring that the generated response meets the privacy protection standard, that is, for each given input x and the corresponding output y, ensure that all outputs are safe. Then the goal of the reinforcement learning training is as follows:
[0037]
[0038] To deal with this constrained optimization problem, the Lagrange multiplier method is adopted. A non - negative Lagrange multiplier λ is introduced, and the constraint condition is incorporated into the optimization goal, thus transformed into the following unconstrained problem:
[0039]
[0040] In the formula, represents the expected benefit objective function, represents the expected cost objective function; λ represents the degree of penalty for violating the security constraint. If the response generated by the model is safer, then λ will decrease, otherwise it will increase, thereby dynamically adjusting the penalty intensity for privacy protection.
[0041] The above - mentioned method for privacy protection of the generated content of the large - language model based on RLHF of the present invention can adaptively and dynamically identify and avoid potential sensitive information, thereby realizing efficient and accurate privacy protection, and at the same time effectively avoiding the common performance degradation problem in traditional privacy protection methods, and has the following beneficial effects:
[0042] 1) By decoupling the utility model and the cost model, the privacy protection goal and the quality goal of the generated content are modeled separately, and the requirements of these two aspects are deeply integrated in model training, significantly reducing the conflict between privacy protection and model performance.
[0043] 2) Design a multi-dimensional privacy evaluation mechanism and a gating network based on a mixture of experts model, which can accurately evaluate the generated content from multiple dimensions such as compliance, sensitivity, and correctness. The model can dynamically activate the appropriate expert network according to the characteristics of the input content and achieve a balance between privacy protection and the quality of the generated content.
[0044] 3) Model the dynamic optimization of privacy protection strategies based on the human privacy preference of regret theory to achieve a balance between privacy protection and generation quality. Through the regret value feedback mechanism, the model can adaptively adjust its generation behavior, improve the efficiency of privacy protection, and reduce long-term privacy risks and training costs. In addition, this mechanism enhances the multi-dimensional optimization ability of the model, supports the collaborative improvement of goals such as harmlessness, helpfulness, and privacy, and improves user trust and practical application effects. Brief Description of the Drawings
[0045] Figure 1 It is a training flowchart of a method for protecting the privacy of the generated content of a large language model based on human feedback reinforcement learning according to an embodiment of the present invention;
[0046] Figure 2 It is a structural diagram of the cost model in an embodiment of the present invention. Detailed Embodiments
[0047] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0048] As Figure 1 shown, the present invention provides a method for protecting the privacy of the generated content of a large language model based on human feedback reinforcement learning. By decoupling the benefit model and the cost model, the understanding of privacy protection is improved; a multi-dimensional scoring head and a gating network are designed to dynamically analyze the weights of the input content in each dimension of privacy protection; combined with regret theory, the traditional preference probability calculation is optimized, and the benefit model and the cost model are better trained, improving the accuracy of privacy protection calculation; finally, the instruction fine-tuning model is trained through reinforcement learning, and the dual guarantee of privacy protection and the quality of the generated content is successfully achieved; specifically, it includes the following steps:
[0049] S1. Fine-tune the base model based on supervised instruction: Based on the instruction fine-tuning dataset labeled for specific downstream generative tasks, train an open-source base model as the instruction fine-tuning model. The open-source base models include the GPT-2 model, Llama model, Baichuan model, InternLM model, etc.; for instruction fine-tuning, use the Lora instruction fine-tuning method provided by the LLAMA-Factory framework.
[0050] S2. Decouple usefulness and harmlessness: Decompose the original reward model into a benefit model and a cost model. Specifically, remove the last embedding layer of the pre-trained language model based on the Transformer structure (such as GPT-2, Llama, etc.), and retain its original Transformer architecture and weights to construct a decoupled model. On this basis, use two models (the instruction fine-tuning trained base model obtained in step S1, one model to construct the benefit model and one to construct the cost model) to add an independent linear mapping layer at the output end of the last Transformer layer respectively as the benefit model and the cost model. The linear mapping layer is used to receive the hidden state output by the last Transformer layer as input and project it into a scalar value to evaluate the usefulness and privacy of the input text respectively. During the training process, the benefit model is supervised and trained using usefulness annotation data (such as scores for dimensions such as correctness and coherence), while the cost model is separately optimized using privacy-related annotation data (such as whether it contains sensitive information). Through this design, the benefit model and the cost model can independently evaluate the performance of the input text in different dimensions, thus achieving an effective decoupling of usefulness and harmlessness. This decoupling method not only ensures the balance between the usefulness and privacy of the generated content, but also provides greater flexibility for subsequent customized optimization.
[0051] S3. Preference Understanding of the Extended Cost Model: The linear mapping layer of the extended cost model is used to endow the cost model with a multi-dimensional refined scoring head structure, and a gating network is added to dynamically analyze the weight of each dimension for privacy protection in the input content. Specifically, the single linear mapping layer in the cost model obtained in step S2 is extended to five independent linear mapping layers, which respectively score compliance, sensitivity, correctness, objectivity, and credibility. Each linear mapping layer extracts features from the hidden state of the last layer of the Transformer and generates a scoring value for the corresponding dimension, so as to realize the multi-dimensional refined evaluation of privacy protection. In addition, in the scoring head structure of the cost model obtained in step S2, a gating network based on the design idea of the mixture of experts model is introduced. The gating network dynamically analyzes the input content and generates weight coefficients for the above five dimensions. These weight coefficients are used to measure the importance of different dimensions for privacy protection under specific input content, and then dynamically adjust the overall scoring method of the cost model. By combining the multi-dimensional scoring head with the gating network, the cost model can not only evaluate the privacy risk of the generated content in a fine-grained manner, but also flexibly allocate dimension weights according to the characteristics of the input content, so as to more comprehensively and accurately model and understand the privacy protection preferences.
[0052] S4. Training the Multi-Dimensional Scoring Cost Model: Train the multi-dimensional scoring cost model based on the labeled dataset containing multi-dimensional scores and multi-dimensional weights. During the training process, first calculate the mean square error between each dimension score output by the multi-dimensional scoring cost model and the labeled score, and the mean square error between each dimension weight generated by the multi-dimensional scoring cost model and the labeled weight, and use these as the main optimization objectives to obtain the following loss function with a regularization term:
[0053]
[0054] In the formula, represents the scoring error loss of the i-th dimension, represents the weight error loss of the i-th dimension; d is the number of dimensions of the score, and λ is the regularization coefficient, which controls the influence degree of the regularization term on the total loss; The regularization term of the score, The regularization term of the weight.
[0055] Then, the loss function with a regularization term is used to penalize the situation where the variances of the score prediction and weight prediction of the model on a single sample are too large. Through the constraint of this regularization term, the model can avoid generating unreasonable extreme values in a single sample, so as to output a smoother and more stable score and weight distribution. This training method aims to prompt the model to generate a distribution that better meets the actual needs while maintaining high prediction accuracy, and improve the robustness and applicability of the multi-dimensional scoring model.
[0056] S5. Optimize the traditional preference probability calculation formula: Combine the weighted score with the Bradley-Terry preference model to improve the traditional preference probability calculation formula. Specifically, dynamically analyze the input content through the multi-dimensional scoring cost model obtained in step S4 to generate the weights of each dimension, and calculate and obtain the weighted score Score weight . This weighted score Score weight Integrates the scores and weights of multiple dimensions and can more comprehensively reflect the comprehensive performance of the input content in terms of privacy protection. Traditional preference prediction is based on the Bradley-Terry model, that is, given the input x, the model believes that y win The output is better than y loss The probability of the output is used to describe how the model judges which output is more in line with human preferences during the training process. In the present invention, the weighted score Score weight is used to replace the single score S(y|x) in the traditional model, so that the preference calculation can more accurately express the comprehensive performance of the generated content in terms of privacy protection preferences. The new preference probability calculation formula is:
[0057]
[0058] By introducing the weighted score generated by dynamic analysis, the model can more effectively judge and preferentially select answers that meet human needs in both privacy protection and other preference goals, thereby improving the performance and practicality of the model in privacy protection tasks.
[0059] S6. Construct a new human privacy preference probability calculation method: Combine multi-dimensional scores with regret theory to construct a human privacy preference probability method. In this embodiment, regret is defined as the deviation between each segment and the optimal decision, which measures the loss of each segment in terms of expected return. This kind of decision-making is more in line with human privacy protection awareness. Specifically, regret can be decomposed into partial return and state value change. Among them, partial return represents the direct loss of each segment in terms of expected return, that is, the difference between the reward brought by each state-action pair and the reward brought by the optimal strategy, and the state value change represents the change between the start and end state values of each segment, which reflects the degree of movement of each segment in the state space. In the regret-based preference model, a model based on the Boltzmann rational distribution is usually used. The core idea of this model is that the probability that the expert preference σ + is better than σ - is based on the discounted partial return of each trajectory segment: where r E is the score of the reward model on each state-action pair, which is used to calculate the quality of each trajectory, and γ t is the discount factor at time step t, and the sum of the discounted returns represents the total reward value of the trajectory segment over time.
[0060] In reinforcement learning, the state (State, s) usually represents the characteristics of the environment or input content, while the action (Action, a) represents the decision or behavior taken in the current state. In the cost model obtained in step 3, the state s can be represented as the score of each dimension, namely Score i , specifically, the predicted score of the model. For each input σ, the score predicted by the model on the i-th dimension is r i (σ). r i (σ + ) is the score of a better answer on the i-th dimension, and r i (σ - ) is the score of a worse answer on the i-th dimension. These scores constitute the state. The action is represented by the weights output by the gating network. The weight of the gating network on each dimension is γ i , then the gating weights will adjust the scores of each dimension. γ i represents the gating weight of the i-th dimension, and γ i ∈[0,1], so it can be mapped through the Sigmoid function γ i =σ(output i ).
[0061] Advantage represents how good or bad a certain action (i.e., the score after weight adjustment) is relative to a baseline (unadjusted score). The advantage can be calculated by the following formula: The advantage A i (s, a) is the difference between the return after choosing a certain action a in a certain state s and the baseline return (i.e., the average score). Among them, the baseline return is the average score of all samples on the i-th dimension. Then, the advantage of each dimension is weighted by the gating weights In this way, the advantage reflects the influence of each dimension by adjusting the weights.
[0062] For each sample σ + and σ - , calculate its weighted total advantage (i.e., the sum of weighted advantages), and use the discount factor to calculate the future weighted advantage: where T is the sequence length (time step).
[0063] For each pair of samples in each batch, use the Boltzmann distribution to calculate the preference, that is, the probability of regret, between the "better" and "worse" answers:
[0064]
[0065] Exponentiate the sum of the weighted advantages of each dimension, and then calculate the preference probability between the two weighted answers.
[0066] S7, Training Benefit Model and Cost Model: Use preferences to train the training benefit model and cost model on the dataset. Among them, the benefit model aims to evaluate the usefulness of the generated content, and when training, the model parameters are adjusted by maximizing the preference probability. Specifically, the model determines which of the two candidate outputs is better based on the preference data, and constructs the loss function for training accordingly:
[0067]
[0068] By minimizing the loss function, the benefit model can more accurately predict the quality of the generated content, thereby improving its performance in terms of task completion, etc.
[0069] The goal of the cost model is to evaluate the privacy protection performance of the generated content, and by introducing the design of regret loss, the gap between the generated content and the ideal privacy protection goal is quantified. The regret loss can be expressed as:
[0070]
[0071] To better utilize this information to train the cost model, the original pairwise comparison loss is corrected by introducing a classification term:
[0072]
[0073] In the formula, represents the expected value, D C represents the training dataset (x, y win , y loss ) and (x, y win , y loss , s win , s loss ) sampled from the dataset are triples and quintuples, where x is the input data, y win and y loss are the answers with better privacy protection and less privacy protection for the input data respectively, s win and s loss are the scaling factors for the classification losses of the two answers respectively; C ψ (y win , x) is the probability function of the cost model for the input x and the output answer y win with better privacy protection, C ψ (y win , x) is the probability function of the cost model for the input x and the output answer y loss with less privacy protection; logσ is the logarithmic form of the logistic sigmoid function, used to convert the probability to logits; is the regret loss introduced above.
[0074] The classification loss term is used to explicitly classify the privacy risks of the generated content, and through the joint optimization with the regret loss, improve the model's ability in fine-grained privacy assessment.
[0075] Finally, the comprehensive loss function effectively balances the two optimization objectives, enabling the cost model to conduct more accurate dynamic assessment of privacy protection. Construct the loss function for training the cost model:
[0076]
[0077] S8. Fine-tune the model with reinforcement learning training instructions: Based on the described benefit model and cost model, combined with the new human privacy preference probability calculation method, fine-tune the model with training instructions through proximal policy optimization. The training objective is to maximize the usefulness reward of the generated content while ensuring that the generated responses meet the privacy protection standards, that is, for each given input x and corresponding output y, all outputs do not contain privacy issues. Therefore, the reinforcement learning training objective can be expressed as:
[0078]
[0079] In the formula, denotes sampling an input x from the data distribution D (i.e., the prompt dataset used for RL training); y ∼ π θ (·|x) denotes sampling an answer y from the answer distribution generated by the current LLM policy π θ ; denotes the expectation, and R φ (y, x) denotes the reward value given by the helpfulness reward model, and C ψ (y, x) denotes the privacy-protective cost evaluated by the cost model.
[0080] Since the optimization problem of strictly satisfying the harmlessness constraint usually significantly increases the difficulty of reinforcement learning training, a new hyperparameter d is introduced to rewrite the original hard constraint into an expected form, making the objective function more optimizable. Specifically, the objective function includes the following two parts:
[0081] denotes the expected usefulness reward;
[0082] denotes the expected harmlessness cost objective, and the probability of generating unsafe responses is controlled by the hyperparameter d.
[0083] The rewritten objective function becomes the following form:
[0084] To further address this constrained optimization problem, the Lagrange multiplier method is adopted to transform the harmlessness constraint into a form of constraint penalty. By introducing non - negative Lagrange multipliers, the usefulness objective and the harmlessness constraint are unified into an optimization objective. The transformed objective function:
[0085]
[0086] In the formula, λ represents the degree of penalty for violating the safety constraint. When the model generates a more secure response, λ will gradually decrease; conversely, if the generated response has more privacy issues, λ will increase, thereby dynamically adjusting the penalty for privacy protection.
[0087] 1. Experimental Setup
[0088] The dataset used in the simulation experiment of the present invention is PKU - SafeRLHF in PKU - Alignment. It contains more than 30,000 pairs of expert comparison data. Each data pair contains two answers to a question, as well as safety meta - labels and preferences based on helpfulness and harmlessness. This dataset provides a detailed ranking for the answers through independent helpfulness and harmlessness evaluations, aiming to improve the safety of the language model output through this two - dimensional ranking system. The dataset is used for research purposes, with a particular emphasis on research to reduce the harmfulness of the model, and contains content that may be offensive or harmful.
[0089] 2. Result Analysis
[0090] The simulation experiment of the present invention adopts the present invention and an existing related technology Safe RLHF training framework. The related technology Safe RLHF training framework refers to the training framework of reinforcement learning from human feedback proposed by Dai J et al. in "Safe RLHF: Safe Reinforcement Learning from Human Feedback, 2023", which decouples human preferences for helpfulness and harmlessness, allows training separate reward and cost models, formalizes the safety problem of LLMs as an optimization task of maximizing the reward function while satisfying the specified cost constraint, and dynamically adjusts the balance between the two objectives during the fine - tuning process.
[0091] To verify the effect of the simulation experiment of the present invention, the GPT-2 and GPT-Neo models were used for verification respectively. First, the accuracy of obtaining correct human privacy preferences was evaluated by using the validation set in the benefit model and the cost model; secondly, the target model after completing the entire training framework and the model only fine-tuned by supervised instructions were used to answer the prompts that induce privacy leakage, and the GPT-4 scores were compared. Since the comparison of two models was involved, the Elo score was calculated, which was shown as the increased value of the score of the target model after completing the entire training framework compared with the model only fine-tuned by supervised instructions. All the results were plotted into the following table:
[0092]
[0093] As can be seen from the above table, significant improvements were shown in both the cost model and the GPT-4 score in the present invention. Specifically, for the GPT-2 model, the effect of adopting the present invention was better than the traditional Safe-RLHF method, where the accuracy of the cost model increased by 1.57%, and the GPT-4 score increased by 14.78 points; for the GPT-Neo model, the accuracies of the benefit model and the cost model increased by 0.74% and 0.97% respectively, and the GPT-4 score increased by 26.53 points. As the number of model parameters increased, the effect of the present invention became more significant, especially the improvement in the GPT-4 score was particularly prominent. The present invention combines a multi-dimensional scoring system and the dynamic weight analysis of a gated network, enabling the model to more comprehensively understand and respond to human preferences, especially in the field of privacy protection, showing stronger adaptability and accuracy. By improving the modeling of the probability of human privacy preferences, the ability of the model to protect privacy data can be more effectively enhanced in reinforcement learning.
[0094] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for protecting the privacy of content generated by large language models based on RLHF, characterized in that: It includes the following steps: S1. Fine-tune the base model based on supervised instructions: Train an open-source base model as an instruction fine-tuning model based on an instruction fine-tuning dataset annotated for specific downstream generative tasks. S2. Decouple usefulness and privacy: Decompose the original reward model into a benefit model and a cost model. S3. Expand the preference understanding of the cost model: Expand the linear layer of the cost model so that the cost model has a multi-dimensional refined scoring head structure, and add a gating network to dynamically analyze the weight of each dimension for privacy protection for the input content. S4. Train a multi-dimensional scoring cost model: Train a multi-dimensional scoring cost model based on an annotated dataset containing multi-dimensional scores and multi-dimensional weights. S5. Optimize the traditional preference probability calculation formula: Improve the traditional preference probability calculation formula based on the weighted score and the Bradley-Terry preference model. The weighted score is the weighted score Score after obtaining the dynamic analysis weight of the input content based on the multi-dimensional scoring cost model obtained in step S4. weight ; S6. Construct a new method for calculating the probability of human privacy preference: Construct a new method for calculating the probability of human privacy preference by combining multi-dimensional scoring with regret theory. In step S6, regret is defined as the deviation between each trajectory segment and the optimal decision, which is used to measure the loss in its expected return; regret is decomposed into partial return and state value change. The former is the difference between the reward corresponding to each state-action and the optimal reward, and the latter is the value difference between the start and end states of the trajectory segment. The state represents the scores of the multi-dimensional scoring cost model in each dimension, and the action is represented by the weights output by the gating network, which is used to adjust the scores of each dimension. S7. Train the benefit model and the cost model: Use preferences to train the benefit model and the cost model on the dataset; the benefit model is used to evaluate the usefulness of the generated content; the cost model is used to evaluate the privacy protection performance of the generated content, and by introducing the design of regret loss, the gap between the generated content and the ideal privacy protection goal is quantified. S8. Train the instruction fine-tuning model through reinforcement learning: Based on the benefit model and the cost model and the new method for calculating the probability of human privacy preference, train the instruction fine-tuning model through proximal policy optimization.
2. The method for protecting the privacy of the content generated by the large language model based on RLHF according to claim 1, characterized in that: In step S2, remove the last un-embedded layer of the large language model pre-trained based on Transformer, and add an additional linear layer in the last Transformer layer to construct the benefit model and the cost model; given any text, the benefit model and the cost model will assign a scalar reward value to the last token, which are used as indicators to evaluate the usefulness and privacy of the answer respectively.
3. The method for protecting the privacy of the content generated by the large language model based on RLHF according to claim 1, characterized in that: In step S3, expand an additional linear layer of the cost model obtained in step S2 into five separate linear layers to perform fine-grained scoring on compliance, sensitivity, correctness, objectivity, and credibility respectively. And add a gating network designed based on the idea of a mixture of experts at an additional linear layer of the cost model obtained in step S2 to dynamically analyze the impact of each dimension on privacy protection for the input content and obtain different weights.
4. A method for protecting the privacy of content generated by a large language model based on RLHF according to claim 1, characterized in that: In the step S4 described above, first calculate the mean square error between each dimension score output by the multi-dimensional scoring cost model and the labeled score, as well as the mean square error between each dimension weight generated by the multi-dimensional scoring cost model and the labeled weight, and use this as the main optimization objective to obtain the following loss function with a regularization term In the formula, represents the scoring error loss of the i-th dimension, represents the weight error loss of the i-th dimension; d is the number of dimensions of the scoring, and λ is the regularization coefficient, which controls the influence degree of the regularization term on the total loss; is the regularization term of the scoring, is the regularization term of the weight; Then, use the loss function to penalize the situation where the variance of the score prediction and weight prediction of the model on a single sample is too large, and avoid generating unreasonable extreme values in a single sample.
5. A method for protecting the privacy of content generated by a large language model based on RLHF, characterized in that: In the said step S5, the weighted score Score weight is used to replace the score of a single evaluation, and the new preference probability calculation formula is as follows: where P(y win > y loss |x) represents the probability that the preference model judges y win to be better than y loss ; Score weight (y win |x) is the numerical value of the weighted score for the input x and the output y win ; Score weight (y loss |x) is the numerical value of the weighted score for the input x and the output y loss ; exp is used to convert the score into a probability, forming a softmax function to ensure that the sum of probabilities is 1.
6. A method for protecting the privacy of content generated by a large language model based on RLHF, characterized in that: In the step S6, in the multi-dimensional scoring cost model obtained in the step S4, first calculate how good or bad the score adjusted by a certain action weight is relative to the benchmark, that is, the negative regret, through the following formula: where r i (σ + ) is the score of a better answer in the i-th dimension, and the benchmark return is the average score of all samples in the i-th dimension; Then, the advantages of each dimension are weighted by the gating weights where γ i is the weight of the better answer in the i-th dimension; combined with the discount factor ω, the future weighted total advantage is calculated: where t is the time step and T is the sequence length; Then, for each pair of samples in each batch, the Boltzmann distribution is used to calculate the probability of preference, i.e., regret, between the "better" and "worse" answers, and the sum of the weighted advantages for each dimension is exponentiated: where P(σ + >σ - ) represents the probability that for a given sample pair, the model judges that the better answer is superior to the worse answer; represents the cumulative weighted advantage value of the better answer, represents the cumulative weighted advantage value of the worse answer; exp is used to convert the score into a probability, forming a softmax function to ensure that the sum of probabilities is 1; Finally, the preference probability between the two weighted answers is calculated.
7. A method for protecting the privacy of content generated by a large language model based on RLHF as claimed in claim 1, characterized in that: In the said step S7, by maximizing the probability of y win being better than y loss , the parameter φ of the benefit model is adjusted to construct the loss function for training the benefit model: In the formula, represents the expected value, D R represents the training data set of the benefit model, (x, y win , y loss ) is a triple sampled from the data set, where x is the input data, and y win and y loss are better and worse answers to the input data respectively; R φ (y win , x) is the probability function of the benefit model for the input x and the better answer y win , and R φ (y loss , x) is the probability function of the benefit model for the input x and the worse answer y loss ; logσ is the logarithmic form of the logistic sigmoid function, which is used to convert the probability to logits.
8. The method for protecting the privacy of content generated by a large language model based on RLHF according to claim 6, wherein: In the step S7 described above, the regret loss is expressed as: where B represents the total number of samples; by introducing a classification term to correct the original pairwise comparison loss, a loss function for cost model training is constructed: In the formula, represents the expected value, and D C represents the training data set (x, y win , y loss ), and (x, y win , y loss , s win , s loss ) are triples and quintuples sampled from the data set, where x is the input data, and y win and y loss are the more privacy-protected and less privacy-protected answers to the input data respectively, and s win and s loss are the scaling factors for the classification losses of the two answers respectively; C ψ (y win , x) is the probability function of the cost model for the input x and the more privacy-protected answer y win , and C ψ (y win , x) is the probability function of the cost model for the input x and the less privacy-protected answer y loss ; logσ is the logarithmic form of the logistic sigmoid function, which is used to convert the probability to logits; is the regret loss introduced above.
9. The method for protecting the privacy of content generated by a large language model based on RLHF according to claim 1, wherein: In the step S8, the goal of the reinforcement learning training is to maximize the usefulness reward of the generated content while ensuring that the generated response complies with the privacy protection standard; the Lagrange multiplier method is adopted, a non-negative Lagrange multiplier λ is introduced, and the constraint conditions are incorporated into the reinforcement learning training, thus transforming it into the following unconstrained problem: In the formula, represents the expected benefit objective function, represents the expected cost objective function; λ represents the degree of penalty for violating safety constraints. If the response generated by the model is safer, then λ will decrease, and vice versa.
Citation Information
Patent Citations
Method for evaluating participation degree of gated hybrid expert network based on space-time Transform
CN118918510A
Large model security alignment method based on dynamic constraint reinforcement learning
CN119539057A