Noise adding method and system for large language model fine tuning scene and storage medium
By screening the cryogenic layer and fine-tuning layer according to the contribution and sensitivity of each layer in the large language model, and adding differentiated noise to the fine-tuning layer, the balance between model availability and privacy protection in the fine-tuning scenario of the large language model is solved, reducing the computing overhead and improving the model accuracy.
Patent Information
- Application Number
- CN202510499486.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to achieve a balance between model availability, efficient model training and privacy protection in large language model fine-tuning scenarios, especially when using differential privacy to protect data privacy, the calculation overhead is high and the model accuracy is low.
By computing the importance score for each layer based on the contribution and sensitivity of each layer to accuracy according to the large language model, and filtering it into a frozen layer or a fine-tuning layer, differential noise is added to different fine-tuning layers based on the importance score.
It realizes the improvement of model availability while reducing computing overhead, and effectively protecting data privacy, achieving a balance between model accuracy and privacy protection.
Smart Images

Figure CN120354902A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and particularly relates to a noise addition method, system and storage medium for the fine-tuning scenario of large language models. Background Art
[0002] Large Language Models (LLMs) have been widely used in the fields of natural language processing, computer vision, etc. Such models usually contain hundreds of billions of parameters and can fully capture the complex features between input data. Classification tasks are one of the typical applications of large language models, such as sentiment analysis, image recognition and other tasks. Due to the large number of parameters in LLMs, the training cost is relatively high. To reduce the training cost, the industry has proposed a pre-training - fine-tuning framework, that is, a large amount of public data sets are used for feature extraction during the pre-training process, and a small amount of private data sets are used to further train the model during fine-tuning, so as to improve the efficiency while ensuring the accuracy of the model on specific tasks.
[0003] However, the private data used during the fine-tuning process may contain sensitive information, and there is a risk of private data leakage due to the membership inference attack (MIA) of the model. The membership inference attack means that the attacker can use the model parameters to judge whether a piece of data is in the training set, which will lead to serious privacy leakage problems. For example, a certain model is used for diabetes diagnosis, and the attacker can judge whether a user has diabetes based on whether the user's data is in the training set.
[0004] Combining differential privacy (DP) with large language model training to protect the privacy of individual data is one of the technical directions for addressing the above issues. Specifically, differential privacy technology ensures that the model does not remember or disclose the specific information of individual data by adding noise during the training process, thereby enabling the training of an effective model while protecting user privacy. However, in the context of large language models, using differential privacy to protect data privacy faces two technical challenges: on the one hand, due to the large number of parameters in large language models, if noise is added to all parameters, it will result in significant time and computational overhead; on the other hand, this technology faces the problem that model usability and privacy protection cannot be achieved simultaneously. Specifically, if more noise is added during the training process, stronger privacy protection can be achieved, but excessive noise will significantly reduce metrics such as the accuracy of the model; if less noise is added, the model has a higher accuracy, but the effect of privacy protection is poor. Therefore, achieving a balance between model usability and privacy protection and minimizing the computational overhead during training is an issue of widespread concern in the industry. Existing solutions can be divided into two categories:
[0005] 1) The first type of method optimizes all parameters in the model. During the training process, the importance of each layer is quantified based on the uncertainty and value range of different layers, and thus differential noise is added to the gradients of different layers of the model according to the importance scores, that is where S is the importance score, C is the sensitivity of the model gradient, that is, the difference between the upper and lower limits of the model gradient values, σ is the noise variance, I represents the identity matrix, represents a Gaussian distribution with mean a and variance b. That is, the more important the layer, the less noise is added; the less important the layer, the more noise is added to it.
[0006] 2) The second type of method only adds noise to specific layers of the model. This type of solution believes that the bottom layers of the model are used for feature extraction, while the upper layers can capture global context and abstract semantics. Therefore, only the last few layers are divided into fine-tuning layers, and the other layers are frozen layers. The frozen layers are the layers that do not participate in training, and the fine-tuning layers participate in model updates during the fine-tuning process, and noise is added to the fine-tuning layers during the model parameter update process It should be noted that in this method, the noise distributions added to different fine-tuning layers are the same.
[0007] Although existing solutions have achieved data privacy protection to a certain extent during the fine-tuning process, they have not simultaneously achieved fine-tuning layer screening and fine-grained noise allocation, making it difficult to effectively achieve a balance among model usability, efficient model training, and privacy protection.
[0008] The first type of method adds differential noise to different layers based on the importance of the layers, achieving high model accuracy and privacy protection. However, this type of method adds noise to all layers, resulting in a large computational overhead during the training process due to the noise addition operation. In addition, when quantifying the importance, this type of solution only calculates based on the sensitivity and uncertainty of the model parameters. This quantization method is relatively coarse-grained and is not sufficient to fully characterize the importance of different layers, resulting in inappropriate added noise and affecting the model performance.
[0009] The second type of method achieves efficient training because this type of method only adds noise to specific layers. However, in this type of method, the noise distributions added to all fine-tuning layers are the same and relatively coarse-grained, resulting in insufficient effectiveness in protecting data privacy while achieving high performance. Summary of the Invention
[0010] The object of the present invention is to propose a noise addition method, system, and storage medium for the fine-tuning scenario of large language models to achieve a balance among model usability, efficient model training, and privacy protection, and to solve the imbalance among model usability, efficient model training, and privacy protection in the prior art.
[0011] The present invention provides a noise addition method for the fine-tuning scenario of large language models, including the following steps:
[0012] Obtain the importance score of each layer in the large language model according to the contribution degree of each layer to the accuracy of the large language model and the sensitivity of each layer;
[0013] Screen the layers of the large language model according to the importance scores of different layers, use the layers with importance scores greater than or equal to the preset importance threshold as frozen layers, and use the layers with importance scores less than the preset importance threshold as fine-tuning layers;
[0014] Add differential noise to different fine-tuning layers based on the importance scores.
[0015] Preferably, a probing model is used to obtain the contribution degree of each layer to the accuracy of the large language model. The input of the probing model is the output of the neurons in each layer of the large language model. The classification task of the probing model is the same as the classification task of the large language model. The contribution degree of each layer to the model accuracy is obtained according to the accuracy of the probing model. The higher the accuracy of the probing model, the greater the contribution degree of the input layer.
[0016] Preferably, labels are set for the data samples according to the data characteristics of the neurons. The large language model is trained using the flipped labels of each layer. The neurons in each layer are perturbed using the model poisoning attack on the data samples, and the perturbed neurons are fixed. The resilience of the model is tested using the dataset with unflipped labels, and the sensitivity of each layer is evaluated.
[0017] Preferably, the evaluation of the sensitivity of each layer includes two stages: the model poisoning attack on the fixed layer and the difference evaluation. The model poisoning attack on the fixed layer is used to perturb the neurons of the fixed layer, and the difference evaluation is used to evaluate the corresponding change rate of the model accuracy after perturbing each layer.
[0018] Preferably, the model poisoning attack on the fixed layer specifically includes the following steps:
[0019] Set the parameters of the large language model of a selected layer to be updatable in sequence, and set the model parameters of other layers to be non-updatable;
[0020] Use the dataset with flipped labels to train the neurons of the selected i-th layer, and perturb the neurons of the i-th layer using the samples of the poisoning attack;
[0021] Then use the dataset without flipped labels for testing to obtain the first accuracy ACC of the large language model 1i .
[0022] Preferably, the difference evaluation includes the following steps:
[0023] In the large language model, fix the neurons of the i-th layer perturbed by any dataset with flipped labels, and set the other neurons except this layer to be trainable;
[0024] Use the dataset without flipped labels to train the large language model;
[0025] Use the dataset without flipped labels to test this model to obtain the second accuracy ACC of the large language model 2i ;
[0026] Evaluate the sensitivity of each layer in the large language model according to the difference between the first accuracy and the second accuracy.
[0027] Preferably, for the i-th layer of the large language model, evaluating the sensitivity of each layer in the large language model according to the difference between the first accuracy and the second accuracy specifically includes:
[0028] Calculate the difference p between the first accuracy and the second accuracy of the i-th layer of the large language model 2i =|ACC 2i -ACC 1i |; The value range of p 2i is [0,1], representing the anti-perturbation ability of this layer, and p 2i is negatively correlated with the importance of this layer;
[0029] Perform the following processing on p 2i to obtain
[0030] p 2i ' = 1 - p2i
[0031] Among them, p 2i ′ is an intermediate variable, and the value range of p 2i ′ is [0, 1], and the change trend of p 2i ′ is positively correlated with the importance of this layer;
[0032] Process p 2i ′ to obtain the sensitivity degree s 2i of the i-th layer:
[0033]
[0034] Among them, p 21 ′, p 22 ′,..., p 2n ′ represent the intermediate variables of the 1st, 2nd,..., nth layers of the large language model, and n is the total number of layers of the large language model.
[0035] Preferably, the contribution degree of each layer to the accuracy of the large language model and the sensitivity degree of each layer are weighted and averaged to obtain the importance score of each layer in the large language model.
[0036] The present invention also provides a noise addition system for the fine-tuning scenario of a large language model, including a processor, and the processor can execute the steps of the above-mentioned noise addition method for the fine-tuning scenario of the large language model.
[0037] The present invention also provides a storage medium, on which a program is stored, and the program can execute the steps of the above-mentioned noise addition method for the fine-tuning scenario of the large language model.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] (1) The present invention proposes a novel layer importance quantification strategy that comprehensively considers the contribution of different layers to accuracy and the sensitivity of different layers. Specifically, regarding the contribution of different layers to accuracy, the present invention applies a probing model strategy, and after processing the accuracy of the probing model, it is used as the contribution value of that layer; regarding the anti-perturbation of different layers, the present invention proposes a model parameter sensitivity quantification scheme based on model poisoning attacks, perturbs the neurons of different layers based on model poisoning attacks, fixes the perturbed neurons, and uses the dataset with unflipped labels to test the resilience of the model, thereby quantifying the sensitivity of different layers. Finally, the present invention combines these two factors to obtain the final model importance score. The present invention improves model usability while achieving privacy protection. Based on the fine-grained noise addition mechanism of layer importance scores, a small amount of noise is added to important layers (fine-tuning layers) to reduce the impact of adding noise on model usability; and more noise is added to unimportant parameters to achieve the purpose of privacy protection.
[0040] (2) Based on the importance scores of different layers, the present invention designs an efficient noise addition strategy. Specifically, the present invention screens the frozen layers and fine-tuning layers according to the importance scores of different layers, and based on the importance scores, adds differentiated noise to different fine-tuning layers instead of adding noise to all layers. Compared with the traditional scheme, the computational overhead caused by adding noise is greatly reduced. Brief Description of the Drawings
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0042] Figure 1 It is the overall framework of the noise addition method for the fine-tuning scenario of large language models according to an embodiment of the present invention.
[0043] Figure 2 It is a schematic diagram of the quantification strategy for the contribution of different layers to the accuracy of large language models according to an embodiment of the present invention.
[0044] Figure 3 It is a schematic diagram of the quantification strategy for the contribution of each layer in the large language model to the accuracy of the large language model according to an embodiment of the present invention.
[0045] Figure 4 It is a flowchart of the noise addition method for the fine-tuning scenario of large language models according to an embodiment of the present invention. Detailed Embodiments
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0047] The present invention provides a noise addition method for the fine-tuning scenario of large language models, including the following steps:
[0048] According to the contribution degree of each layer in the large language model to the accuracy of the large language model and the sensitivity of each layer, obtain the importance score of each layer in the large language model;
[0049] According to the importance scores of different layers, screen the layers of the large language model, use the layers with importance scores greater than or equal to the preset importance threshold as the frozen layers, and use the layers with importance scores less than the preset importance threshold as the fine-tuning layers;
[0050] Add differentiated noise to different fine-tuning layers based on the importance scores.
[0051] Further, use a probing model to obtain the contribution degree of each layer to the accuracy of the large language model. The input of the probing model is the output of the neurons in each layer of the large language model. The classification task of the probing model is the same as the classification task of the large language model. According to the accuracy of the probing model, obtain the contribution degree of each layer to the model accuracy. The higher the accuracy of the probing model, the greater the contribution degree of the input layer.
[0052] Further, set labels for data samples according to the data characteristics of neurons, train the large language model using the flipped labels of each layer, perturb the neurons in each layer using model poisoning attacks on data samples, and fix the perturbed neurons. Use the dataset with unflipped labels to test the resilience of the model and evaluate the sensitivity of each layer.
[0053] Furthermore, the evaluation of the sensitivity of each layer includes two stages: model poisoning attack on the fixed layer and differential evaluation. The model poisoning attack on the fixed layer is used to perturb the neurons in the fixed layer, and the differential evaluation is used to evaluate the corresponding change rate of the model accuracy after perturbing each layer.
[0054] Furthermore, the model poisoning attack on the fixed layer specifically includes the following steps:
[0055] Sequentially set the parameters of the large language model of a selected layer to be updatable, and set the model parameters of other layers to be non-updatable;
[0056] Train the selected neurons in the $i$-th layer using the dataset with flipped labels, and perturb the neurons in the $i$-th layer using the samples of the poisoning attack;
[0057] Then use the dataset with unflipped labels for testing to obtain the first accuracy ACC of the large language model 1i .
[0058] Furthermore, the difference evaluation includes the following steps:
[0059] In the large language model, fix the neurons in the $i$-th layer perturbed by any dataset with flipped labels, and set the neurons other than this layer to be trainable;
[0060] Use the dataset with unflipped labels to train the large language model;
[0061] Use the dataset with unflipped labels to test this model to obtain the second accuracy ACC of the large language model 2i ;
[0062] Evaluate the sensitivity of each layer in the large language model according to the difference between the first accuracy and the second accuracy.
[0063] Furthermore, for the $i$-th layer of the large language model, evaluating the sensitivity of each layer in the large language model according to the difference between the first accuracy and the second accuracy specifically includes:
[0064] Calculate the difference $p$ between the first accuracy and the second accuracy of the $i$-th layer of the large language model 2i $= |ACC$ 2i $- ACC$ 1i $|$; The value range of $p$ 2i is $[0, 1]$, representing the anti-perturbation ability of this layer, and $p$ 2i is negatively correlated with the importance of this layer;
[0065] Perform the following processing on $p$ 2i to obtain
[0066] $p$ 2i $' = 1 - p$ 2i
[0067] where $p$ 2i $'$ is an intermediate variable, the value range of $p$ 2i $'$ is $[0, 1]$, and the change trend of $p$ 2i $'$ is positively correlated with the importance of this layer;
[0068] Perform processing on $p$ 2i $'$ to obtain the sensitivity $s$ of the $i$-th layer 2i :
[0069]
[0070] Among them, p 21 ′, p 22 ′,..., p 2n ′ represent the intermediate variables of the 1st, 2nd,..., nth layers of the large language model, where n is the total number of layers of the large language model.
[0071] Furthermore, the contribution degree of each layer to the accuracy of the large language model and the sensitivity degree of each layer are weighted and averaged to obtain the importance score of each layer in the large language model.
[0072] The present invention also provides a noise addition system for the fine-tuning scenario of a large language model, including a processor, and the processor can execute the steps of the above-mentioned noise addition method for the fine-tuning scenario of the large language model.
[0073] The present invention also provides a storage medium, on which a program is stored, and the program can execute the steps of the above-mentioned noise addition method for the fine-tuning scenario of the large language model.
[0074] Embodiment 1
[0075] As Figure 4 shown, the present invention provides a noise addition method for the fine-tuning scenario of a large language model, including the following steps:
[0076] According to the contribution degree of each layer in the large language model to the accuracy of the large language model and the sensitivity degree of each layer, obtain the importance score of each layer in the large language model;
[0077] According to the importance scores of different layers, screen the layers of the large language model, and use the layers with importance scores greater than or equal to the preset importance threshold as the frozen layers, and use the layers with importance scores less than the preset importance threshold as the fine-tuning layers;
[0078] Add differential noise to different fine-tuning layers based on the importance scores.
[0079] Embodiment 2
[0080] As Figures 1-4 shown, the present invention provides a noise addition method for the fine-tuning scenario of a large language model, including the following steps:
[0081] According to the contribution degree of each layer in the large language model to the accuracy of the large language model and the sensitivity degree of each layer, weight and average the contribution degree of each layer to the accuracy of the large language model and the sensitivity degree of each layer to obtain the importance score of each layer in the large language model;
[0082] Among them, a detection model is used to obtain the contribution degree of different layers to the accuracy of the large language model. The input of the detection model is the output of neurons in each layer of the large language model. The classification task of the detection model is the same as that of the large language model. The contribution degree of each layer to the model accuracy is obtained according to the accuracy of the detection model. The higher the accuracy of the detection model, the greater the contribution degree of the input layer.
[0083] Among them, labels are set for data samples according to the data characteristics of neurons. The large language model is trained using the flipped labels of each layer. The neurons in each layer are perturbed using model poisoning attacks on data samples, and the perturbed neurons are fixed. The resilience of the model is tested using a dataset with unflipped labels, and the sensitivity of each layer is evaluated.
[0084] Specifically, the evaluation of the sensitivity of each layer includes two stages: model poisoning attack on the fixed layer and differential evaluation. The model poisoning attack on the fixed layer is used to perturb the neurons in the fixed layer, and the differential evaluation is used to evaluate the corresponding change rate of the model accuracy after perturbing each layer.
[0085] Among them, the model poisoning attack on the fixed layer specifically includes the following steps:
[0086] Successively set the parameters of the large language model of a selected layer to be updatable, and the model parameters of other layers to be non-updatable;
[0087] Use the dataset with flipped labels to train the neurons in the selected i-th layer, and perturb the neurons in the i-th layer using the samples of the poisoning attack;
[0088] Then use the dataset with unflipped labels for testing to obtain the first accuracy ACC of the large language model 1i .
[0089] Among them, the differential evaluation includes the following steps:
[0090] In the large language model, fix the neurons in the i-th layer perturbed by any dataset with flipped labels, and set the other neurons except this layer to be trainable;
[0091] Use the dataset with unflipped labels to train the large language model;
[0092] Use the dataset with unflipped labels to test this model to obtain the second accuracy ACC of the large language model 2i ;
[0093] Evaluate the sensitivity of each layer in the large language model according to the difference between the first accuracy and the second accuracy.
[0094] Among them, for the i-th layer of the large language model, evaluating the sensitivity of each layer in the large language model according to the difference between the first accuracy rate and the second accuracy rate specifically includes:
[0095] Calculate the difference p between the first accuracy rate and the second accuracy rate of the i-th layer of the large language model 2i =|ACC 2i -ACC 1i |; p 2i ranges from [0, 1], indicating the anti-disturbance ability of this layer, and p 2i is negatively correlated with the importance of this layer;
[0096] Perform the following processing on p 2i to obtain
[0097] p 2i ′=1 - p 2i
[0098] where p 2i ′ is an intermediate variable, p 2i ′ ranges from [0, 1], and the change trend of p 2i ′ is positively correlated with the importance of this layer;
[0099] Process p 2i ′ to obtain the sensitivity degree s of the i-th layer 2i :
[0100]
[0101] where p 21 ′, p 22 ′,..., p 2n ′ represent the intermediate variables of the 1st, 2nd... nth layers of the large language model, and n is the total number of layers of the large language model.
[0102] According to the importance scores of different layers, screen the layers of the large language model, and use the layers with importance scores greater than or equal to the preset importance threshold as the frozen layers, and use the layers with importance scores less than the preset importance threshold as the fine-tuning layers;
[0103] Add differentiated noise to the fine-tuning layers with different importance scores.
[0104] Example 3
[0105] As Figures 1-4 shown, the present invention provides a noise addition method for the fine-tuning scenario of a large language model, including the following steps:
[0106] Step S1, importance score calculation. When calculating the importance score, the present invention comprehensively considers the contribution degree of each layer in the large language model to the accuracy rate of the large language model and the sensitivity degree of each layer.
[0107] Specifically, regarding the contribution of different layers to the accuracy of the large language model, the present invention adopts the idea of a ProbeModel. The ProbeModel is a classifier whose input is the output of each layer of neurons in the LLM model, and its classification task is the same as the LLM task. In the present invention, cat and dog classification is used. The more important a layer is, the higher the accuracy of the ProbeModel of that layer. Thus, the contribution of that layer to the accuracy of the large language model is quantified. The present invention denotes the contribution of the i-th layer to the accuracy of the large language model as s 1i . The output of neurons contains information about the input features. The greater the contribution of a certain layer to the accuracy of the LLM model
[0108] After certain processing of the results of the ProbeModel, it is used as the quantification of the contribution degree of the accuracy of the large language model. The greater the contribution of the neurons of a certain layer to the accuracy of the LLM model, the higher the accuracy of predicting the classification result only through the output of the neurons of that layer. Therefore, the present invention adopts the strategy of the ProbeModel, such as Figure 2 shown. The ProbeModel strategy means taking the output of each layer of the target model as the input to construct a lightweight classifier to judge the contribution of the output of each layer of the model to the overall model. The same as the traditional Probe scheme, taking the data feature D as the input of the large language model, the outputs o1, o2,..., o of each layer of the neural network of the large language model are obtained n , taking o1, o2,..., o n as the input of the ProbeModel, and thus conducting the cat and dog recognition classification task test
[0109] The ProbeModel includes two layers of fully connected neural networks. And use the test set to evaluate the ProbeModel, that is, compare the output of the ProbeModel with the true label y of the data, so as to obtain the accuracy rates of each layer as p 11 , p 12 ,..., p 1n . Different from the traditional Probe scheme, after obtaining the accuracy rates of different layers, in order to make the gap between the accuracy rate contributions of different layers larger, the present invention needs to perform certain processing on p 11 , p 12 ,..., p 1n to obtain s 11 , s 12 ,..., s 1n . Taking p 1i as an example, perform processing on p 1i to obtain s 1i , as the quantification of the input feature information, that is
[0110]
[0111] Among them, min() represents the minimum value function, and max() represents the maximum value function.
[0112] Regarding the sensitivity of different layers, the present invention designs a quantization strategy for the anti-perturbation ability and sensitivity based on the model poisoning attack (Data Poisoning Attack). The present invention perturbs the neurons of different layers based on the model poisoning attack, fixes the perturbed neurons, and uses the dataset with unflipped labels to test the resilience of the model, thereby quantifying the anti-perturbation ability of different layers, and then quantifying the sensitivity. The stronger the anti-perturbation ability of this layer, the less important this layer is, because even if this layer is perturbed, other layers can still make correct predictions; the weaker the anti-perturbation ability of this layer, the more important this layer is, because when this layer is perturbed, it is difficult for other layers to make correct predictions, as Figure 3 shown. The present invention records the sensitivity of the neurons in the i-th layer as s 2i , and then combines s 1i and s 2i through weighted combination to obtain the importance score of the i-th layer, denoted as S i . If the model recovery ability is poor after perturbation, it indicates that the neurons in this layer have a more significant impact on the accuracy of the large language model. The anti-perturbation ability of the neurons in this layer is poor while the sensitivity is stronger, that is, the neurons in this layer are more important.
[0113] The evaluation of the sensitivity of each layer includes two stages: the model poisoning attack on the fixed layer and the difference evaluation. The model poisoning attack on the fixed layer is used to perturb the neurons of the fixed layer, and the difference evaluation is used to evaluate the corresponding change rate of the model accuracy after perturbing each layer. Specifically as follows:
[0114] Perform a model poisoning attack on the neurons of different layers. In the traditional model poisoning attack scheme, first, the labels of the dataset are flipped. For example, the label of a cat picture is changed to a dog, and the label of a dog picture is changed to a cat. Then, the flipped dataset is used to update all model parameters.
[0115] Different from the traditional scheme, in order to test the contribution of different layers to the model accuracy, the present invention sequentially sets the model parameters of a specific layer to be updatable and the other layers to be non-updatable. For example, to quantify the anti-interference ability and sensitivity of the neurons in the i-th layer, first set the neurons in the i-th layer to be updatable and the other neurons to be non-trainable and updatable. Use the dataset with flipped labels to train the neurons in the i-th layer, and perturb the neurons in the i-th layer according to the samples of the poisoning attack. Then use the unflipped dataset for testing to obtain the first accuracy ACC 1i .
[0116] The difference evaluation is used to evaluate the change rate of the model accuracy after perturbing this layer, and this change rate reflects the recovery ability of the model. Specifically, taking the neurons in the $i$-th layer as an example, at this time, the neurons in the $i$-th layer are perturbed by the flipped label dataset. In the present invention, the neurons in this layer are fixed, and the neurons other than this layer are set to be trainable, and the model is trained using the unflipped dataset. This method can be used to evaluate whether the neurons in other layers can correctly predict the sample labels on the premise that the neurons in the $i$-th layer are perturbed.
[0117] Finally, the unflipped dataset is used to test the model to obtain the second accuracy ACC 2i 。
[0118] Finally, calculate $p$ 2i =|ACC 2i -ACC 1i |, the larger $p$ 2i , the stronger the anti-perturbation ability of the $i$-th layer, and the less important the $i$-th layer is, because on the premise that the neurons in this layer are perturbed, other layers can still correctly predict the labels; if $p$ 2i is smaller, it indicates that the anti-perturbation ability of the $i$-th layer is weaker, which means that this layer is more important, because on the premise that the neurons in this layer are perturbed, other layers cannot correctly predict the labels. The value range of $p$ 2i is $[0,1]$, and $p$ 2i is negatively correlated with the importance of this layer. Therefore, the present invention performs the following processing on $p$ 2i :
[0119] $p$ 2i ′=1 - $p$ 2i
[0120] The value range of $p$ 2i ′ is $[0,1]$, and the change trend of $p$ 2i ′ is positively correlated with the importance of this layer. At the same time, in order to magnify the gap in the sensitivity of different layers, $p$ 2i ′ is processed to obtain $s$ 2i , and $s$ 2i is used as the quantification of the sensitivity of the $i$-th layer, that is
[0121]
[0122] For all layers of the large language model, the model poisoning attack operation and the difference evaluation operation of the fixed layer need to be performed to obtain the sensitivity $s$ 21 , $s$ 22 ,..., $s$ 2n 。
[0123] Obtain the contribution $s$ of each layer of neurons to the accuracy of the LLM model11 , s 12 ,..., s 1n , and the sensitivity s of each layer of neurons 21 , s 22 ,..., s 2n After that, s 11 , s 12 ,..., s 1n and s 21 , s 22 ,..., s 2n are weighted and averaged to obtain the final importance score. Taking the i-th layer as an example, the importance score of the i-th layer
[0124] S i = α * s 1i + (1 - α) * s 2i
[0125] where α is a weight parameter, and its value range is [0, 1]. Since the value ranges of α, s 1i and s 2i are both [0, 1], the value range of S i is also [0, 1].
[0126] Step S2, fine-tuning layer and freezing layer screening. The present invention sets a preset importance threshold T and compares the importance score S i with it:
[0127] 1) If the importance score S of the layer i is greater than T, it means that this layer is relatively important. In order to reduce the impact of adding noise on the usability of the model, this layer is set as the freezing layer.
[0128] 2) If the importance S of the layer i is less than T, it means that this layer is not important. In order to protect data privacy, this layer is set as the fine-tuning layer, and fine-tuning and noise addition operations based on S i will be performed later.
[0129] Step S3, noise addition. Different from the traditional scheme, the present invention adds different noise values to different fine-tuning layers according to the magnitude of the importance score S i , (1 - S i ) * N(0, C 2 σ 2 I 2 ). That is, if the importance score s i is larger, less noise is added to the fine-tuning layer. If s i is smaller, more noise is added to this layer. And the DP-SGD parameter update algorithm is used to update the model parameters.
[0130] The following is the pseudocode of the algorithm of the present invention:
[0131]
[0132]
[0133]
[0134] The following is the training and testing process of the model:
[0135] First, divide the training dataset and the test dataset in a ratio of 7:3, denoted as D train and D test . The specific training and testing process includes three parts: detecting model training and testing, quantifying the sensitivity, and fine-tuning the model based on differential privacy noise.
[0136] (1) Detecting model training and testing
[0137] First, use the D train dataset to fine-tune the large language model. Note that no differential privacy noise is added during fine-tuning here, so as to obtain the model parameters without added noise. Then, use the D train dataset as the input to obtain the neuron output of each layer of the large language model, so as to obtain o1, o2,..., o n , use the features of D train as the input for detecting model training, and the label is the label y of D train . Use the backpropagation algorithm to update the model parameters of the detecting model. After training, use the D test dataset as the input to evaluate the detecting model, so as to obtain the accuracy rates of different layers, and then calculate s 11 , s 12 ,..., s 1n .
[0138] (2) Quantifying the sensitivity of different layers
[0139] Flip the labels of D train to obtain D train ′. For each layer, model poisoning attacks and difference evaluation operations need to be performed. Here, taking the i-th layer as an example, use D train ′ as the input, set the model parameters of the i-th layer to be updatable, while the other layers are not updatable, train the perturbed model parameters, and calculate the accuracy rate ACC test with D 1i as the input. Immediately afterwards, set the model parameters of the i-th layer to be non-updatable, while the other layers are set to be updatable, train this model with D train as the input, and calculate the accuracy rate ACC test with D2i , according to ACC 2i and ACC 1i to calculate s based on the difference 2i . For each layer, the above operations are performed to obtain s 21 , s 22 ,..., s 2n .
[0140] (3) Fine-tuning operation based on differential privacy
[0141] Based on S1, S2,..., S n perform the screening of the freezing layer and the fine-tuning layer, and use D train as the training data set. For each fine-tuning layer, add noise to it, and use the DP-SGD algorithm to update the model parameters, so as to obtain a highly available, efficient and data-protected model. And use D test to test this model.
[0142] The above are only optional embodiments of the present invention, and do not limit the patent scope of the present invention. All equivalent structural transformations made under the inventive concept of the present invention by using the content of the specification and drawings of the present invention, or directly / indirectly applied in other related technical fields are included in the patent protection scope of the present invention.
Claims
1. A noise addition method for the fine-tuning scenario of large language models, characterized in that, It includes the following steps: Obtain the importance score of each layer in the large language model according to the contribution degree of each layer to the accuracy of the large language model and the sensitivity of each layer; Screen the layers of the large language model according to the importance scores of different layers, use the layers with importance scores greater than or equal to the preset importance threshold as the frozen layers, and use the layers with importance scores less than the preset importance threshold as the fine-tuning layers; Add differential noise to different fine-tuning layers based on the importance scores.
2. The noise addition method for the large language model fine-tuning scenario according to claim 1, wherein, Use a probing model to obtain the contribution degree of each layer to the accuracy of the large language model. The input of the probing model is the output of the neurons in each layer of the large language model. The classification task of the probing model is the same as that of the large language model. Obtain the contribution degree of each layer to the model accuracy according to the accuracy of the probing model. The higher the accuracy of the probing model, the greater the contribution degree of the input layer.
3. The noise addition method for the large language model fine-tuning scenario according to claim 1, wherein Set labels for data samples according to the data characteristics of neurons, train the large language model using the flipped labels of each layer, perturb the neurons in each layer using the poisoned attack data samples, fix the perturbed neurons, and use the dataset without flipped labels to test the resilience of the model and evaluate the sensitivity of each layer.
4. The noise addition method for large language model fine-tuning scenarios according to claim 3, wherein The evaluation of the sensitivity of each layer includes two stages: model poisoning attack on the fixed layer and differential evaluation. The model poisoning attack on the fixed layer is used to perturb the neurons in the fixed layer, and the differential evaluation is used to evaluate the corresponding change rate of the model accuracy after perturbing each layer.
5. The noise addition method for the large language model fine-tuning scenario according to claim 4, wherein The model poisoning attack on the fixed layer specifically includes the following steps: Set the parameters of the large language model of a selected layer to be updatable in turn, and set the parameters of the models of other layers to be non-updatable; Use the dataset with flipped labels to train the neurons in the selected i-th layer, and perturb the neurons in the i-th layer using the poisoned attack samples; Then, use the dataset with unflipped labels for testing to obtain the first accuracy ACC of the large language model 1i .
6. The noise addition method for the fine-tuning scenario of large language models according to claim 5, characterized in that, The differential evaluation includes the following steps: In the large language model, fix the neurons in the i-th layer perturbed by any flipped label dataset, and set the neurons other than this layer to be trainable; Use the dataset without flipped labels to train the large language model; The model is tested using a dataset with unflipped labels to obtain the second accuracy ACC of the large language model 2i ; Evaluate the sensitivity of each layer in the large language model according to the difference between the first accuracy and the second accuracy.
7. The noise addition method for large language model fine-tuning scenarios according to claim 6, wherein For the i-th layer of the large language model, evaluating the sensitivity of each layer in the large language model according to the difference between the first accuracy and the second accuracy specifically includes: Calculate the difference p between the first accuracy and the second accuracy of the i-th layer of the large language model 2i = |ACC 2i - ACC 1i |; p 2i ranges from [0, 1], indicating the anti-disturbance ability of this layer, and p 2i is negatively correlated with the importance of this layer; For p 2i perform the following processing to obtain p 2i ′ = 1 - p 2i where p 2i ′ is an intermediate variable, and the value range of p 2i ′ is [0, 1], and the change trend of p 2i ′ is positively correlated with the importance of this layer; Process p 2i ′ to obtain the sensitivity level s of the i-th layer 2i : Among them, p 21 ′, p 22 ′,..., p 2n ′ represent the intermediate variables of the 1st, 2nd,..., nth layers of the large language model, where n is the total number of layers of the large language model.
8. The noise addition method for the large language model fine-tuning scenario according to any one of claims 1-7, characterized in that Perform weighted averaging on the contribution degree of each layer to the accuracy of the large language model and the sensitivity of each layer to obtain the importance score of each layer in the large language model.
9. A noise addition system for the fine-tuning scenario of large language models, characterized in that, It includes a processor that can execute the steps of the noise addition method for the large language model fine-tuning scenario according to any one of claims 1-8.
10. A storage medium, characterized in that, A program is stored on the storage medium, and the program can execute the steps of the noise addition method for the large language model fine-tuning scenario according to any one of claims 1-8.