Method for improving safety of large model of end side equipment

By constructing a safe and unsafe QA-based sample training classifier, evaluating the neuron safety and importance of large language models, and eliminating unsafe parameters, solving the robustness and security problems of the end-side device large model, realizing the security improvement of model compression.

CN120258075APending Publication Date: 2025-07-04ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510362835.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The pruning technology of large language models in end-side devices increases the risk of attack by malicious attackers and reduces the robustness of the end-side device model.

Method used

By constructing safe and unsafe QA pair samples, a linear classifier is trained to evaluate the safety and importance of each layer of neurons, prune unimportant and unsafe parameters, retain important and safe parameters, and form a pruned model.

Benefits of technology

On the premise of ensuring model performance, the security of the end-side device large model is improved, the number of parameters is reduced, and the inference speed is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258075A_ABST
    Figure CN120258075A_ABST
Patent Text Reader

Abstract

The invention discloses a method for improving the safety of a large model of end-side equipment, and the method comprises the steps: 1, dividing safe and unsafe QA pair samples into a training set and a test set at random according to a proportion, and inputting the training set and the test set into the large model to obtain an activation value of each layer of neurons of the model; 2, a linear classifier is trained on the training set according to the corresponding activation and labels, and the degree of correlation between neurons of the layer and the model safety is judged through the prediction precision of the verification set; 3, inputting a conventional text into the large model, obtaining the importance of each parameter in a substructure of each layer of neurons according to an activation value, weighting and combining the importance of each parameter with prediction precision to serve as an evaluation index, retaining a certain proportion of parameters according to actual requirements, and setting other parameters to be zero to obtain a pruned model; and step 4, deploying the pruned model to an end side device. According to the method, the safety of the large model of the end side equipment is improved, and it is ensured that model compression does not cause serious safety risks of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a method for improving the security of large models on edge devices. Background Art

[0002] Nowadays, large language models have been widely applied in various edge scenarios such as mobile phones, cars, and embedded devices. The large language model pruning technology can reduce model parameters and accelerate model inference speed without sacrificing too much performance by finding redundant parameters in the large language model. Since the technology was proposed, it has attracted wide attention in the academic and industrial circles. Applying pruning technology in large language models is an emerging direction to solve problems such as high deployment conditions and slow inference speed of large language models. However, as an additional step in the deployment of large language models on edge devices, the large language model pruning technology will increase the attack risk of malicious attackers on the large model system of edge devices and reduce the robustness of the large model on edge devices.

[0003] Therefore, designing a method for improving the security of large models on edge devices has guiding significance for improving the security of large models in edge scenarios. Summary of the Invention

[0004] The purpose of the present invention is to propose a method for improving the security of large models on edge devices in view of the deficiencies of the prior art. This method can calculate the correlation degree between each structure in the large model and the model security, and prune the large model in combination with the importance degree of each parameter of the model, providing guidance for improving the model security of large models on edge devices.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A method for improving the security of large models on edge devices includes the following steps:

[0007] Step 1: Connect the malicious question Q i and its safe reply and the unsafe reply to construct safe QA pair samples and unsafe QA pair samples respectively, so as to obtain the QA pair sample data set I, where N is the total number of samples; and input the samples in I into the base large model M to obtain the activation values H = {h1, h2,..., h i ,..., h L} of each layer of neurons in the model, where h i represents the activation value of the i-th layer neuron in the base large model M, and L is the number of neuron layers in the base large model; randomly select the activation value h i in H and its corresponding label y i to construct the training set D in proportiontrain and the test set D test , ensuring that the number of safe samples and unsafe samples in the training set and the test set is the same;

[0008] Step 2: Based on the activation value h of each sample train on the training set D i and the corresponding label y i train a classifier C, calculate the loss during the training process, and optimize the classifier parameters to improve the classification accuracy of the classifier; for each layer of neurons in the base large model (taking the l-th layer as an example), input all the activation values test in the test set D that belong to the neurons of this layer into the classifier to obtain the predicted label Calculate the prediction accuracy of the classifier C for the neurons of this layer to measure the degree of correlation between the neurons of this layer and the model security;

[0009] Step 3: Input the regular text T = {t1, t2, …, t V}, where V is the total number of samples in T, into the base large model M to obtain the activation values X i = {x i1 , x i2 , …, x iW} of each layer of neurons in the base large model, where i represents the i-th layer of neurons and W represents the number of substructures in each layer of neurons. Obtain the importance of each parameter in each substructure of the neurons of this layer based on the activation values, and perform a weighted sum with the prediction accuracy obtained in Step 2. Use the resulting value as the evaluation index for the importance of the parameter to the model security; and retain a certain proportion of the parameters according to actual needs, and set the other parameters to zero to obtain the pruned model;

[0010] Step 4: Deploy the pruned model to the edge device.

[0011] In the above technical solution, further, the safe QA pair samples and the unsafe QA pair samples in Step 1 are specifically expressed as: where {question}, {safe_answer}, and {unsafe_answer} are respectively replaced with the text of the malicious question Q i , the safe reply and the unsafe reply .

[0012] Further, the value of the label y i is: when the sample is a safe reply, the value of the label y i is 1; when the sample is an unsafe reply, the label yi The value of

[0013] Further, the classifier C is constructed as follows: The classifier C is a linear classifier. The input is a vector with the same shape as the activation values of each layer of neurons in the base large model M, and the output is a constant y, and this constant y ∈ [0, 1]. If it exceeds this range, the adjacent boundary value is taken.

[0014] Further, the measurement of the prediction effect of the classifier C uses the cross-entropy loss function:

[0015]

[0016] where L l (C) is the cross-entropy loss function value corresponding to the neurons in the l-th layer of the base large model. Define {y1, y2, …, y i , …, y N} as the true classification labels of the activation values of the neurons in the corresponding layer of the base large model, and {y1 ′ , y2 ′ , …, y i ′ , …, y ′ N} represents the predicted labels of the activation values of the neurons in the corresponding layer of the base large model output by the classifier C;

[0017] To improve the accuracy of the classifier C in predicting the degree of relevance between a certain layer of neurons in the base large model and the model security, it is necessary to continuously optimize the classifier C by minimizing the loss function:

[0018]

[0019] where θ c is the parameter used in the training of the classifier C.

[0020] Further, for the test set D test The predicted label of the l-th layer obtained The degree of relevance between the l-th layer of neurons and the model security is measured according to the prediction accuracy of the l-th layer of neurons. The prediction accuracy E l of the l-th layer of neurons is the negative value of the cross-entropy loss function:

[0021] E l = -L l (C)

[0022] where the prediction accuracy corresponding to all sub-structures (including but not limited to each parameter matrix in the multi-layer perceptron and the attention mechanism) in the l-th layer of neurons is equal to the prediction accuracy E l of this layer of neurons.

[0023] Further, in step 3, the calculation formula of the index for evaluating the importance of parameters in the sub-structures of each layer of neurons is as follows:

[0024]

[0025] Wherein, is the importance index value of the k-th parameter in the i-th sub-structure, W ik is the value of the k-th parameter in the i-th sub-structure, x ik is the k-th activation value of the i-th sub-structure, and ‖·‖2 represents the 2-norm of the vector.

[0026] Further, in step 3, the evaluation index E ik for evaluating the security and importance of model parameters is calculated as follows:

[0027]

[0028] Wherein, λ is an adjustable trade-off parameter used to adjust the degree of demand for security and importance; is the prediction accuracy of each layer of neurons of the base large model. For all parameters in the i-th sub-structure belonging to the l-th layer of the base large model, E ik The larger it is, the greater the importance of the k-th parameter of the i-th sub-structure of the corresponding layer of neurons in the model.

[0029] Further, a malicious question dataset Q = {q1, q2, …, q N} can be used to jailbreak the large model of the edge device obtained through various large language model jailbreaking methods, test the jailbreaking success rate and the performance of the large model of the edge device, and calculate the pruning robustness score by weighted calculation. The higher the score, the better the pruning effect. The pruning robustness score RobustScore is calculated as follows:

[0030] RobustScore = Score performance - μAttack successful

[0031] Wherein, Score performance represents the performance of the model, Attack successful represents the attack success rate of the model when facing jailbreaking attacks, and μ is a weighted parameter.

[0032] The beneficial effects of the present invention are:

[0033] In view of the current situation of the increased security risks in large model compression technology, the present invention proposes a method for enhancing the security of large models on edge devices. This method can ultimately achieve secure pruning of large models, reduce model parameters while ensuring model security and performance, accelerate model inference speed, and provide guidance for the secure application of model compression technology in the deployment of large models on edge devices.

[0034] The method of the present invention aims to reduce the parameters of the large language model, reduce the model size, and accelerate the inference speed while ensuring model performance and model security by finding the parts with relatively low security relevance in the large language model and pruning them on the premise of considering the impact of internal parameters on performance. The present invention trains a classifier and innovatively uses the prediction accuracy of the classifier on each layer as the evaluation criterion for the security of the corresponding layer of the large language model, achieving a high classification accuracy. Different from conventional methods that only consider parameter importance, the present invention also considers the impact of the parameter on model security and takes both factors into account in the evaluation index to avoid pruning parameters that affect model security. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a flowchart of the method for enhancing the security of the large model on the edge device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The following further describes the technical solutions of the present invention in combination with the specific embodiments and drawings of the present invention, but the present invention is not limited to these embodiments.

[0037] The flowchart shown in the drawings is only an exemplary illustration and does not necessarily include all steps. For example, some steps can be decomposed, while some steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0038] The present invention proposes a method for enhancing the security of large models on edge devices. First, randomly divide the secure and insecure QA pair samples into a training set and a test set in proportion, input them into the base large model to obtain the activation values of each layer of neurons in the model; train a linear classifier on the training set according to the corresponding activation values and labels, and use the prediction accuracy of the test set to judge the relevance of the neurons in this layer to model security; input conventional text, use the weighted combination of the product of the second-order norm of the activation value and the model parameters and the prediction accuracy as the evaluation index, retain a certain proportion of parameters according to actual needs, and set the remaining parameters to zero to obtain the pruned model; then deploy the pruned model to the edge device. Use the dangerous question dataset to test the jailbreak success rate and performance of the large model on the edge device, and calculate the robustness score of the pruning. The higher the score, the better the pruning effect.

[0039] AsFigure 1 As shown in the figure, the method of the present invention specifically includes the following steps:

[0040] Step 1: Connect the malicious question Q i and its safe reply and the unsafe reply to construct a safe QA pair sample and an unsafe QA pair sample to obtain the QA pair sample dataset I, where N is the total number of samples. For a given malicious question Q and the corresponding reply A, the corresponding sample pair is "Question: {Q}; Answer: {A}.", where {Q} and {A} are replaced with the corresponding malicious text Q and reply A text. Specifically, where {question}, {safe_answer}, and {unsafe_answer} are respectively replaced with the text of the malicious question Q i , the safe reply and the unsafe reply .

[0041] Input the samples in I into the base large model M to obtain the activation values H = {h1, h2,..., h i ,..., h L} of each layer of neurons in the model, where h i represents the activation value of the i-th layer of neurons in the base large model M, and L is the number of neuron layers in the base large model; randomly select the activation value h i in H and its corresponding label y i (if the sample is a safe reply, the value of the label y i is 1, otherwise it is 0) to construct the training set D train and the test set D test , ensuring that the number of safe samples and unsafe samples in the training set and the test set is the same.

[0042] Step 2: Train a classifier C on the training set D train according to the activation value h i of each sample and the corresponding label y i . During the training process, calculate the loss, optimize the classifier parameters, and improve the classification accuracy of the classifier; for each layer of neurons in the base large model (taking the l-th layer as an example), input all the activation values test in the test set D that belong to this layer of neurons into the classifier to obtain the predicted label and calculate the prediction accuracy of the classifier C on this layer of neurons to measure the correlation degree between this layer of neurons and the model security.

[0043] The classifier C is a linear classifier. The input is a vector with the same shape as the activation values of each layer of neurons in the base large model M, and the output is a constant y, where y ∈ [0, 1]. If it exceeds this range, the adjacent boundary value is taken. The cross-entropy loss function is used to measure the prediction effect of the classifier C:

[0044]

[0045] Among them, L l (C) is the cross-entropy loss function value corresponding to the neurons in the l-th layer of the base large model. Define {y1, y2, …, y i , …, y N} as the true classification labels of the activation values of the neurons in the corresponding layer of the base large model, and {y1 ′ , y2 ′ , …, y i ′ , …, y ′ N} represents the predicted labels of the activation values of the neurons in the corresponding layer of the base large model output by the classifier C. To improve the accuracy of the classifier in predicting the degree of relevance of a certain layer of the base large model to model security, the classifier needs to continuously optimize by minimizing the loss function:

[0046]

[0047] Among them, θ c is the parameter used in the training of the classifier C.

[0048] For the test set D test The predicted label of the l-th layer obtained The degree of relevance of the neurons in the l-th layer to model security is measured according to the prediction accuracy of the neurons in the l-th layer. The prediction accuracy E l of the neurons in the l-th layer is the negative value of the cross-entropy loss function:

[0049] E l = -L l (C)

[0050] Among them, the prediction accuracy corresponding to all sub-structures in the neurons of the l-th layer is equal to the prediction accuracy E l of the neurons in this layer.

[0051] Step 3: Input the regular text T = {t1, t2, …, t V}, where V is the total number of samples in T, into the base large model M to obtain the activation values X i = {x i1 , x i2 , …, x iW}, where i represents the i-th layer of neurons, W represents the number of substructures in each layer of neurons. Based on the activation values, the importance of each parameter in each substructure of this layer of neurons is obtained. Then, the index for evaluating the importance of parameters in the substructure is calculated as follows:

[0052]

[0053] where, is the importance index value of the l-th parameter in the i-th substructure, W ik is the value of the k-th parameter in the i-th substructure, x ik is the k-th activation value of the i-th substructure, and ‖·‖2 represents the 2-norm of the vector.

[0054] The is weighted and summed with the prediction accuracy obtained in step 2, and the resulting value is used as the evaluation index for the security and importance of this parameter. The calculation formula is as follows:

[0055]

[0056] where λ is an adjustable trade-off parameter used to adjust the degree of demand for security and importance. is the prediction accuracy of each layer of neurons in the base large model. For all parameters in the i-th substructure belonging to the l-th layer of the base large model, E ik the larger it is, the greater the importance of the k-th parameter in the i-th substructure of the corresponding layer of neurons in the model.

[0057] According to the actual needs, a certain proportion of parameters are retained, and the other parameters are set to zero to obtain the pruned model.

[0058] Step 4: Deploy the pruned model to the edge device. The edge device large model has good security while maintaining the original performance.

[0059] Using a malicious question dataset Q = {q1, q2,..., q n}, the obtained edge device large model is jailbroken through various large language model jailbreaking methods, and the jailbreaking success rate and performance of the model are tested. The robustness score of pruning is calculated as follows:

[0060] RobustScore = Score performance - μAttack successful

[0061] where Score performance represents the performance of the model, and Attack successfulIt represents the attack success rate of the model when facing jailbreak attacks, μ is a weighted parameter, and the higher the robustness score, the better the pruning effect.

[0062] Taking the base model Llama-2-13B-chat with a pruning ratio of 50% as an example, the base large model is processed using the conventional pruning method and the method of the present invention, and the security and performance of the models before and after the processing are compared. As shown in Table 1, after being processed using the conventional pruning method, compared with the original base large model, the accuracy of the processed model remains basically unchanged, but the security performance is significantly reduced; and compared with the end-side device large model obtained by the conventional pruning method, the end-side device large model obtained by the method of the present invention has an average attack success rate of 8.7% when facing three jailbreak attack methods, and the performance on 7 zero-sample tasks has only decreased by 0.32%. It shows that compared with the conventional pruning method, the method of the present invention can better alleviate the decline in model security caused by model compression.

[0063] Table 1

[0064]

[0065] The seven zero-shot tasks include: Boolq, Rte, Hellaswag, Winogrande, Arc_easy, Arc_challenge, and Openbookqa.

[0066] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.

Claims

1. A method for enhancing the security of large models on edge devices, characterized in that Including the following steps: Step 1: Connect the malicious problem Q i and its safe reply with the unsafe reply to construct safe QA pair samples and unsafe QA pair samples Thereby obtaining the QA pair sample dataset I, where N is the total number of samples; and input the samples in I into the base large model M to obtain the activation values H = {h1, h2, …, h i , …, h L} of each layer of neurons in the model, where h i represents the activation value of the i-th layer neuron of the base large model M, and L is the number of layers of neurons in the base large model; randomly select the activation value h i in H and its corresponding label y i to construct the training set D train and the test set D test to ensure that the number of safe samples and unsafe samples in the training set and the test set is the same; Step 2: On the training set D train train a classifier C according to the activation value h i of each sample and the corresponding label y i ; For each layer of neurons in the base large model, all the activation values belonging to the neurons in this layer in the test set D test are input into the classifier to obtain the predicted labels, and the prediction accuracy of the classifier C for the neurons in this layer is calculated to measure the degree of relevance between the neurons in this layer and the model security; Step 3: Input the regular text T = {t1, t2, …, t v}, where V is the total number of samples in T, into the base large model M to obtain the activation values X i = {x i1 , x i2 , …, x iw}, where i represents the i-th layer of neurons and W represents the number of sub-structures in each layer of neurons. Obtain the importance of each parameter in each sub-structure of this layer of neurons based on the activation values, and perform weighted summation with the prediction accuracy obtained in Step 2. Take the resulting value as the evaluation index of the importance and security of this parameter for the model; And retain a certain proportion of parameters according to actual needs, and set other parameters to zero to obtain the pruned model; Step 4: Deploy the pruned model to the edge device.

2. The security improvement method for the large model of the edge device according to claim 1, wherein, The safety QA pairs of samples in Step 1 and the unsafe QA pairs of samples are specifically expressed as follows: where {question}, {safe_answer}, and {unsafe_answer} are respectively replaced with the text of malicious question Q i , safe answer and unsafe answer for replacement.

3. The security improvement method for the large model of the edge device according to claim 1, characterized in that, Label y i takes the value: when the sample is a safe reply, label y i takes the value 1; when the sample is an unsafe reply, label y i takes the value 0.

4. The security improvement method for the large model of the edge device according to claim 1, characterized in that, The classifier C in the said step 2 is a linear classifier, the input is a vector with the same shape as the activation value of each layer of neurons in the base large model M, the output is a constant y, and this constant y ∈ [0, 1]. If it exceeds this range, the adjacent boundary value is taken.

5. The method for improving the security of the large model of the edge device according to claim 4, wherein, The prediction effect of the said classifier C is measured using the cross-entropy loss function: Among them, L l (C) is the cross-entropy loss function value corresponding to the l-th layer neurons of the base large model. Define {y1, y2, …, y i , …, y N} as the true classification labels of the activation values of the neurons in the corresponding layer of the base large model, and {y1 ′ , y2 ′ , …, y i ′ , …, y ′ N} represents the predicted labels of the activation values of the neurons in the corresponding layer of the base large model output by the classifier C; To improve the accuracy of the classifier C in predicting the degree of relevance of a certain layer of neurons in the base large model to model security, it is necessary to continuously optimize the classifier C by minimizing the loss function: where θ c is a parameter used in the training of classifier C.

6. The security improvement method for the large model of the edge device according to claim 5, wherein, For the test set D test The predicted label of the l-th layer obtained The correlation degree between the l-th layer neuron and the model security is measured according to the prediction accuracy of the l-th layer neuron. The prediction accuracy E of the l-th layer neuron l The calculation formula is as follows: E l = -L l (C) Among them, the prediction accuracy corresponding to all substructures in the l-th layer of neurons is equal to the prediction accuracy E of the neurons in this layer. l Equal.

7. The security improvement method for the large model of the edge device according to claim 6, wherein, In step 3, the calculation formula of the index for evaluating the importance of parameters in the sub-structures of neurons in each layer is as follows: Among them, is the importance index value of the k-th parameter in the i-th sub-structure, W ik is the value of the k-th parameter in the i-th sub-structure, x ik is the k-th activation value of the i-th sub-structure, and ‖·‖2 represents the 2-norm of the vector.

8. A method for improving the security of a large model on an edge device according to claim 7, characterized in that In step 3, the evaluation index E for judging the security and importance of model parameters ik has the following calculation formula: where λ is an adjustable trade-off parameter used to adjust the degree of demand for security and importance; is the prediction accuracy of the neurons in each layer of the base large model.