Software security vulnerability mining method for natural language processing
By preprocessing and generating a training subset with triggers, combining NLP model and backdoor vulnerability mining technology, the problem of detecting backdoor vulnerabilities of NLP model under the condition of mixing a small number of backdoor samples is solved, and efficient and hidden backdoor vulnerability detection is achieved.
Patent Information
- Application Number
- CN202510101215.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to detect whether there are backdoor vulnerabilities in the NLP model under the condition of mixing a small number of backdoor samples, and the existing methods are easily noticed by users when detecting backdoor vulnerabilities. As the amount of training data increases, it is unrealistic to mix a large number of backdoor samples.
By preprocessing the vulnerability code dataset sample, the backdoor trigger of the vulnerability code is determined, the training subset with triggers is generated, the model is generated for backdoor vulnerability code detection, the detection metrics of the vulnerability code are calculated, and the backdoor vulnerability is mined.
It is realized that when a small number of backdoor vulnerabilities are injected, it can effectively detect whether there are backdoor vulnerabilities in the NLP model, and the detection method is highly effective and concealed, avoiding the problem of inclusion of a large number of backdoor samples.
Smart Images

Figure CN120012109A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a software security vulnerability mining method for natural language processing. Background Art
[0002] At present, the field of Natural Language Processing (NLP) is developing rapidly, and NLP models are widely used in real-world scenarios. With the widespread application of NLP in key mission areas, people have begun to pay attention to the security issues of NLP models. In 2014, researchers discovered that there were security vulnerabilities in neural network models. By making slight perturbations to samples, the predictions of neural network models could be made wrong. Subsequently, a large number of researchers began to pay attention to the security issues of neural networks, constantly exploring the security vulnerabilities of neural network models and proposing defense methods.
[0003] In 2017, researchers confirmed that there may be backdoor vulnerabilities in deep neural network models, and proposed the BadNets method in the image field to confirm that deep neural network models for image classification have backdoor vulnerabilities. In 2019, researchers began to consider that NLP models have backdoor vulnerabilities and proposed the Insent method. By inserting an irrelevant sentence into the text to generate a backdoor training set, the NLP model was retrained using the backdoor training set. It was found that there was a probability of nearly 96% that the model would be misclassified, which ultimately showed that the NLP model had serious backdoor vulnerabilities. In addition, methods such as Insent and RIPPLES generate poorly concealed backdoor samples by inserting meaningless sentences and uncommon words. This sample generation method is easy to detect and defend. At present, there are few studies on software security vulnerability detection and related defense of text backdoors, but given that NLP models have a wide range of practical application prospects, such as machine translation, grammar detection, content filtering, and fraud detection systems, it is particularly important and urgent to evaluate and detect security vulnerabilities in NLP models.
[0004] In the past three years, large language models have developed very rapidly. Training models requires a lot of time and cost. Researchers have proposed that pre-trained models can be used to fine-tune the models according to downstream tasks. At the same time, most users use data sets obtained from the Internet to fine-tune pre-trained models. In the above process, the sources of the pre-trained models and data sets used by users are unclear. Attackers can easily obtain training data sets and pre-trained models, and even manipulate the training process of the models, which poses a greater risk of backdoors being injected into the NLP model. What's more serious is that the backdoor vulnerability of the model will only be activated by poisoned samples with backdoor triggers, and the prediction of normal samples is normal, which makes it difficult for users to discover backdoor vulnerabilities in the model, and there are huge challenges in detecting and defending backdoor vulnerabilities.
[0005] Most model backdoors are implemented by mixing backdoor samples into the training set, but current methods can only successfully insert backdoors when 20%-40% of backdoor samples are mixed in. Such a large number of backdoor samples will be easily noticed by ordinary users; and as the amount of training data increases, it is unrealistic to mix in a large number of backdoor samples. This invention aims to study whether the existence of backdoor vulnerabilities in NLP models can be detected under the condition of mixing in a small number of backdoor samples. Summary of the invention
[0006] The purpose of the present invention is to provide a software security vulnerability mining method for natural language processing to detect whether a text classification program trained with a small number of poisoned samples has a backdoor vulnerability.
[0007] The technical solution adopted by the present invention is a software security vulnerability mining method for natural language processing, comprising the following steps:
[0008] Step S1, preprocessing vulnerability code dataset samples;
[0009] Step S2, determining the backdoor trigger of the vulnerability code;
[0010] Step S3, generating a trigger for the vulnerability code;
[0011] Step S4, generating a training subset with triggers;
[0012] Step S5, generating a model for backdoor vulnerability code detection;
[0013] Step S6, calculating the detection index of the vulnerability code and mining the backdoor vulnerability.
[0014] Furthermore, the vulnerability code dataset of step S1 includes an independent training set and a test set. First, the label t is set as the target label. pre Extract the sample set labeled as the target label as the target label training subset to be preprocessed in, represents the target label training subset to be preprocessed, N represents The number of samples in express The last sample in ; then the sample is converted into a numerical value through the BERT model and normalized. Specifically:
[0015]
[0016] in, represents the target label sample to be preprocessed, x i represents the target label sample after preprocessing, express The mean of express The variance of i express The value corresponding to the conversion of the i-th word in, R represents The number of words in a R express The value converted from the last word in ;
[0017] The target label training subset to be preprocessed Preprocessing is performed to obtain the preprocessed target label training subset D t , perform the same operation on other samples and finally obtain the preprocessed training set D.
[0018] Furthermore, in step S2, the target label training subset D obtained after preprocessing in step S1 is t , extract the important feature set of each preprocessed target label sample through the class activation mapping method. The specific steps are as follows:
[0019] S21, obtain the features of a single sample, and convert the preprocessed target label sample x i Input into the neural network, extract the value and weight information of the last layer output in the neural network, and obtain the preprocessed target label sample x through weighted calculation i The eigenvalue of is as follows:
[0020]
[0021] Among them, CAM i Represents the preprocessed target label sample x i The characteristic value of represents the weight of the kth neuron of the last layer of the neural network for the target label t, C i,k represents the output of the kth neuron of the last layer of the neural network for the i-th word, and A represents the number of neurons in the last layer of the neural network;
[0022] S22, CAM of the feature value of the obtained preprocessed target label sample i Filter and select the value greater than CAM i The feature set of the intermediate value is used as the important feature set F of the preprocessed target label sample i , and then select the words with the most occurrences from the important feature set as the feature words f of the preprocessed target label sample i , the specific formula is as follows:
[0023] x i ={w1,w2,…,w B}
[0024] M i =median(CAM i )
[0025] F i ={w j}where CAM i [j]>M i andw j ∈x i
[0026] f i =max(count(F i ))
[0027] Among them, x i represents the target label sample after preprocessing, median(·) represents the median of the obtained vector, CAM i Represents x i The eigenvalue of is a vector, M i Indicates CAM i The median of , F i Represents x i The important feature set, w j Represents x i The jth word in CAM i [j] indicates CAM i The jth value in count(F i ) indicates the i The number of occurrences of all words in the , max(·) means selecting the word with the largest number of occurrences, and B represents the preprocessed target label sample x i Quantity, w B Represents x i The last word of
[0028] The preprocessed target label training subset D t Each sample in performs the operations of S21 to S22, and finally the feature words of all the preprocessed target label samples constitute a feature word set The details are as follows:
[0029]
[0030] in, Indicates D t The feature word set, f N express The last characteristic word of .
[0031] Furthermore, in step S3, the characteristic word set of the target label training subset is The number of occurrences of all words in are counted and sorted in descending order. The first m frequent words are selected as the trigger set of the vulnerability code dataset on the target label t. The specific set is as follows:
[0032]
[0033] Among them, Tri t Represents the trigger set, Indicates trigger set Tri t The last word in the trigger set Tri t The number of trigger words in the .
[0034] Furthermore, the specific steps of step S4 are as follows:
[0035] S41: According to the modification ratio p, randomly select some samples of non-target labels in the training set D of the vulnerability code dataset as the training subset D to be triggered p , D p The non-target label training samples in are denoted as x p , the specific calculation formula for the modified ratio p is:
[0036]
[0037] in, Indicates the training subset D to be triggered p The number of samples, N D Represents the number of samples in the training set D;
[0038] S42: Select a replacement position, execute the operation of step S2 on the non-target label training sample, and select a non-target label training sample with a feature value greater than x p The word corresponding to the feature of the median feature value is used as an important feature set F for non-target label samples p , for the important feature set F p The occurrence counts of all words in are counted and sorted in descending order, and the first b frequent words are selected as non-target label training samples x p The feature word set is expressed as:
[0039] M p =median(CAM p )
[0040] F p ={w a}where CAM p [a]>M p
[0041] f p ={wp1 ,w p2 ,…,w pb}
[0042] Among them, CAM p Represents non-target label training sample x p The eigenvalue is a vector, median(·) means taking the median, M p Indicates CAM p The median of , F p Represents non-target label training sample x p The important feature set, w a Represents non-target label training sample x p The ath word in CAM p [a] indicates CAM p The ath value in p b Represents non-target label training sample x p The last replacement position in f p Non-target label training sample x p The feature word set, w pb Represents non-target label training sample x p The word at the last replacement position in ;
[0043] f p The word in the non-target label training sample x p The position is taken as the replacement position, recorded as Pos p ={p1,p2,…,p b}, where b represents the replacement position Pos p the number of
[0044] S43: Based on the trigger set Tri on the target tag t obtained in step S3 t and the non-target label training sample x obtained in S42 p Replacement position Pos p , using Tri t The trigger words in iteratively replace the non-target label training samples x p Replacement position Pos p The corresponding words on the left are repeated until the specified perturbation amount is reached, and the first candidate vulnerability test sample is generated, denoted as x' p1 ; The selected triggers and replacement positions are randomly combined, and the perturbation amount is calculated as follows:
[0045] Pr(x p )=N tri / len(x p )
[0046] Among them, N trirepresents the number of replacement positions that have been replaced, len(·) represents the length, Pr(·) represents the perturbation amount, x p Represents non-target label training samples;
[0047] S44: Calculate the perplexity change value ΔPPL(x' p1 ,x p ), ΔPPL(x' p1 ,x p ) exceeds the set threshold, the replacement is canceled, and the replacement operation in step S43 is re-executed to select other combinations to replace the non-target label training samples to generate new candidate vulnerability test samples; the j-th replacement operation is performed, and the generated candidate vulnerability test samples The calculated perplexity change value If the set threshold is not exceeded, the candidate vulnerability test sample generated for the jth time For the final vulnerability test sample x' p , and the final vulnerability test sample x' p The label is set to the target label t. Specifically, the calculation formula is:
[0048]
[0049] Among them, w i Represents candidate vulnerability test samples The i-th word in p(w i |w <i ) means that according to the context w <i ={w1,w2,…,w i-1 Prediction i The probability of Represents non-target label training sample x p The candidate vulnerability test samples generated when performing the jth replacement in S43, N p express The number of words in express The perplexity value, PPL(x p ) represents the non-target label training sample x p The confusion value, Represents the candidate vulnerability test sample generated for the jth time The perplexity change value of
[0050] S45: Training subset D to be triggered p Each sample in performs the operations of steps S42 to S44, and all the final vulnerability test samples generated constitute the training subset D' with triggers p .
[0051] Furthermore, the specific steps of generating a backdoor vulnerability code detection model in step S5 are as follows:
[0052] S51, the training subset D' with trigger obtained in S4 p Add to the training set D, and get the training set with trigger D'=D+D' p ;
[0053] S52, initially setting a learning rate, a training batch, and a batch size according to the environment settings;
[0054] S53, input the training set D' with triggers generated in S51 into the BERT model F(·) for training, and obtain the BERT model F'(·) for backdoor vulnerability detection.
[0055] Furthermore, the specific steps of step S6 are as follows:
[0056] S61, select samples in the test set T whose labels are not the target labels as the test subset T of non-target labels p , T p All samples in the test set perform the operation of step S4 to obtain the test subset T' with trigger p ={x'1,x'2,…,x' Q}, where Q represents the test subset T' with triggers p The number of samples in x' Q Indicates T' p The last sample in ;
[0057] S62, the test subset T' with trigger p Put it into the BERT model F'(·) for backdoor vulnerability detection for prediction, and get the predicted label set of the test subset with triggers, as follows:
[0058]
[0059]
[0060] Among them, T' p represents a test subset with triggers, Indicates T' p The predicted label set, Q represents T' p The number of samples, F'(·) represents the BERT model for backdoor vulnerability detection, represents the last sample x' of the test subset with trigger Q Predicted labels in the BERT model F'(·) for backdoor vulnerability detection;
[0061] S63, statistics of predicted label sets for the test subset with triggers All labels in the _ are the number of target labels t, and the attack success rate ASR on the vulnerability code dataset and BERT model is calculated. The specific formula is as follows:
[0062]
[0063] Among them, x' q Represents a test subset T' with triggers p The qth sample in Represents x' q Predicted labels in the BERT model F'(·) for backdoor vulnerability detection, t represents the target label, and Q represents the test subset T' with triggers p The number of samples, I1(·) represents the calculation of the test subset T' with trigger p The predicted label of the qth sample in Whether it is a function of the target label, ASR represents the attack success rate;
[0064] S64, input the test set T into the BERT model F'(·) for backdoor vulnerability detection to obtain the predicted label set of the test set Determine the predicted label set for the test set The original label set of each label in the test set Whether the corresponding labels are consistent, the classification accuracy CACC of the BERT model for backdoor vulnerability detection in clean samples is calculated by counting the number of consistent labels. The calculation formula is as follows:
[0065]
[0066] in, represents the predicted label of the zth sample in the test set, represents the original label of the zth sample in the test set, H represents the number of samples in the test set T, represents the last predicted label of the prediction set, represents the last original label of the prediction set, CACC represents the classification accuracy of the model under clean samples, and I2(·) represents the function that calculates whether the predicted label and the original label of the zth sample in the test set are the same;
[0067] S65: Evaluate test subset T' with triggers p The specific formula for the fluency change is as follows:
[0068]
[0069] Among them, x q is the test subset T of non-target labelsp The qth non-target label test sample in q Represents a test subset T' with triggers p The qth non-target label test sample in , Q represents the test subset T' with trigger p The number of samples in , ΔPPL(·) represents the fluency change value;
[0070] S66: Calculate the test subset T of non-target labels p Non-target labeled test samples x in q , test subset T' with trigger p The test sample x' with trigger q Similarity, non-target label test sample x q and the test sample x' with trigger q Through the USE model, we get the vector representation v q 、v' q , calculate the cosine similarity of the two vectors, and then take the average to get the test subset T' with triggers p The USE value of is as follows:
[0071] v q =F USE (x q )
[0073] v' q =F USE (x' q )
[0074]
[0075] Among them, F USE (·) represents the USE model, cos(·) represents the cosine similarity, and v q Represents the non-target label test sample x q By USE Model F USE (·) The vector representation obtained is v' q Represents a test sample x' with a trigger q By USE Model F USE (·) represents the vector obtained, ||·|| represents the length of the vector, and USE(·) represents the test subset T' with triggers p The cosine similarity of p represents the test subset of the SST-2 test set T whose labels are non-target labels, T' p represents a test subset with a trigger, Q represents T' p The number of samples in ;
[0076] S67: Use syntax error variation value to evaluate the test subset T' with triggers p The quality of , first use the grammar checker to check the non-target label test sample x q and the test sample x' with trigger q The syntax errors are counted; then the test sample x' with trigger is calculated q Test sample x with non-target label q The change value of syntax errors is averaged to obtain the test subset T' with triggers p The syntax error change value is as follows:
[0078] ΔG(x' q ,x q )=G(x' q )-G(x q )
[0079]
[0080] Among them, x q The test subset T represents the non-target labels p The qth sample in T p represents the test subset of non-target labels in the SST dataset, T' p represents a test subset with triggers, x' q represents a test sample with a trigger, Q represents T' p The number of samples in G(x q ) represents the non-target label test sample x q The number of syntax errors, G(x' q ) represents the test sample x' with trigger q The number of grammatical errors, ΔG represents the non-target label test sample x q and the test sample x' with trigger q Syntax error change value, ΔG(·) represents the test sample x' with trigger q Test sample x with non-target label q The change in syntax errors, ΔGE(·) represents the test subset T' with triggers p Syntax error changing value.
[0081] The beneficial effects of the present invention are:
[0082] 1. The present invention can verify whether a program has a backdoor vulnerability by injecting a small amount of backdoor vulnerability test samples, and is superior to similar technologies.
[0083] 2. The present invention combines NLP model, CAM technology and backdoor vulnerability mining, which has important application value and research value for studying backdoor vulnerability detection and defense of text programs.
[0084] 3. This invention studies the problems existing in the current backdoor software security vulnerability detection methods in the field of natural language processing, and provides a new research idea for backdoor attack detection and defense of NLP models. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0086] Figure 1 It is a general flow chart of an embodiment of the present invention.
[0087] Figure 2 It is a process schematic diagram of an embodiment of the present invention.
[0088] Figure 3 It is a schematic diagram of the process of generating vulnerability test samples according to an embodiment of the present invention. DETAILED DESCRIPTION
[0089] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0090] Example 1
[0091] The embodiment of the present invention provides a software security vulnerability mining method for natural language processing, such as Figures 1 to 3 As shown, the following steps are included:
[0092] Step S1, preprocess the vulnerability code dataset samples, use the SST-2 dataset as the vulnerability code dataset, which contains independent training sets and test sets. Set label t as the target label, in the training set D to be preprocessed pre Extract the sample set labeled as the target label as the target label training subset to be preprocessed After preprocessing, we get the preprocessed target label training subset D t .in, represents the target label training subset to be preprocessed, N represents The number of samples in express The last sample in .
[0093] Since the samples currently obtained are English texts, the samples are converted into numerical values through the BERT (Bidirectional Encoder Representations from Transformers) model, and then normalized. Specifically:
[0094]
[0095] in, represents the target label sample to be preprocessed, x i represents the target label sample after preprocessing, express The mean of express The variance of i express The value corresponding to the conversion of the i-th word in, R represents The number of words in a R express The value to convert the last word in .
[0096] The target label training subset to be preprocessed Preprocessing is performed to obtain the preprocessed target label training subset D t , perform the same operation on other samples and finally obtain the preprocessed training set D.
[0097] Step S2, determine the backdoor trigger of the vulnerability code, and obtain the preprocessed target label training subset D based on step S1 t , extract the important feature set of each preprocessed target label sample through the Class Activation Mapping (CAM) method;
[0098] S21, obtain the features of a single sample. i Input into the neural network, extract the value and weight information of the last layer output in the neural network, and obtain the preprocessed target label sample x through weighted calculation i The eigenvalue of is as follows:
[0099]
[0100] Among them, CAM i Represents the preprocessed target label sample x i The characteristic value of represents the weight of the kth neuron of the last layer of the neural network for the target label t, C i,k represents the output of the kth neuron of the last layer of the neural network for the ith word, and A represents the number of neurons in the last layer of the neural network.
[0101] S22, CAM of the feature value of the obtained preprocessed target label sample i Filter and select the value greater than CAM i The feature set of the intermediate value is used as the important feature set F of the preprocessed target label sample i , and then select the words with the most occurrences from the important feature set as the feature words f of the preprocessed target label sample i , the specific formula is as follows:
[0102] x i ={w1,w2,…,w B}
[0103] M i =median(CAM i )
[0104] F i ={w j}where CAM i [j]>M i andw j ∈x i
[0105] f i =max(count(F i ))
[0106] Among them, x i represents the target label sample after preprocessing, median(·) represents the median of the obtained vector, CAM i Represents x i The eigenvalue of is a vector, M i Indicates CAM i The median of , F i Represents x i The important feature set, w j Represents x i The jth word in CAM i [j] indicates CAM i The jth value in count(F i ) indicates the i The number of occurrences of all words in the , max(·) means selecting the word with the largest number of occurrences, and B represents the preprocessed target label sample x i Quantity, wB Represents x i The last word of .
[0107] The preprocessed target label training subset D t Each sample in performs the operations of S21 to S22, and finally the feature words of all the preprocessed target label samples constitute the feature word set f Dt , as follows:
[0108]
[0109] in, Indicates D t The feature word set, f N express The last characteristic word of .
[0110] Step S3, generate a trigger for the vulnerability code. The feature word set of the target label training subset The occurrence counts of all words in are counted and sorted in descending order. The first m frequent words are selected as the trigger set of the SST-2 dataset on the target label t. The specific set is as follows:
[0111]
[0112] Among them, Tri t Represents the trigger set, Indicates trigger set Tri t The last word in the trigger set Tri t The number of trigger words in the .
[0113] Step S4, generate a training subset D' with triggers p .
[0114] S41, according to a certain modification ratio p, randomly select some samples of non-target labels in the training set D of the vulnerability code dataset as the training subset D to be triggered p , D p The non-target label training samples in are denoted as x p The specific calculation formula for the modified ratio p is:
[0115]
[0116] in, Indicates the training subset D to be triggered p The number of samples, N D Represents the number of samples in the training set D.
[0117] S42, select a replacement position. Perform feature extraction of step S2 on the non-target label training sample, and select the non-target label training sample with a feature value greater than x p The word corresponding to the feature of the median feature value is used as an important feature set F for non-target label samples p For the important feature set F p The occurrence counts of all words in are counted and sorted in descending order, and the first b frequent words are selected as non-target label training samples x p The feature word set is expressed as:
[0118] M p =median(CAM p )
[0119] F p ={w a}where CAM p [a]>M p
[0120] f p ={w p1 ,w p2 ,…,w pb}
[0121] Among them, CAM p Represents non-target label training sample x p The eigenvalue is a vector, median(·) means taking the median, M p Indicates CAM p The median of F p Represents non-target label training sample x p The important feature set, w a Represents non-target label training sample x p The ath word in CAM p [a] indicates CAM p The ath value in (i.e., the feature value of the ath word in the target label training sample), p b Represents non-target label training sample x p The last replacement position in f p Non-target label training sample x p The feature word set, w pb Represents non-target label training sample x p The word in the last replacement position in .
[0122] f p The word in the non-target label training sample x p The position is taken as the replacement position, recorded as Pos p ={p1,p2,…,p b}, the position may be continuous or discontinuous, where b represents the replacement position Pos p The number is a preset parameter and can be changed. In the present invention, b=20.
[0123] S43, based on the trigger set Tri on the target tag t obtained in step S3 t and the non-target label training sample x obtained in S42 p Replacement position Pos p , using Tri t The trigger words in iteratively replace the non-target label training samples x p Replacement position Pos p The corresponding words on the left are repeated until the specified perturbation amount is reached, and the first candidate vulnerability test sample is generated, denoted as x' p1 ; The selected triggers and replacement positions are randomly combined, and the perturbation amount is calculated as follows:
[0124] Pr(x p )=N tri / len(x p )
[0125] Among them, N tri represents the number of replacement positions that have been replaced, len(·) represents the length, Pr(·) represents the perturbation amount, x p Represents non-target label training samples.
[0126] S44, calculate the perplexity change value ΔPPL(x' p1 ,x p ), if ΔPPL(x' p1 ,x p ) exceeds the set threshold (the threshold is set to 50 in this embodiment), the replacement is canceled, and the replacement operation in step S43 is executed again, and other combinations are selected to replace the non-target label training samples to generate new candidate vulnerability test samples; if the candidate vulnerability test sample x' generated when the replacement in S43 is executed for the jth time pj Calculated The perplexity change value If the set threshold is not exceeded, the candidate vulnerability test sample generated for the jth time For the final vulnerability test sample x' p , and the final vulnerability test sample x' p The label is set to the target label t. Specifically, the candidate vulnerability test sample generated for the jth time The perplexity change value The calculation method is:
[0127]
[0128]
[0129] Among them, w i Represents candidate vulnerability test samples The i-th word in p(w i |w <i ) means that according to the context w <i ={w1,w2,…,w i-1 Prediction i The probability of Represents non-target label training sample x p The candidate vulnerability test samples generated when performing the jth replacement in S43, N p express The number of words in express The perplexity value, PPL(x p ) represents the non-target label training sample x p The confusion value, Represents the candidate vulnerability test sample generated for the jth time The perplexity change value.
[0130] S45: Training subset D to be triggered p Each sample in performs the operations of steps S42 to S44, and all the final vulnerability test samples generated constitute the training subset D' with triggers p .
[0131] Step S5, generating a BERT model for backdoor vulnerability code detection.
[0132] S51, training subset D' with triggers obtained based on S4 p , D' p Add to the training set D, and get the training set with trigger D'=D+D' p ;
[0133] S52, according to the environment setting, the initial setting learning rate lr = 2e-5 (i.e., 2 -5 ), training batch epoch = 10 and batch size batch_size = 8;
[0134] S53, inputting the training set D' with triggers generated in S51 into the BERT model F(·) for training, and obtaining a BERT model for backdoor vulnerability detection, denoted as F'(·);
[0135] Step S6, calculate the detection index of the vulnerability code and mine the backdoor vulnerability. Based on the BERT model F'(·) for backdoor vulnerability detection obtained in step S5, determine whether the detection F'(·) can classify and predict the test sample with the backdoor trigger as the target label.
[0136] S61, select samples in the test set T whose labels are not the target labels as the test subset T of non-target labels p , T p Execute step S4 for all samples in to obtain the test subset T' with trigger p ={x'1,x'2,…,x' Q}, where Q represents the test subset T' with triggers p The number of samples in x' Q Indicates T' p The last sample in .
[0137] S62, test subset T' with trigger p Put it into the BERT model F'(·) for backdoor vulnerability detection for prediction, and get the predicted label set of the test subset with triggers, as follows:
[0138]
[0139] Among them, T' p represents a test subset with triggers, Indicates T' p The predicted label set, Q represents T' p The number of samples, F'(·) represents the BERT model for backdoor vulnerability detection, represents the last sample x' of the test subset with trigger Q Predicted labels in the BERT model F'(·) for backdoor vulnerability detection;
[0140] S63, statistics of predicted label sets for the test subset with triggers All labels in are the number of target labels t, and the attack success rate ASR on the vulnerability code dataset and BERT model is calculated;
[0141]
[0142] Among them, x' q Represents a test subset T' with triggers p The qth sample in Represents x' q Predicted labels in the BERT model F'(·) for backdoor vulnerability detection, t represents the target label, Q represents the test subset T' with triggers pThe number of samples, I1(·) represents the calculation of the test subset T' with trigger p The predicted label of the qth sample in Is it a function of the target label? If is the target label, I1(·) is 1. If it is not a target label, I1(·) is 0; ASR represents the attack success rate. If ASR is greater than 30%, it means that the test subset T' with trigger p More than 30% of the test samples had incorrect label predictions, indicating that the model had a backdoor vulnerability.
[0143] S64, input the test set T into the BERT model F'(·) for backdoor vulnerability detection to obtain the predicted label set of the test set Determine the predicted label set for the test set The original label set of each label in the test set Whether the corresponding labels are consistent, the classification accuracy CACC of the BERT model for backdoor vulnerability detection in clean samples is calculated by counting the number of consistent labels. The calculation formula is as follows:
[0144]
[0145] in, represents the predicted label of the zth sample in the test set, represents the original label of the zth sample in the test set, H represents the number of samples in the test set T, represents the last predicted label of the prediction set, represents the last original label of the prediction set, CACC represents the classification accuracy of the model under clean samples, and I2(·) represents the function that calculates whether the predicted label and the original label of the zth sample in the test set are the same. If they are the same, I2(·) is 1, and if they are not the same, I2(·) is 0.
[0146] S65, Evaluate the test subset T' with triggers p The specific formula for the fluency change is as follows:
[0147]
[0148] Among them, x q is the test subset T of non-target labels p The qth non-target label test sample in q Represents a test subset T' with triggers p The qth non-target label test sample in , Q represents the test subset T' with trigger p is the number of samples in , and ΔPPL(·) represents the change value of fluency.
[0149] S66, calculate the test subset T of non-target labels p Non-target labeled test samples x in q , test subset T' with trigger p The test sample x' with trigger q Similarity, non-target label test sample x q and the test sample x' with trigger q The vector representation v is obtained through the USE (Universal Sentence Encoder) model. q and v' q , calculate the cosine similarity of the two vectors, and then take the average to get the test subset T' with triggers p The larger the USE value, the larger the test subset T' with triggers. p The lower the similarity with the original sample, the easier it is to detect the backdoor vulnerability of the software.
[0150] v q =F USE (x q )
[0152] v' q =F USE (x' q )
[0153]
[0154] Among them, F USE (·) represents the USE model, cos(·) represents the cosine similarity, and v q Represents the non-target label test sample x q By USE Model F USE (·) The vector representation obtained is v' q Represents a test sample x' with a trigger q By USE Model F USE (·) represents the vector obtained, ||·|| represents the length of the vector, and USE(·) represents the test subset T' with triggers p The cosine similarity of p represents the test subset of the SST-2 test set T whose labels are non-target labels, T' p represents a test subset with a trigger, Q represents T' p The number of samples in .
[0155] S67, using syntax error variation value to evaluate the test subset T' with trigger p First, we use the grammar checker to check the non-target label test sample xq and the test sample x' with trigger q The syntax errors are counted. Then the test sample x' with trigger is calculated. q Test sample x with non-target label q The change value of syntax errors is averaged to obtain the test subset T' with triggers p The syntax error change value is as follows:
[0157] ΔG(x' q ,x q )=G(x' q )-G(x q )
[0158]
[0159] Among them, x q The test subset T represents the non-target labels p The qth sample in T p represents the test subset of non-target labels in the SST dataset, T' p represents a test subset with triggers, x' q represents a test sample with a trigger, Q represents T' p The number of samples in G(x q ) represents the non-target label test sample x q The number of syntax errors, G(x' q ) represents the test sample x' with trigger q The number of grammatical errors, ΔG represents the non-target label test sample x q and the test sample x' with trigger q Syntax error change value, ΔG(·) represents the test sample x' with trigger q Test sample x with non-target label q The change in syntax errors, ΔGE(·) represents the test subset T' with triggers p Syntax error changing value.
[0160] The text quality of the test subset with triggers is evaluated by calculating the fluency change value, grammatical error change value, and similarity with the test subset with non-target labels of each detection method. The lower the text quality, the easier it is to defend against the software backdoor vulnerability detected by the test subset with triggers, and the more likely it is to cause the backdoor vulnerability detection to fail.
[0161] Experimental verification
[0162] In order to more clearly demonstrate the technical effect of the present invention, a series of experiments are conducted to verify its performance and compare it with existing methods.
[0163] In order to verify the effectiveness of the present invention, the present invention will perform software vulnerability detection on the three commonly used models BERT, RoBERTa, and DistilBERT models to detect whether the current model has vulnerabilities after training with a training set with a small number of vulnerability samples. The specific experimental results are shown in Table 2:
[0164] Table 2 Performance on different datasets and models
[0165]
[0166] It proves that the present invention detects that the three models have backdoor vulnerabilities after being trained using a training set with a small number of vulnerability tests, indicating that the robustness of these models is poor. The ASR of the SST-2 dataset on the DistilBERT model is relatively low, and it can be seen that the robustness of the DistilBERT model is better than other models. The performance on the SST-2 dataset is lower than that on the IMDB dataset. It can be seen that the backdoor vulnerability of SST-2 is smaller, which may be due to the relatively short text length of the SST-2 dataset, indicating that in actual scenarios, the use of short text training models is more secure. The present invention is compared with other backdoor methods, and the results are shown in Table 3:
[0167] Table 3 Performance of different methods on the SST-2 dataset
[0168]
[0169]
[0170] From the results in Table 2, it can be seen that the ASR index of the present invention is the highest, and the three indicators for evaluating the quality of backdoor sample generation have advantages over other methods. For example, the fluency of the generated samples in the ΔPPL index is higher than that of the three methods of RIPPLES, EP and BITE, and the USE index is higher than that of the SOS and StyleBkd methods. Among them, the StyleBkd method is better in the indicators of sample generation quality evaluation (i.e., ΔPPL, USE, ΔGE), but its results in the backdoor effectiveness evaluation indicators CACC and ASR are not good; while the EP method is the opposite. It shows that the current backdoor method is incompatible with the effectiveness of the backdoor and the generation quality of the backdoor samples. The method of the present invention takes into account the quality of the generated backdoor samples on the basis of achieving the effectiveness of the backdoor. The experimental results prove that the present invention can detect the existence of backdoor vulnerabilities in the NLP model under the condition of mixing in a small number of vulnerability test sample sets, and the backdoor samples have high effectiveness and concealment.
[0171] This paper proves that when a small amount of backdoor samples are injected, the model can also be implanted with a backdoor, and there is a backdoor vulnerability, especially for the currently popular large language model, a backdoor can be implanted when a small amount of backdoor samples are injected. This paper explores the defects of current text attacks. The high effectiveness and high concealment of text backdoor samples are incompatible, and the quality of text backdoor sample generation is poor. Future defense can be considered from the quality of sample generation, and attention should be paid to the defense effect under low poisoning rate.
[0172] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0173] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A software security vulnerability mining method for natural language processing, characterized in that: The following steps are involved: Step S1, preprocessing vulnerability code dataset samples; Step S2, determining the backdoor trigger of the vulnerability code; Step S3, generating a trigger for the vulnerability code; Step S4, generating a training subset with triggers; Step S5, generating a model for backdoor vulnerability code detection; Step S6, calculating the detection index of the vulnerability code and mining the backdoor vulnerability.
2. According to the method for mining software security vulnerabilities for natural language processing in claim 1, it is characterized in that: The vulnerability code dataset in step S1 includes an independent training set and a test set. First, label t is set as the target label. pre Extract the sample set labeled as the target label as the target label training subset to be preprocessed in, represents the target label training subset to be preprocessed, N represents The number of samples in express The last sample in ; then the sample is converted into a numerical value through the BERT model and normalized. Specifically: in, represents the target label sample to be preprocessed, x i represents the target label sample after preprocessing, express The mean of express The variance of i express The value corresponding to the conversion of the i-th word in, R represents The number of words in a R express The value converted from the last word in ; The target label training subset to be preprocessed Preprocessing is performed to obtain the preprocessed target label training subset D t , perform the same operation on other samples and finally obtain the preprocessed training set D.
3. The method for mining software security vulnerabilities for natural language processing according to claim 1, characterized in that: In step S2, the target label training subset D obtained after preprocessing in step S1 is t , extract the important feature set of each preprocessed target label sample through the class activation mapping method. The specific steps are as follows: S21, obtain the features of a single sample, and convert the preprocessed target label sample x i Input into the neural network, extract the value and weight information of the last layer output in the neural network, and obtain the preprocessed target label sample x through weighted calculation i The eigenvalue of is as follows: Among them, CAM i Represents the preprocessed target label sample x i The characteristic value of represents the weight of the kth neuron of the last layer of the neural network for the target label t, C i,k represents the output of the kth neuron of the last layer of the neural network for the i-th word, and A represents the number of neurons in the last layer of the neural network; S22, CAM of the feature value of the obtained preprocessed target label sample i Filter and select the value greater than CAM i The feature set of the intermediate value is used as the important feature set F of the preprocessed target label sample i , and then select the words with the most occurrences from the important feature set as the feature words f of the preprocessed target label sample i , the specific formula is as follows: x i ={w1,w2,…,w B } M i =median(CAM i ) F i ={w j }where CAM i [j]>M i and w j ∈x i f i =max(count(F i )) Among them, x i represents the target label sample after preprocessing, median(·) represents the median of the obtained vector, CAM i Represents x i The eigenvalue of is a vector, M i Indicates CAM i The median of , F i Represents x i The important feature set, w j Represents x i The jth word in CAM i [j] indicates CAM i The jth value in count(F i ) indicates the i The number of occurrences of all words in the , max(·) means selecting the word with the largest number of occurrences, and B represents the preprocessed target label sample x i Quantity, w B Represents x i The last word of The preprocessed target label training subset D t Each sample in performs the operations of S21 to S22, and finally the feature words of all the preprocessed target label samples constitute a feature word set The details are as follows: in, Indicates D t The feature word set, f N express The last characteristic word of .
4. The method for mining software security vulnerabilities for natural language processing according to claim 1, characterized in that: In step S3, the feature word set of the target label training subset is The number of occurrences of all words in are counted and sorted in descending order. The first m frequent words are selected as the trigger set of the vulnerability code dataset on the target label t. The specific set is as follows: Among them, Tri t Represents the trigger set, Indicates trigger set Tri t The last word in the trigger set Tri t The number of trigger words in the .
5. The method for mining software security vulnerabilities for natural language processing according to claim 1, characterized in that: The specific steps of step S4 are as follows: S41: According to the modification ratio p, randomly select some samples of non-target labels in the training set D of the vulnerability code dataset as the training subset D to be triggered p , D p The non-target label training samples in are denoted as x p , the specific calculation formula for the modified ratio p is: in, Indicates the training subset D to be triggered p The number of samples, N D Represents the number of samples in the training set D; S42: Select a replacement position, execute the operation of step S2 on the non-target label training sample, and select a non-target label training sample with a feature value greater than x p The word corresponding to the feature of the median feature value is used as an important feature set F for non-target label samples p , for the important feature set F p The occurrence counts of all words in are counted and sorted in descending order, and the first b frequent words are selected as non-target label training samples x p The feature word set is expressed as: M p =median(CAM p ) F p ={w a }where CAM p [a]>M p Among them, CAM p Represents non-target label training sample x p The eigenvalue is a vector, median(·) means taking the median, M p Indicates CAM p The median of , F p Represents non-target label training sample x p The important feature set, w a Represents non-target label training sample x p The ath word in CAM p [a] indicates CAM p The ath value in p b Represents non-target label training sample x p The last replacement position in f p Non-target label training sample x p The feature word set, Represents non-target label training sample x p The word at the last replacement position in ; f p The word in the non-target label training sample x p The position is taken as the replacement position, recorded as Pos p ={p1,p2,…,p b }, where b represents the replacement position Pos p the number of S43: Based on the trigger set Tri on the target tag t obtained in step S3 t and the non-target label training sample x obtained in S42 p Replacement position Pos p , using Tri t The trigger words in iteratively replace the non-target label training samples x p Replacement position Pos p The corresponding words on the previous page are generated until the specified perturbation amount is reached, and the first candidate vulnerability test sample is generated, which is recorded as The selected triggers and replacement positions are randomly combined, and the perturbation amount is calculated as follows: Pr(x p )=N tri / len(x p , Among them, N tri represents the number of replacement positions that have been replaced, len(·) represents the length, Pr(·) represents the perturbation amount, x p Represents non-target label training samples; S44: Calculate the perplexity change value of the first candidate vulnerability test sample When the set threshold is exceeded, the replacement is canceled, and the replacement operation in step S43 is performed again, and other combinations are selected to replace the non-target label training samples to generate new candidate vulnerability test samples; the jth replacement operation is performed, and the generated candidate vulnerability test samples The calculated perplexity change value If the set threshold is not exceeded, the candidate vulnerability test sample generated for the jth time For the final vulnerability test sample x' p , and the final vulnerability test sample x' p The label is set to the target label t. Specifically, the calculation formula is: Among them, w i Represents candidate vulnerability test samples The i-th word in p(w i |w <i ) means that according to the context w <i ={w1,w2,…,w i-1 Prediction i The probability of Represents non-target label training sample x p The candidate vulnerability test samples generated when performing the jth replacement in S43, N p express The number of words in express The perplexity value, PPL(x p ) represents the non-target label training sample x p The confusion value, Represents the candidate vulnerability test sample generated for the jth time The perplexity change value of S45: Training subset D to be triggered p Each sample in performs the operations of steps S42 to S44, and all the final vulnerability test samples generated constitute the training subset D' with triggers p .
6. The method for mining software security vulnerabilities for natural language processing according to claim 1, characterized in that: The specific steps of step S5 to generate a backdoor vulnerability code detection model are as follows: S51, the training subset D' with trigger obtained in S4 p Add to the training set D, and get the training set with trigger D'=D+D' p ; S52, initially setting a learning rate, a training batch, and a batch size according to the environment settings; S53, input the training set D' with triggers generated in S51 into the BERT model F(·) for training, and obtain the BERT model F'(·) for backdoor vulnerability detection.
7. The method for mining software security vulnerabilities for natural language processing according to claim 1, characterized in that: The specific steps of step S6 are as follows: S61, select samples in the test set T whose labels are not the target labels as the test subset T of non-target labels p , T p All samples in the test set perform the operation of step S4 to obtain the test subset T' with trigger p ={x'1,x'2,…,x' Q }, where Q represents the test subset T' with triggers p The number of samples in x' Q Indicates T' p The last sample in ; S62, test subset T' with trigger p Put it into the BERT model F'(·) for backdoor vulnerability detection for prediction, and get the predicted label set of the test subset with triggers, as follows: Among them, T' p represents a test subset with triggers, Indicates T' p The predicted label set, Q represents T' p The number of samples, F'(·) represents the BERT model for backdoor vulnerability detection, represents the last sample x' of the test subset with trigger Q Predicted labels in the BERT model F'(·) for backdoor vulnerability detection; S63, statistics of predicted label sets for the test subset with triggers All labels in the _ are the number of target labels t, and the attack success rate ASR on the vulnerability code dataset and BERT model is calculated. The specific formula is as follows: Among them, x' q Represents a test subset T' with triggers p The qth sample in Represents x' q Predicted labels in the BERT model F'(·) for backdoor vulnerability detection, t represents the target label, Q represents the test subset T with triggers ’ p The number of samples, I1(·) represents the calculation of the test subset T' with trigger p The predicted label of the qth sample in Whether it is a function of the target label, ASR represents the attack success rate; S64, input the test set T into the BERT model F'(·) for backdoor vulnerability detection to obtain the predicted label set of the test set Determine the predicted label set for the test set The original label set of each label in the test set Whether the corresponding labels are consistent, the classification accuracy CACC of the BERT model for backdoor vulnerability detection in clean samples is calculated by counting the number of consistent labels. The calculation formula is as follows: in, represents the predicted label of the zth sample in the test set, represents the original label of the zth sample in the test set, H represents the number of samples in the test set T, represents the last predicted label of the prediction set, represents the last original label of the prediction set, CACC represents the classification accuracy of the model under clean samples, and I2(·) represents the function that calculates whether the predicted label and the original label of the zth sample in the test set are the same; S65: Evaluate test subset T' with triggers p The specific formula for the fluency change is as follows: Among them, x q is the test subset T of non-target labels p The qth non-target label test sample in q Represents a test subset T with triggers ’ p The qth non-target label test sample in , Q represents the test subset T with trigger ’ p The number of samples in , ΔPPL(·) represents the fluency change value; S66: Calculate the test subset T of non-target labels p Non-target labeled test samples x in q , test subset T with trigger ’ p Test sample x with trigger ‘ q Similarity, non-target label test sample x q and a test sample x with a trigger ‘ q Through the USE model, we get the vector representation v q 、v' q , calculate the cosine similarity of the two vectors, and then take the average to get the test subset T with triggers ’ p The USE value of is as follows: v q =F USE (x q ) v' q =F USE (x‘ q ) Among them, F USE (·) represents the USE model, cos(·) represents the cosine similarity, and v q Represents the non-target label test sample x q By USE Model F USE (·) The vector representation obtained is v' q Represents a test sample x with a trigger ‘ q By USE Model F USE (·) represents the vector obtained, ||·|| represents the length of the vector, and USE(·) represents the test subset T with triggers ’ p The cosine similarity of p represents the test subset of the SST-2 test set T whose labels are non-target labels, T ’ p represents a test subset with a trigger, Q represents T ’ p The number of samples in ; S67: Using syntax error variation values to evaluate a test subset T with triggers ’ p The quality of , first use the grammar checker to check the non-target label test sample x q and a test sample x with a trigger ‘ q grammatical errors, count the grammatical errors that occur; then calculate the test sample x with triggers ‘ q Test sample x with non-target label q The change value of syntax errors is averaged to obtain the test subset T' with triggers p The syntax error change value is as follows: ΔG(x‘ q ,x q )=G(x' q )-G(x q ) Among them, x q The test subset T represents the non-target labels p The qth sample in T p represents the test subset of non-target labels in the SST dataset, T ’ p represents a test subset with triggers, x ‘ q represents a test sample with a trigger, Q represents T ’ p The number of samples in G(x q ) represents the non-target label test sample x q The number of syntax errors, G(x' q ) represents the test sample x with trigger ‘ q The number of grammatical errors, ΔG represents the non-target label test sample x q and a test sample x with a trigger ‘ q Syntax error change value, ΔG(·) represents the test sample x with trigger ‘ q Test sample x with non-target label q The change in syntax errors, ΔGE(·) represents the test subset T with triggers ’ p Syntax error changing value.
Citation Information
Cited By
Large model training method and device based on sample difficulty dynamic perception
CN121637058A