Method and apparatus for constructing natural and covert backdoor attacks using text features
By building a trigger candidate thesaurus and low-frequency word selection, and modifying the tags using black box conditions, the problem of weak trigger concealment in text backdoor attacks is solved, and a natural hidden backdoor attack with high success rate and low cost is achieved.
Patent Information
- Application Number
- CN202310734163.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-06-20
AI Technical Summary
The existing text backdoor attack methods have the problem of weak trigger concealment in the model training stage and the trigger stage, and the specific words inserted are easily detected, resulting in poor attack effect.
By constructing a trigger candidate thesaurus and low-frequency additional data sets, low-frequency words are selected as backdoor triggers, labels are modified under black box conditions, poisoned sample data sets are constructed and model training is performed, forming a natural hidden backdoor attack method.
It realizes that without modifying the text content, improves the attack success rate and concealment, reduces the false trigger rate and attack cost, and improves the stability and security of text backdoor attacks.
Smart Images

Figure CN116561587B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information security, and particularly relates to a method for constructing a natural and hidden backdoor attack by using text features, and also relates to a device for constructing a natural and hidden backdoor attack by using text features. Background Art
[0002] In recent years, text pre-training models have been widely applied in various fields based on natural language processing and have become one of the focuses of research by computer scientists around the world. However, the application of pre-training models is inseparable from information security. Due to the large scale of pre-training models, they are usually outsourced to a third party for training or downloaded from a third party for application during use. If the third party is untrusted, there will be significant security risks. Therefore, the security issues of pre-training models have also received extensive attention from researchers.
[0003] Backdoor attack is an emerging security threat against artificial intelligence models. Backdoor attacks usually inject triggers into victim models during the model training process, so that the victim models can work normally when facing normal text inputs during the test stage. However, if the input contains pre-designed trigger features, the victim models will trigger the backdoor and output specific results. In practical applications, if a malicious third party injects a backdoor into a face recognition system, it can correctly recognize general faces during daily use. However, when encountering a face wearing a preset color glasses (trigger), regardless of which person the face wearing the glasses actually corresponds to, the victim model will identify it as a specific person.
[0004] Since there is not much difference in the performance of the backdoor-injected model and the normal model when facing normal inputs without trigger features, it is very difficult for model users to realize the existence of the backdoor, which makes backdoor attacks highly concealed and harmful.
[0005] By studying text backdoor attack techniques, security vulnerabilities of natural language processing models can be deeply explored, and the security and robustness of natural language processing models can be further improved to reduce the risks of natural language processing-based models being put into practical applications.
[0006] Current text backdoor attack methods mainly use a certain specific rare word inserted additionally as a trigger. Although these methods have achieved a relatively high success rate of backdoor attacks, their concealment is poor. The inserted words will significantly damage the grammar and fluency of the original text, and these words are easily detected by various methods, resulting in the failure of the attack. As a result, it is difficult to guarantee the attack effect on text backdoor attack models and accurately discover the weaknesses of the models. Summary of the Invention
[0007] The first object of the present invention is to provide a method for constructing a natural and concealed backdoor attack using text features, which solves the problem of weak concealment of triggers in the model training stage and trigger stage of existing text backdoor attacks.
[0008] The second object of the present invention is to provide a device for constructing a natural and concealed backdoor attack using text features.
[0009] The first technical solution adopted by the present invention is a method for constructing a natural and concealed backdoor attack using text features, specifically as follows:
[0010] Step 1: Use the clean training dataset D c to train the pre-trained model to obtain a clean model F c , where the corresponding class label of the text in D c is y;
[0011] Step 2: Extract features from the clean training dataset D c to construct a trigger candidate word library C w ;
[0012] Step 3: Use 3 - 5 alternative datasets different from the clean training dataset D c to construct a low-frequency additional dataset word library E w ;
[0013] Step 4: Sort the trigger candidate word library C w through the low-frequency additional dataset word library E w , select the 3 - 5 words with the lowest word frequencies and randomly select one word from these 3 - 5 words as the backdoor trigger W t ;
[0014] Step 5: Under the black-box condition, use the backdoor trigger W t to construct a text poisoned sample dataset D p ;
[0015] Step 6: Construct a non-target class enhanced dataset D e ;
[0016] Step 7: Obtain a trained backdoor model;
[0017] Step 8: Construct a backdoor text sample test set and input it into the trained backdoor model in Step 7 to obtain the model backdoor trigger result.
[0018] The feature of the present invention also lies in that
[0019] Step 2 is specifically implemented according to the following steps:
[0020] Step 2.1: Set the target class label as x, and form the dataset D from all the data in the clean training dataset D where the label y ≠ x. c Extract the text features of D. The text feature extraction method is calculated using the inverse document frequency. c_nontarget Extract the text features of D. c_nontarget The text feature extraction method is calculated using the inverse document frequency. The inverse document frequency. The formula expression of the inverse document frequency is:
[0021] Step 2.2: Since it is related to the poisoning ratio, therefore, select the words from D, and sort them in ascending order. These words construct the trigger candidate library C. Since it is related to the poisoning ratio, therefore, Select the words from D. c_nontarget Select the words from D. And sort them in ascending order. These words construct the trigger candidate library C. And sort them in ascending order. These words construct the trigger candidate library C. w .
[0022] Step 3 is specifically implemented according to the following steps:
[0023] Use 3 - 5 alternative datasets different from the clean training dataset D. Collect all the sentences in all the alternative datasets, calculate the frequency of the words appearing in all the sentences, c Use 3 - 5 alternative datasets different from the clean training dataset D. Collect all the sentences in all the alternative datasets, calculate the frequency of the words appearing in all the sentences, And sort them in descending order according to the frequency. The frequency calculation formula is: The frequency calculation formula is: All the words in the text are sorted in ascending order according to the value of the word, forming the low - frequency additional dataset word library E. All the words in the text are sorted in ascending order according to the value of the word, forming the low - frequency additional dataset word library E. w .
[0024] Step 4 is specifically implemented according to the following steps:
[0025] Step 4.1: Calculate the trigger scores of all the words in the trigger candidate library C. Query the words in C in the low - frequency additional dataset word library E and find the corresponding values of the words. Calculate the trigger scores of each word in the trigger candidate library C. w Calculate the trigger scores of all the words in the trigger candidate library C. Query the words in C in the low - frequency additional dataset word library E and find the corresponding values of the words. Calculate the trigger scores of each word in the trigger candidate library C. w Query the words in C in the low - frequency additional dataset word library E and find the corresponding values of the words. Calculate the trigger scores of each word in the trigger candidate library C. w Query the words in C in the low - frequency additional dataset word library E and find the corresponding values of the words. Calculate the trigger scores of each word in the trigger candidate library C. Calculate the trigger scores of each word in the trigger candidate library C. w Calculate the trigger scores of each word in the trigger candidate library C. Sort them in ascending order according to the trigger scores to get the re - sorted C′. w Sort them in ascending order according to the trigger scores to get the re - sorted C′. The calculation formula of each word's trigger score is:
[0026] Step 4.2: Select the first 3 - 5 words from the re - sorted C′, and randomly select one from these 3 - 5 words as the final backdoor trigger W. w Select the first 3 - 5 words from the re - sorted C′, and randomly select one from these 3 - 5 words as the final backdoor trigger W. t .
[0027] Step 5 is specifically implemented according to the following steps:
[0028] Under the black-box condition, search in D c_nontarget When the sentence in D c_nontarget contains the backdoor trigger W t Change the label of the sentence to the target class label x, and the modified dataset is the poisoned sample dataset D p ; The text poisoned sample dataset D p is obtained by modifying the labels of the sentences containing the text trigger words in the clean training dataset D c .
[0029] Step 6 is specifically implemented according to the following steps:
[0030] Since constructing the poisoned text dataset D p will cause a reduction in the non-target class dataset, which will affect the performance of the model. Split the sentences in D c_nontarget that do not contain the trigger into different clauses, and mark the labels of the clauses as the same as the original sentence to form the non-target class augmented dataset D e ;
[0031] The construction of the clauses in Step 6 is specifically as follows: Split a sentence by commas, and the split sentences can form different clauses.
[0032] Step 7 is specifically implemented according to the following steps:
[0033] Combine the poisoned sample dataset D p , the non-target class augmented dataset D e and the unchanged dataset in the clean training dataset D c as the training set of the backdoor model to retrain the clean model F c to obtain the trained backdoor model;
[0034] The construction of the training set of the backdoor model in Step 7 is specifically as follows: For each additional poisoned sample, select a clause with the same label as before modification from D e and add it to the training set of the backdoor model for model training together with other unchanged datasets.
[0035] Step 8 is specifically implemented according to the following steps:
[0036] The construction of the backdoor text sample test dataset in Step 8 is specifically as follows:
[0037] S1: Use the text that does not contain the trigger W t as the normal sample test set;
[0038] S2: To ensure the concealment of the triggering stage, trigger W is used t to form a poisoned sample test set by sentence construction;
[0039] S3: Combine the normal sample test set with the poisoned sample test set to form a backdoor text sample test set. The normal sample test set can be correctly classified in the backdoor model, and the label of the backdoor sample test set can only be the target class label x.
[0040] The second technical solution adopted by the present invention is to construct a natural and concealed backdoor attack device using text features, including:
[0041] A pre-trained model training module that uses the clean training dataset D c to train the pre-trained model to obtain a clean model F c , where the corresponding class label of the text in D c is y;
[0042] A feature extraction module for extracting features from the clean training dataset D c to construct a trigger candidate word library C w ;
[0043] A low-frequency additional dataset word library construction module that uses 3-5 alternative datasets different from the clean training dataset D c to construct a low-frequency additional dataset word library E w ;
[0044] A sorting module that sorts the trigger candidate word library C w through the low-frequency additional dataset word library E w to select the 3-5 words with the lowest word frequencies and randomly select one word from these 3-5 words as the backdoor trigger W t ;
[0045] A text poisoned sample dataset construction module for constructing a text poisoned sample dataset D t under black box conditions using the backdoor trigger W p ;
[0046] A non-target class enhancement dataset construction module for constructing a non-target class enhancement dataset D e ;
[0047] A backdoor model training module for obtaining a trained backdoor model;
[0048] A backdoor trigger result module for constructing a backdoor text sample test set and inputting it into the trained backdoor model to obtain the model backdoor trigger result.
[0049] The beneficial effects of the present invention are:
[0050] (1) The key point of the method of the present invention is to use the words contained in the target dataset as triggers. Different from other methods that require inserting text, it does not need to modify the text at all, can completely hide the backdoor trigger while having a high attack success rate; it solves the problem of weak concealment of the trigger existing in the existing text backdoor attack in the model training stage and the trigger stage; the method of the present invention aims to optimize the current text backdoor attack based on the pre-trained model, while reducing the attack cost, improving the attack success rate of the text backdoor attack.
[0051] (2) The innovation of the method of the present invention also lies in being able to reasonably analyze the text features, using the low-frequency features of words to design a trigger selection method. Different from the high false trigger caused by inserting sentence-level triggers, while ensuring the attack effect, the selected trigger can effectively prevent false triggers.
[0052] (3) The key point of the method of the present invention also lies in the relatively low attack cost. Different from other methods that require optimizing the model, the present invention only needs to modify the label after finding the word used as the trigger to complete the attack on the pre-trained model, reducing the attack cost.
[0053] (4) Compared with the prior art, the method of the present invention uses natural text as the trigger, without modifying the text content, and has the advantages of strong concealment, low false trigger rate, low attack cost and good attack effect.
[0054] (5) The purpose of the method of the present invention is to use the characteristics of the text itself to find a low-frequency word in the original dataset, and complete the backdoor injection without making any modifications to the text itself, thereby improving the stability of the text backdoor attack. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flowchart of the method for constructing a natural and concealed backdoor attack using text features according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0056] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0057] The present invention provides a method for constructing a natural and concealed backdoor attack using text features, as Figure 1 shown, including the following specific steps:
[0058] Step 1: Use the clean training dataset D c to train the pre-trained model to obtain a clean model F c , where the corresponding class label of the text in D c is y;
[0059] Step 2: Perform feature extraction on the clean training dataset D c to construct a trigger candidate word library C w ;
[0060] Step 2 is specifically implemented according to the following steps:
[0061] Step 2.1: Set the target class label as x, and form a dataset D from all the data in the clean training dataset D c where the label y ≠ x. Extract the text features of D c_nontarget . The text feature extraction method is to use the inverse document frequency c_nontarget for calculation. The formula expression of the inverse document frequency is:
[0062] Step 2.2: Since is related to the poisoning ratio, therefore select words from D c_nontarget and sort them in ascending order. These words are used to construct the trigger candidate word library C . w .
[0063] Step 3: Use 3 - 5 alternative datasets different from the clean training dataset D c to construct a low - frequency additional dataset word library E w ;
[0064] Step 3 is specifically implemented according to the following steps:
[0065] Use 3 - 5 alternative datasets different from the clean training dataset D c . Collect all the sentences in all the alternative datasets together, calculate the frequency of the words appearing in all the sentences and sort them in descending order of frequency. The frequency calculation formula is: All the words in the text are sorted in ascending order according to the value to form the low - frequency additional dataset word library E w .
[0066] Step 4: Sort the trigger candidate word library C w through the low - frequency additional dataset word library E w , select the 3 - 5 words with the lowest word frequencies and randomly select one word from these 3 - 5 words as the backdoor trigger W t ;
[0067] Step 4 is specifically implemented according to the following steps:
[0068] Step 4.1: Calculate the trigger scores of all words in the trigger candidate word library C w in C w query the words in the low-frequency additional dataset word library E w and find the corresponding tf wi value of the word, and calculate the trigger score of each word in the trigger candidate word library C w Sort the trigger scores from low to high to obtain the re-sorted C′ The trigger score calculation formula for each word is: w The formula is:
[0069] Step 4.2: Select the first 3 - 5 words in the re-sorted C′ w and randomly select one of these 3 - 5 words as the final backdoor trigger W t .
[0070] Step 5: Under the black-box condition, use the backdoor trigger W t to construct the text poisoned sample dataset D p ;
[0071] Step 5 is specifically implemented as follows: Under the black-box condition, search in D c_nontarget When the sentence in D c_nontarget contains the backdoor trigger W t , modify the label of the sentence to the target class label x, and the modified dataset is the poisoned sample dataset D p ; The text poisoned sample dataset D p is obtained by modifying the labels of the sentences containing the text trigger words in the clean training dataset D c ;
[0072] Step 6: Construct the non-target class augmented dataset D e , Since constructing the poisoned text dataset D p will cause a reduction in the non-target class dataset and thus affect the model's performance, split the sentences in D c_nontarget that do not contain the trigger into different clauses, and mark the labels of the clauses as the same as the original sentence to form the non-target class augmented dataset D e ;
[0073] The construction of the clauses in Step 6 is specifically as follows: Split a sentence by commas, and the split sentences can form different clauses.
[0074] Step 7: Combine the poisoned sample dataset D p , the non-target class augmented dataset D e and the clean training dataset Dc Combine the unmodified data sets in it as the training set of the backdoor model to retrain the clean model F c to obtain a trained backdoor model;
[0075] In step 7, the construction of the training set of the backdoor model is specifically as follows: for each additional poisoned sample, select a clause from D e that is consistent with the label before modification and add it to the training set of the backdoor model for model training together with other unmodified data sets.
[0076] Step 8: Construct a backdoor text sample test set and input it into the trained backdoor model in step 7 to obtain the model backdoor trigger result. The backdoor text sample test set includes a normal sample test set and a poisoned sample test set.
[0077] The construction of the backdoor text sample test data set in step 8 is specifically as follows:
[0078] S1: Use the text that does not contain the trigger W t as the normal sample test set;
[0079] S2: To ensure the concealment of the triggering stage, use the trigger W t to construct sentences to form a poisoned sample test set;
[0080] S3: Combine the normal sample test set and the poisoned sample test set to form a backdoor text sample test set. The normal sample test set can be correctly classified in the backdoor model, and the label of the backdoor sample test set can only be the target class label x.
[0081] Example 1
[0082] A method for constructing a natural and concealed backdoor attack using text features includes the following specific steps:
[0083] Step 1: Use the clean training data set D c to train the pre-trained model to obtain a clean model F c , where the corresponding class label of the text in D c is y; among them, the clean training data set D c is the IMDb data set;
[0084] Step 2: Extract features from the clean training data set D c to construct a trigger candidate word library C w ;
[0085] Step 2 is specifically implemented according to the following steps:
[0086] Step 2.1: Set the target class label as x, and use the clean training data set D cAll the data with tags y≠x in it form the dataset D c_nontarget , extract D c_nontarget 's text features. The text feature extraction method is to use the inverse document frequency DF wi to calculate. The formula expression of the inverse document frequency is:
[0087] Step 2.2: Since is related to the poisoning ratio, therefore select from D c_nontarget the words and sort them in ascending order according to . These words construct the trigger candidate word library C . w .
[0088] Step 3: Use 5 alternative datasets different from the clean training dataset D c to construct the low-frequency additional dataset word library E w ;
[0089] Step 3 is specifically implemented according to the following steps:
[0090] Use 5 alternative datasets different from the clean training dataset D c to collect all the sentences in all alternative datasets together and calculate the frequency of the words appearing in all the sentences and sort them in descending order according to the frequency. The frequency calculation formula is: All the words in the text are sorted in ascending order according to the value of the word to form the low-frequency additional dataset word library E w .
[0091] Step 4: Sort the trigger candidate word library C w through the low-frequency additional dataset word library E w , select the 3 words with the lowest word frequency and randomly select one word from these 3 words as the backdoor trigger W t ;
[0092] Step 4 is specifically implemented according to the following steps:
[0093] Step 4.1: Calculate the trigger scores of all the words in the trigger candidate word library C w . Query the words in C w in the low-frequency additional dataset word library E w and find the corresponding value of the word, and calculate the trigger scores of each word in the trigger candidate word library C w Sort C' after reordering in ascending order of trigger scores w , the trigger score of each word The calculation formula is:
[0094] Step 4.2: Select the first 3 - 5 words in the re - ordered C' w and randomly select one of these 3 - 5 words as the final backdoor trigger W t .
[0095] Step 5: Under the black - box condition, use the backdoor trigger W t to construct a text - poisoned sample dataset D p ;
[0096] Step 5 is specifically implemented as follows: Under the black - box condition, search in D c_nontarget . When the sentence in D c_nontarget contains the backdoor trigger W t , modify the label of the sentence to the target class label x, and the modified dataset is the poisoned sample dataset D p ; The text - poisoned sample dataset D p is obtained by modifying the labels of the sentences containing the text trigger words in the clean training dataset D c .
[0097] Step 6: Construct a non - target class augmented dataset D e . Since constructing the poisoned text dataset D p will cause a reduction in the non - target class dataset, which affects the model's performance, split the sentences in D c_nontarget that do not contain the trigger into different clauses, and mark the labels of the clauses as the same as the original sentence to form the non - target class augmented dataset D e ;
[0098] The construction of clauses in Step 6 is specifically as follows: Split a sentence by commas, and the split sentences can form different clauses.
[0099] Step 7: Combine the poisoned sample dataset D p , the non - target class augmented dataset D e and the unchanged dataset in the clean training dataset D c as the training set of the backdoor model to retrain the clean model F c to obtain the trained backdoor model;
[0100] The construction of the training set of the backdoor model in Step 7 is specifically as follows: For each additional poisoned sample, from D eSelect a clause that is consistent with the label before modification and add it to the backdoor model training set for model training together with other unmodified data sets.
[0101] Step 8: Construct a backdoor text sample test set and input it into the trained backdoor model in Step 7 to obtain the model backdoor trigger result. The backdoor text sample test set includes a normal sample test set and a poisoned sample test set.
[0102] The construction of the backdoor text sample test data set in Step 8 is specifically as follows:
[0103] S1: Use the text that does not contain the trigger W t as the normal sample test set;
[0104] S2: To ensure the concealment of the trigger phase, use the trigger W t to construct sentences to form the poisoned sample test set;
[0105] S3: Combine the normal sample test set and the poisoned sample test set to form the backdoor text sample test set. The normal sample test set can be correctly classified in the backdoor model, and the label of the backdoor sample test set can only be the target class label x.
[0106] Example 2
[0107] Except for selecting the SST-2 data set as the clean training data set D in Step 1, the other steps are the same as those in Example 1; c
[0108] After the above steps, the attack effects of Examples 1 and 2 are obtained as shown in Table 2. Table 2 shows the attack effects of the IMDb and SST-2 datasets on the BERT (Bidirectional Encoder Representation from Transformers) model, where both SST-2 and IMDb are sentiment analysis datasets. It can be seen from Table 2 that under the black-box condition, for the IMDb dataset, the method of the present invention can achieve an attack success rate of 97.83%, which is higher than the attack success rate of 96.34% of the baseline method; for the SST-2 dataset, the method of the present invention can achieve an attack success rate of 100%, which is higher than the attack success rate of 95.5% of the baseline method; among them, the baseline method is "Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, and Maosong Sun. Hidden killer: Invisible textual backdoor attacks with syntactic trigger. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL / IJCNLP 2021, (Volume 1: Long Papers), Virtual 376 Event, August 1-6, 2021, pages 443–453. Association for Computational Linguistics, 2021b. 377 doi:10.18653 / v1 / 2021.acl-long.37. URL https: / / doi.org / 10.18653 / v1 / 2021.acl-long.378 37."
[0109] Table 2 Comparison of attack effects of the method of the present invention and the baseline method on different datasets
[0110]
[0111] Table 3 shows the effectiveness of the attack method proposed in the present invention in terms of stealth. By comparing three indicators, namely PPL, GE, and UES, with the baseline method, where PPL represents the fluency of the text and GE represents the grammar errors in the text. The lower these two indicators are, the smaller the impact of the trigger on the text and the stronger the stealth. UES is the similarity between the text after adding the trigger and the original text. The larger the value, the higher the similarity and the stronger the stealth of the trigger. It can be seen from Table 3 that the attack method proposed in the present invention has strong stealth and is superior to the baseline method. Thus, it can be seen that the method of the present invention solves the problem of weak stealth of the trigger existing in the existing text backdoor attack in the model training stage and the trigger stage.
[0112] Table 3 Comparison of the stealth effects of the trigger of the method of the present invention and the baseline method for different data sets
[0113]
[0114]
[0115] Example 3
[0116] Construct a natural stealth backdoor attack device using text features, including:
[0117] Pre-trained model training module, using the clean training data set D c Train the pre-trained model to obtain a clean model F c , where the corresponding class label of the text in D c is y;
[0118] Feature extraction module, used to extract features from the clean training data set D c to construct a trigger candidate word library C w ;
[0119] Low-frequency additional data set word library construction module, using 3-5 alternative data sets different from the clean training data set D c to construct a low-frequency additional data set word library E w ;
[0120] Sorting module, through the low-frequency additional data set word library E w sort the trigger candidate word library C w select the 3-5 words with the lowest word frequency and randomly select one word from these 3-5 words as the backdoor trigger W t ;
[0121] Text poisoned sample data set construction module, used to construct a text poisoned sample data set D t under the black box condition using the backdoor trigger W p ;
[0122] Construct a non-target class augmented dataset module for constructing a non-target class augmented dataset D e ;
[0123] Backdoor model training module for obtaining a trained backdoor model;
[0124] Backdoor trigger result module for constructing a backdoor text sample test set and inputting it into the trained backdoor model to obtain the model backdoor trigger result.
Claims
1. A method for constructing a natural and covert backdoor attack using text features, characterized in that Specifically: Step 1: Use the clean training dataset D c to train the pre-trained model and obtain the clean model F c , where the corresponding class label of the text in D c is y; Step 2: Perform feature extraction on the clean training dataset D c to construct a trigger candidate word library C w ; Step 2 is specifically implemented according to the following steps: Step 2.1: Set the target class label as x, and form the dataset D from all the data in the clean training dataset D c where the label y ≠ x, and extract the text features of D c_nontarget The text features are extracted by using the inverse document frequency c_nontarget for calculation. The formula expression of the inverse document frequency is as follows: Step 2.2: Select from D c_nontarget the words, and sort them in ascending order according to . These words construct the trigger candidate word library C ; w ; Step 3: Use 3 - 5 alternative datasets different from the clean training dataset D c in the field to construct a low - frequency additional dataset thesaurus E w ; Step 4: Through the low-frequency additional dataset word library E w Sort the trigger candidate word library C w Select the 3-5 words with the lowest word frequencies and randomly select one word from these 3-5 words as the backdoor trigger W t ; Step 5: Under the black-box condition, use the backdoor trigger W t Construct the text poisoned sample dataset D p ; Step 6: Construct the non-target class augmented dataset D e ; Step 7: Obtain the trained backdoor model; Step 8: Construct a backdoor text sample test set and input it into the trained backdoor model in Step 7 to obtain the model backdoor trigger result; The backdoor attack method is used to deeply explore the security vulnerabilities of natural language processing models and further improve the security and robustness of natural language processing models.
2. The method for constructing a natural and concealed backdoor attack using text features according to claim 1, wherein Step 3 is specifically implemented according to the following steps: Use 3 - 5 alternative datasets different from the clean training dataset D c In the field, collect the sentences in all alternative datasets together, and calculate the frequency of words appearing in all sentences And sort them in descending order of frequency. The frequency The calculation formula is: All words in the text are sorted from low to high according to the value to form the low-frequency additional dataset word library E w .
3. The method for constructing a natural and hidden backdoor attack using text features according to claim 2, wherein Step 4 is specifically implemented according to the following steps: Step 4.1: Calculate the trigger score of all words in the trigger candidate dictionary C w Query the words in C w in the low-frequency additional dataset dictionary E w and find the corresponding value. Calculate the trigger score of each word in the trigger candidate dictionary C w Sort C in ascending order of trigger score to obtain the re-sorted C' , and the trigger score of each word w is calculated as follows: The calculation formula is: Step 4.2: Select the first 3-5 words in the re-ordered C′ w and randomly select one of these 3-5 words as the final backdoor trigger W t .
4. The method for constructing a natural and concealed backdoor attack using text features according to claim 1, wherein Step 5 is specifically implemented according to the following steps: Under the black-box condition, search in D c_nontarget When the sentence in D c_nontarget contains the backdoor trigger W t change the label of the sentence to the target class label x, and the modified dataset is the poisoned sample dataset D p ; The text poisoned sample dataset D p is obtained by modifying the labels of the sentences containing the text trigger words in the clean training dataset D c .
5. The method for constructing a natural and concealed backdoor attack using text features according to claim 1, wherein Step 6 is specifically implemented according to the following steps: Split the sentences in D c_nontarget that do not contain triggers into different clauses, and mark the labels of the clauses as the same as the original sentences to form a non-target class enhanced dataset D e ; The construction of clauses in Step 6 is specifically as follows: Split a sentence by commas, and the split sentences can form different clauses.
6. The method for constructing a natural and concealed backdoor attack using text features according to claim 1, wherein Step 7 is specifically implemented according to the following steps: Combine the poisoned sample dataset D p , the non-target class augmented dataset D e with the unchanged dataset in the clean training dataset D c and use it as the training set of the backdoor model to retrain the clean model F c to obtain the trained backdoor model; The construction of the training set for the backdoor model in step 7 is specifically as follows: for each additional poisoned sample, select a clause from D e that is consistent with the label before modification and add it to the training set of the backdoor model for model training together with other unmodified data sets.
7. The method for constructing a natural and concealed backdoor attack using text features according to claim 6, wherein Step 8 is specifically implemented according to the following steps: The construction of the backdoor text sample test data set in Step 8 is specifically as follows: S1: Use the text that does not contain the trigger W t as the normal sample test set; S2: Use trigger W t to construct a poisoned sample test set by forming sentences; S3: Combine the normal sample test set and the poisoned sample test set to form the backdoor text sample test set.
8. A backdoor attack device for implementing the backdoor attack method of constructing a natural and concealed backdoor attack using text features as described in claim 1, characterized in that, Including: Pre-trained model training module, using the clean training dataset D c Train the pre-trained model to obtain the clean model F c , where D c The corresponding class label of the text is y; A feature extraction module for performing feature extraction on the clean training dataset D c to construct a trigger candidate word library C w ; Low-frequency additional dataset thesaurus construction module, using 3-5 alternative datasets different from the clean training dataset D c in the domain to construct the low-frequency additional dataset thesaurus E w ; Sorting module, through the low-frequency additional dataset thesaurus E w Sort the trigger candidate thesaurus C w Select the 3-5 words with the lowest word frequencies and randomly select one word from these 3-5 words as the backdoor trigger W t ; A text poisoned sample dataset construction module, which is used to construct a text poisoned sample dataset D under the black box condition using the backdoor trigger W t p ; Construct a non-target class augmented dataset module for constructing a non-target class augmented dataset D e ; The backdoor model training module is used to obtain the trained backdoor model; The backdoor trigger result module is used to construct the backdoor text sample test set and input it into the trained backdoor model to obtain the model backdoor trigger result.
Citation Information
Patent Citations
Label-consistent text backdoor attack method
CN113946687A
Method for performing text backdoor attack by using punctuations
CN114936594A