Text backdoor defense method and system based on SHAP values
By using the SHAP values method to analyze and process the feature words in text sentences, the existing text backdoor defense methods have solved the problem of detecting and defending hidden backdoor attacks and having a great impact on clean samples, and efficient text backdoor defense and semantic integrity reservation are achieved.
Patent Information
- Application Number
- CN202510184818.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-16
AI Technical Summary
The existing text backdoor defense methods have problems such as detecting backdoor trigger algorithms, difficulty in defending against hidden sentence-level backdoor attacks, and a greater impact on the classification prediction of clean samples.
Using the SHAP values method, the SHAP values of each characteristic word in the pending sentence are obtained through the SHAP interpreter, and the suspicious words are accurately positioned, and based on the comparison results of the SHAP values value and the preset threshold, the suspicious words are deleted or replaced, and a new sentence is obtained.
Accurate detection and defense of text backdoor attacks is achieved, reducing the impact on clean samples, and retaining the semantic integrity of the original samples.
Smart Images

Figure CN120012762A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data security technology, and in particular to a text backdoor defense method and system based on SHAP values. Background Art
[0002] With the rapid development of machine learning applications in natural language processing, backdoor attacks remain a key issue in computer security, with significant research potential. In natural language processing, backdoor attacks pose a serious threat to model and dataset security. Current research focuses on detecting and addressing backdoors, and on preventing them by eliminating them.
[0003] As backdoor attacks in the text field have gradually gained attention, methods for defending against text backdoors have been developed. However, these methods still have many flaws. For example, the current ONION method has the following flaws:
[0004] (1) The algorithm for detecting backdoor triggers has limitations and cannot accurately locate the backdoor triggers implanted by attackers, resulting in large errors;
[0005] (2) Since this method relies on using GPT-2 to detect outlier words to find backdoors, it is difficult to defend against more subtle sentence-level backdoor attacks;
[0006] (3) Existing methods generally use deletion methods to remove backdoors, which causes excessive damage to clean samples and seriously affects the model's classification prediction of clean samples. Summary of the Invention
[0007] In response to the problems existing in the prior art, the present invention provides a text backdoor defense method and system based on SHAPvalues.
[0008] The present invention provides a text backdoor defense method based on SHAP values, comprising:
[0009] Use the SHAP interpreter to obtain the SHAP values of each feature word in the sentence to be processed;
[0010] Taking the first preset number of feature words with the largest SHAP values in the sentence to be processed as suspect words;
[0011] The SHAP values of the suspected words are compared with a preset threshold, and the suspected words are deleted or replaced according to the comparison result to obtain a new sentence.
[0012] According to a text backdoor defense method based on SHAP values provided by the present invention, before using a SHAP interpreter to obtain the SHAP values of each feature word in a sentence to be processed, the method further includes:
[0013] Preprocessing the original sentence sample, wherein the preprocessing includes removing illegal characters in the original sentence sample;
[0014] Use the poisoning model to perform classification prediction on the preprocessed original sentence sample to obtain the predicted category and prediction vector of the original sentence sample;
[0015] The SHAP interpretable is constructed according to the predicted category and the predicted vector of the original sentence sample.
[0016] According to a text backdoor defense method based on SHAP values provided by the present invention, the first preset number when defending against word-level backdoor attacks on the sentence to be processed is greater than the first preset number when defending against sentence-level backdoor attacks on the sentence to be processed.
[0017] According to a text backdoor defense method based on SHAP values provided by the present invention, when defending the sentence to be processed against word-level backdoor attacks, the SHAP values of the suspected words are compared with a preset threshold, and the suspected words are deleted or replaced according to the comparison result to obtain a new sentence, including:
[0018] When the SHAP values of the suspected word are greater than the preset threshold, performing word replacement on the suspected word;
[0019] When the SHAP values of the suspected word are less than or equal to the preset threshold, the suspected word is deleted.
[0020] According to a text backdoor defense method based on SHAP values provided by the present invention, when defending the sentence-level backdoor attack on the sentence to be processed, the SHAP values of the suspected words are compared with a preset threshold, and the suspected words are deleted or replaced according to the comparison result to obtain a new sentence, including:
[0021] According to the SHAP values of the suspicious words in the sentence to be processed, the suspicious words are divided into two groups, wherein the SHAP values of the suspicious words in one group are greater than the SHAP values of the other group of suspicious words;
[0022] If the SHAP values of the group of suspicious words are greater than the preset threshold, the suspicious words are replaced; otherwise, the suspicious words are deleted;
[0023] The other group of suspicious words is deleted.
[0024] According to a text backdoor defense method based on SHAP values provided by the present invention, the step of replacing the suspicious word includes:
[0025] Masking the suspected words in the sentence to be processed;
[0026] Concatenate the original sentence to be processed and the masked sentence to be processed and input them into the BERT model to obtain the vocabulary probability distribution at the mask;
[0027] The word with the highest probability is selected as the replacement word at the mask according to the vocabulary probability distribution.
[0028] According to a text backdoor defense method based on SHAP values provided by the present invention, the steps of deleting or replacing the suspicious words include:
[0029] Based on the SHAP values and contextual semantics of the suspected words, the suspected words are selectively deleted and replaced to maintain the overall semantic consistency of the sentence to be processed.
[0030] The present invention also provides a text backdoor defense system based on SHAP values, comprising:
[0031] The backdoor trigger search module is used to obtain the SHAPvalues value of each feature word in the sentence to be processed using the SHAP interpreter;
[0032] A SHAP values analysis module, configured to select a first preset number of feature words with the largest SHAP values in the sentence to be processed as suspect words;
[0033] The suspicious word processing module is used to compare the SHAP values of the suspicious words with a preset threshold, and delete or replace the suspicious words according to the comparison result to obtain a new sentence.
[0034] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the text backdoor defense method based on SHAPvalues as described above is implemented.
[0035] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described text backdoor defense methods based on SHAP values.
[0036] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned text backdoor defense methods based on SHAP values.
[0037] The text backdoor defense method and system based on SHAP values provided by the present invention can accurately locate suspicious words in the input sample by calculating the contribution of each feature word in the input sample using the SHAP method, thereby effectively detecting backdoor triggers in the sample; select K feature words with the largest contribution as suspicious words based on the number K of suspicious words, compare the SHAP values of the suspicious words with a threshold λ, and process the suspicious words in the sample based on the comparison results by combining deletion and word replacement. The deletion operation removes the backdoor trigger in the sample, effectively achieving text backdoor defense while retaining the semantic integrity of the original sample through word replacement. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 This is a flowchart of the text backdoor defense method based on SHAP values provided by the present invention;
[0040] Figure 2 This is a complete flowchart of the text backdoor defense method based on SHAP values provided by the present invention;
[0041] Figure 3 This is a schematic diagram of the overall framework of the text backdoor defense method based on SHAP values provided by the present invention;
[0042] Figure 4 This is a visualization diagram of SHAP values in the text backdoor defense method based on SHAP values provided by the present invention;
[0043] Figure 5 This is a schematic diagram comparing the effects of three suspicious word processing methods in the text backdoor defense method based on SHAP values provided by the present invention;
[0044] Figure 6 It is a structural diagram of the text backdoor defense system based on SHAP values provided by the present invention;
[0045] Figure 7 This is an ASR effect diagram comparing the text backdoor defense method based on SHAP values provided by the present invention with other existing defense methods;
[0046] Figure 8 This is a CACC effect diagram comparing the text backdoor defense method based on SHAP values provided by the present invention with other existing defense methods. DETAILED DESCRIPTION
[0047] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0048] The following combination Figure 1 The present invention describes a text backdoor defense method based on SHAP values, comprising:
[0049] Step 101, using the SHAP (SHapley Additive exPlanations) interpreter to obtain the SHAPvalues of each feature word in the sentence to be processed;
[0050] First, we search for backdoor triggers. During the verification phase, we remove these triggers from the sentences being processed to achieve effective defense. We use the SHAP interpreter to analyze the SHAP values of each feature word in the sentences being processed. The SHAP values of each feature word reflect the contribution of the feature word to the prediction results of the poisoning detection model.
[0051] Step 102: taking a first preset number of feature words with the largest SHAP values in the sentence to be processed as suspect words;
[0052] Next, the SHAP values of the feature words are analyzed. After obtaining the SHAP values of each feature word in the sentence to be processed, a judgment operation is performed before processing the backdoor trigger to minimize the degree of sentence modification. A number of suspected words, K, and a threshold, λ, are set. After analyzing the SHAP values of each feature word, the specified number, K, of suspected words are accurately located. These suspected words are then processed accordingly based on the threshold, λ.
[0053] Step 103: Compare the SHAP values of the suspected word with a preset threshold, and delete or replace the suspected word according to the comparison result to obtain a new sentence.
[0054] Finally, the suspicious words are processed. Three possible solutions are available for suspicious words: deletion, word replacement, and hierarchical deletion and word replacement. Different processing methods can be used for different backdoor attack methods.
[0055] The present invention calculates the contribution of each feature word in the input sample by using the SHAP method, which can accurately locate the suspicious words in the input sample, thereby effectively detecting backdoor triggers in the sample; according to the number K of suspicious words, the K feature words with the largest contribution are selected as suspicious words, and the SHAP values of the suspicious words are compared with the threshold λ. According to the comparison results, the suspicious words in the sample are processed by combining deletion and word replacement. The deletion operation removes the backdoor triggers in the sample, effectively realizing text backdoor defense while retaining the semantic integrity of the original sample through word replacement.
[0056] Based on the above embodiment, before using the SHAP interpreter to obtain the SHAP values of each feature word in the sentence to be processed, this embodiment further includes:
[0057] Preprocessing the original sentence sample, wherein the preprocessing includes removing illegal characters in the original sentence sample;
[0058] Use the poisoning model to perform classification prediction on the preprocessed original sentence sample to obtain the predicted category y and predicted vector p of the original sentence sample;
[0059] The SHAP interpretable is constructed according to the predicted category and the predicted vector of the original sentence sample.
[0060] This paper introduces the concept of SHAP values to find suspicious words. The interpretability technology can also be used to extract the feature values in the sample during the process of explaining the model. Among them, the Shapley regression value is the feature importance of the linear model when multicollinearity exists. This method requires all feature subsets to be Retrain the model on , where F is the set of all features. It assigns an importance value to each feature, indicating the impact of including the feature on the model prediction. A model f s∪{i} is trained with existing features, and another model f s is trained with the retained features. Then, in the current input f s∪{i} (x)-f s(x). Since the effect of the retained feature depends on the other features in the model, for all possible subsets Calculate the previous differences. Then calculate the Shapley values and use them as feature attributes. They are weighted averages of all possible differences, as follows:
[0061]
[0062] like Figure 2 and Figure 3 As shown, in order to better detect the backdoor trigger in the sentence, the SHAP interpreter is used. Since the true label of the sample cannot be accurately obtained in the verification stage, the present invention cleans and preprocesses the original sentence sample, and then uses the poisoning model to classify and predict the original sentence sample x to obtain the prediction result y and the prediction vector p. The present invention uses y and p to construct the SHAP interpreter. The sentence to be processed is passed into the SHAP interpreter to obtain the SHAPvalues value of each feature word. Since the backdoor trigger has an obvious guiding effect on the prediction result of the model, one or more feature words with the largest SHAP values are likely to be triggers, and they are set as suspicious words. The suspicious words are processed to obtain normal sentences. The SHAP values visualization diagram is shown as follows Figure 4 shown.
[0063] On the basis of the above embodiment, in this embodiment, the first preset number when defending the sentence to be processed against a word-level backdoor attack is larger than the first preset number when defending the sentence to be processed against a sentence-level backdoor attack.
[0064] When defending against word-level backdoor attacks, a small number of suspicious words are processed. When defending against sentence-level backdoor attacks, the number of suspicious words K is appropriately increased.
[0065] Based on the above embodiment, in this embodiment, when defending the sentence to be processed against word-level backdoor attacks, the SHAP values of the suspected words are compared with a preset threshold, and the suspected words are deleted or replaced according to the comparison result to obtain a new sentence, including:
[0066] When the SHAP values of the suspected word are greater than the preset threshold λ, performing word replacement on the suspected word;
[0067] When the SHAP values of the suspected word are less than or equal to the preset threshold, the suspected word is deleted.
[0068] When defending against word-level backdoor attacks, the number of suspected words K is small, and the SHAP values of the suspected words are directly compared with the preset threshold λ. If the SHAP value of a suspected word is greater than the preset threshold λ, the suspected word is replaced; otherwise, the suspected word is deleted.
[0069] Based on the above embodiment, in this embodiment, when defending the sentence-level backdoor attack against the sentence to be processed, the SHAP values of the suspicious words are compared with a preset threshold, and the suspicious words are deleted or replaced according to the comparison result to obtain a new sentence, including:
[0070] According to the SHAP values of the suspicious words in the sentence to be processed, the suspicious words are divided into two groups, wherein the SHAP values of the suspicious words in one group are greater than the SHAP values of the other group of suspicious words;
[0071] If the SHAP values of the group of suspicious words are greater than the preset threshold, the suspicious words are replaced; otherwise, the suspicious words are deleted;
[0072] The other group of suspicious words is deleted.
[0073] When defending against sentence-level backdoor attacks, the number of suspected words, K, is appropriately increased. Furthermore, different SHAPvalues are processed in different levels. For one or more suspected words with the largest SHAP values, their SHAPvalues are compared with a preset threshold. If the SHAPvalue is greater than the threshold, the suspected word is replaced; otherwise, it is deleted. All other suspected words with lower SHAPvalues are deleted.
[0074] To further limit excessive sentence modification, this embodiment sets a threshold λ. By comparing the SHAPvalues with λ, the deletion or replacement operation is selected. When faced with multiple suspicious words, the present invention adopts a hierarchical processing approach.
[0075] Based on the above embodiments, the step of replacing the suspected word in this embodiment includes:
[0076] Masking the suspected words in the sentence to be processed;
[0077] Concatenate the original sentence to be processed and the masked sentence to be processed and input them into the BERT model to obtain the vocabulary probability distribution at the mask;
[0078] The word with the highest probability is selected as the replacement word at the mask according to the vocabulary probability distribution.
[0079] Suspected word replacements are generated using Bidirectional Encoder Representations from Transformers (BERT). When generating candidate word replacements, the semantics of the word itself and the context of the sentence are considered simultaneously, resulting in replacements that better align with the original text.
[0080] Specifically, the suspicious words in the original sentence are masked so as to be input into BERT to predict the masked tokens. Let S = w0,…,w L As a set of discrete tokens, for a suspected word W in sentence S, mask W to obtain a new sentence S1. The original sequence S and S1 are concatenated into a sentence pair and fed into BERT to obtain the probability distribution p(·|S,S'\{w}) of the words at the masked position. This method not only considers the word itself but also the semantics of the context. Finally, the word with the highest probability is selected as the replacement word.
[0081] In the deletion part, the suspect word is directly deleted from the original sentence to obtain a new sentence.
[0082] Based on the above embodiments, the steps of deleting or replacing the suspicious words in this embodiment include:
[0083] Based on the SHAP values and contextual semantics of the suspected words, the suspected words are selectively deleted and replaced to maintain the overall semantic consistency of the sentence to be processed.
[0084] While deletion can effectively remove triggers from malicious samples, it can also significantly damage clean samples. Word replacement can effectively preserve the fluency of clean sentences, but it rarely alters triggers in malicious samples. We selectively process suspicious words based on the values obtained from SHAP, combining deletion and word replacement.
[0085] During the suspicious word detection phase, the number of suspicious words is kept small to avoid excessive sentence modification that could lead to loss of semantics and other information. During word replacement, for shorter triggers (e.g., word-level triggers), the number of suspicious words K is set to 1; for longer triggers (e.g., sentence-level triggers), the number of suspicious words K is increased appropriately. Experiments have found that a K value of 3 achieves the best results.
[0086] The present invention provides a text backdoor defense method based on SHAP values, which specifically includes the following steps:
[0087] 1. Data preparation:
[0088] Collect text data, such as text data on social media platforms, including posts, comments, etc. Or product review data on e-commerce platforms.
[0089] Use the poisoning model to classify and predict text data to obtain the predicted category and predicted vector.
[0090] 2. SHAP value calculation:
[0091] Using the predicted categories and predicted vectors, a SHAP interpreter is constructed to analyze the feature words in each text sample.
[0092] Calculate the SHAP value of each feature word and identify the feature word that contributes most to the model prediction results.
[0093] 3. Processing of suspicious words:
[0094] Deletion: Delete suspicious words from the text sample and generate new output text.
[0095] Replacement: Use the BERT model to mask the suspected words, input them into BERT to predict replacement candidates for the masked words, and select the word with the highest probability for replacement.
[0096] Selective deletion and replacement: Based on the SHAP value and contextual semantics, the suspicious words are selectively deleted and replaced to maintain the overall semantic consistency of the text. Figure 5 shown.
[0097] Social media platforms are often subject to malicious attacks, where attackers use specific words to manipulate public opinion or spread false information. The following example illustrates how to use the method of the present invention to detect and process malicious trigger words in social media text.
[0098] Original text: "This is fake news, don't believe it!"
[0099] Suspect word detected: "fake news"
[0100] Processed text:
[0101] Delete: "This is one, don't believe it!"
[0102] Replace: "This is misleading information, don't believe it!"
[0103] Selectively delete and replace: "This is possibly untrue information, please don't believe it!"
[0104] The processed text maintains the overall semantics of the original text and removes malicious trigger words that may guide public opinion, which contributes to the content health of social media platforms.
[0105] Product reviews on e-commerce platforms can be subject to malicious attacks, where attackers use specific words to influence a product's reputation. The following example illustrates how to detect and address malicious trigger words in product reviews.
[0106] Original review: "This product is horrible and doesn't work at all!"
[0107] Suspect word detected: "bad"
[0108] Processed comments:
[0109] Delete: "This product is terrible and cannot be used at all!"
[0110] Replace: "This product is not good and cannot be used at all!"
[0111] Selective deletion and replacement: "There is something wrong with this product and it doesn't work at all!"
[0112] The processed reviews still retain users' negative feedback, but remove extreme or malicious trigger words that may affect the product's reputation, which helps to improve the fairness and credibility of the e-commerce platform's evaluation system.
[0113] The experimental datasets used in this paper are shown in Table 1. SST-2 is a sentiment tree library containing 6,921 training samples, 873 validation samples, and 1,822 test samples. The average length of each sample is 17 words. OLID (Offensive Language Identification Dataset) is a hierarchical dataset used to identify the types and targets of offensive text in social media, annotated as positive or negative. AG News is a dataset containing over 1 million news items.
[0114] Table 1 Experimental dataset
[0115]
[0116] This paper selected a pre-trained language model, BERT (bert-baseuncased). The model uses a code base from the Transformers library. The model was tested using both cleansed and contaminated samples.
[0117] This paper primarily targets three text backdoor attack methods: (1) BadNet, a typical word-level backdoor attack; (2) InsertSent, a sentence-level backdoor attack; and (3) Syntactic. This paper is not limited to defending against insertion-based text backdoor attacks; it also attempts to explore the threats posed by more challenging syntactic structure attacks. During the training phase, the poisoning rate of the BERT training set was set to 20%.
[0118] For a fair comparison, we selected two commonly used baseline methods. The first is ONION, which can be completed without user involvement in the model training process. ONION uses outlier detection to detect and eliminate outliers, providing excellent protection against word-level backdoor attacks. The second is IMBERT, the latest known backdoor defense method. IMBERT uses gradient and self-attention mechanisms to detect suspicious words and achieves defense by removing them. IMBERT can defend against both word-level and sentence-level backdoor attacks.
[0119] To measure the quality of generated samples, this paper employs various automatic evaluation metrics. The Attack Success Rate (ASR) evaluates the performance of backdoor attack methods. The Decrement of the Attack Success Rate (ASR) is a core metric for measuring the success of defense methods. The Clean Accuracy Rate (CACC) measures the accuracy of the backdoor model on the original clean test set. Ideally, performance degradation on clean data should be minimal, which is also the fundamental principle of backdoor attacks.
[0120] This paper uses the HuggingFace code library and the pre-trained model bert-baseuncased to train the poisoning model required by this paper on a dataset mixed with poisoned samples. The batch size, learning rate, and poisoning rate are set to 32, 2e-5, and 20, respectively.
[0121] This paper first evaluates whether SHAP can identify triggers from poisoned inputs. BadNet and InsertSent, as insertion-based backdoor attack methods, have relatively obvious triggers. Therefore, this paper primarily evaluates whether SHAP, when used as a detector, can accurately locate the triggers generated by these two attack methods. In the experiment, this paper uses SHAP Explain to perform quantitative analysis on each poisoned sample, identifying the feature words in each sample that have the greatest impact on the model's prediction results.
[0122] The present invention evaluates the detection accuracy of SHAP by analyzing whether these feature words are attacker-specific triggers. In Table 2, the present invention found that SHAP can accurately locate triggers with an accuracy of over 90%. Whether facing word-level attacks BadNet or sentence-level attacks InsertSent, SHAP can perform well. In the experiments here, only one trigger was used as the evaluation criterion, just to prove the effectiveness of using SHAP. When facing multiple triggers, the present invention will appropriately adjust the size of the number of suspected words K, which will be explained in detail in subsequent experiments.
[0123] Table 2 SHAP detection trigger success rate
[0124]
[0125] This paper has verified that SHAP can effectively locate triggers in malicious samples. To achieve effective defense, this paper requires processing these triggers. Currently, there are two existing processing methods: deletion and word replacement. For example, ONION and IMBERT both use deletion to remove triggers from samples, achieving effective defense. Similarly, previous papers have demonstrated that word replacement can be used to defend against insertion-based backdoor attacks.
[0126] Here, we first evaluated these two approaches. We used the BadNet and InsertSent attacks on the Agnews dataset to generate two training sets containing poisoned data, and then trained them using the BERT pre-trained model to generate two poisoned models. We then used SHAP to detect suspicious words and, after identifying them, adopted two approaches: deletion and replacement.
[0127] The present invention uses processed samples to evaluate the two poisoning models. The experimental results are shown in Table 3. When facing the sentence-level InsertSent attack, the method of deleting suspicious words can greatly reduce the attack success rate ASR, with a reduction of up to 98.2%. However, the impact on clean samples is also relatively large, and the clean accuracy rate is reduced by 8.7%. When the replacement method is used to process suspicious words, the impact on clean samples is minimal, and the clean accuracy rate is only reduced by 0.2%, but the defense effect on poisoned samples is far inferior to the deletion method mentioned above. Therefore, in order to ensure that the clean accuracy rate can be higher while reducing the ASR, combined with the contribution values of each feature word provided by SHAP to the present invention, the present invention proposes a new processing method that combines deletion and replacement.
[0128] Table 3. Processing suspicious words using deletion and replacement operations
[0129]
[0130] The present invention lists the defense effects of three defense methods against three poisoning attacks in three data sets in Table 4. It can be seen that when facing the word-level backdoor attack BadNet, the three defense methods of the present invention have good defense effects. ONION uses outlier word detection to detect triggers, IMBERT uses gradient and self-attention mechanism to detect triggers, and the method of the present invention uses SHAP to detect triggers. All three methods can effectively identify more obvious poisoned triggers in poisoned samples. The method of the present invention has the best defense effect on the SST and OLID data sets, and is only slightly inferior to IMBERT on AGNEWS. In the face of sentence-level backdoor attacks, the triggers in malicious samples are usually hidden in the text in the form of sentences, which makes ONION unable to find the malicious part of the sentence through outlier words, so ONION cannot defend against this type of poisoning attack. However, IMBERT and the method of the present invention can still detect multiple triggers in the sample. In terms of defense effect, the present invention can still outperform IMBERT on the two data sets of SST and OLID.
[0131] Table 4 ASR experimental results of the method of the present invention
[0132]
[0133] Here, we discuss the effectiveness of various defense methods in terms of clean sample accuracy. The experimental results are shown in Table 5, where "origin" represents the prediction accuracy of clean samples without the model's processing. Experiments show that our method achieves the highest accuracy in most cases, and performs best on the BadNets-attacked SST dataset, with an accuracy drop of only 0.4%. Our method occasionally underperforms ONION. We believe that ONION detects triggers by detecting outliers. In clean samples, these detected outliers are typically not core words, so deleting them will not significantly impact the model's prediction results, effectively ensuring clean sample accuracy. Our method, on the other hand, detects the core words that contribute the most to a sentence. Although we use word replacement to process these core words in an attempt to preserve the original semantic integrity of the sentence, the replacement words used can occasionally have semantic deviations, affecting the model's judgment. Compared to the IMBERT method, our method consistently achieves higher sample accuracy. This demonstrates that, compared to simply deleting suspicious words, our method for handling suspicious words better preserves the semantic integrity of clean samples.
[0134] Table 5 CACC experimental results of the method of the present invention
[0135]
[0136] In this work, we propose a novel defense method for preventing textual backdoor attacks based on insertion, which performs sample filtering during the verification phase. This method encompasses two aspects: suspicious word detection and suspicious word processing. Experiments on various datasets demonstrate the effectiveness of our method. Compared with two strong baselines, our method not only reduces the attack success rate but also significantly preserves the semantic integrity of clean samples.
[0137] The text backdoor defense system based on SHAP values provided by the present invention is described below. The text backdoor defense system based on SHAP values described below and the text backdoor defense method based on SHAP values described above can refer to each other.
[0138] like Figure 6 As shown, the system includes:
[0139] The backdoor trigger search module is used to obtain the SHAPvalues value of each feature word in the sentence to be processed using the SHAP interpreter;
[0140] A SHAP values analysis module, configured to select a first preset number of feature words with the largest SHAP values in the sentence to be processed as suspect words;
[0141] The suspicious word processing module is used to compare the SHAP values of the suspicious words with a preset threshold, and delete or replace the suspicious words according to the comparison result to obtain a new sentence.
[0142] Backdoor attacks can cause the model to misclassify input data, which will directly affect applications such as fake news detection, toxic content filtering, and public opinion mining, and will cause a series of economic losses. The text backdoor defense system developed by the present invention achieves defense functions by filtering poisoned samples during the model verification phase. When the user uses the model to perform a classification task, the system can be deployed in the classification task in advance. After the system receives the information input by the user, the system will process the input information received, remove the threatening part, and generate a sample with the same semantics but without backdoor information, which is passed into the classification model. In this way, even if the classification model used has been poisoned, the input sample can be classified normally after being processed by the system, thereby achieving a defensive effect.
[0143] Figure 7 This is an ASR effect diagram comparing the present invention with other existing defense methods. It can be found that the SHAPDBA of the present invention can have a lower ASR attack success rate on the same level. Figure 8This is a CACC advantage effect diagram of the present invention compared with other existing defense methods. It can be found that the SHAPDBA of the present invention can achieve a higher CACC clean sample accuracy rate on the same level.
[0144] The effectiveness of this defense method was evaluated on three different datasets and three backdoor attack methods. The experimental results show that the method not only significantly reduces the success rate of backdoor attacks, but also has better prediction accuracy for clean samples than the baseline.
[0145] The expected benefits and commercial value of this invention's technical solution are as follows: Backdoor attacks can cause models to misclassify input data, directly impacting applications such as fake news detection, toxic content filtering, and public opinion mining. Leaving these attacks unchecked can lead to economic losses. Developing an effective backdoor defense system can mitigate these attacks and ensure the safety of models for commercial use.
[0146] The technical solution of the present invention fills the technical gap in the industry at home and abroad. It is the first defense method to use explainability technology to defend against text backdoor attacks. The technical solution of the present invention solves the technical problems that people have always been eager to solve but have never been able to achieve success. Previous methods achieved backdoor defense by deleting detected triggers. Although this can reduce the success rate of backdoor attacks, it will destroy the semantic integrity of clean samples to a large extent. The method of the present invention adds a word replacement processing method, which is combined with the deletion method. It can not only effectively defend against text backdoor attacks, but also ensure the semantic integrity of clean samples.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A text backdoor defense method based on SHAP values, characterized in that: include: Use the SHAP interpreter to obtain the SHAP values of each feature word in the sentence to be processed; Taking a first preset number of feature words with the largest SHAP values in the sentence to be processed as suspicious words; The SHAP values of the suspected word are compared with a preset threshold, and the suspected word is deleted or replaced according to the comparison result to obtain a new sentence.
2. The text backdoor defense method based on SHAP values according to claim 1 is characterized in that: Before using the SHAP interpreter to obtain the SHAP values of each feature word in the sentence to be processed, it also includes: Preprocessing the original sentence sample, wherein the preprocessing includes removing illegal characters in the original sentence sample; Use the poisoning model to perform classification prediction on the preprocessed original sentence sample to obtain the predicted category and prediction vector of the original sentence sample; The SHAP interpretable is constructed according to the predicted category and the predicted vector of the original sentence sample.
3. The text backdoor defense method based on SHAP values according to claim 1, characterized in that: The first preset number when defending the sentence to be processed against a word-level backdoor attack is greater than the first preset number when defending the sentence to be processed against a sentence-level backdoor attack.
4. The text backdoor defense method based on SHAP values according to claim 3 is characterized in that: When defending the sentence to be processed against word-level backdoor attacks, the SHAP values of the suspected word are compared with a preset threshold, and the suspected word is deleted or replaced according to the comparison result to obtain a new sentence, including: When the SHAP values of the suspected word are greater than the preset threshold, performing word replacement on the suspected word; When the SHAP values of the suspected word are less than or equal to the preset threshold, the suspected word is deleted.
5. The text backdoor defense method based on SHAP values according to claim 3 is characterized in that: When defending the sentence-level backdoor attack against the sentence to be processed, the SHAP values of the suspected word are compared with a preset threshold, and the suspected word is deleted or replaced according to the comparison result to obtain a new sentence, including: According to the SHAP values of the suspicious words in the sentence to be processed, the suspicious words are divided into two groups, wherein the SHAP values of one group of suspicious words are greater than the SHAP values of the other group of suspicious words; When the SHAP values of the group of suspicious words are greater than the preset threshold, the suspicious words are replaced, otherwise the suspicious words are deleted; The other group of suspicious words is deleted.
6. The text backdoor defense method based on SHAP values according to any one of claims 1 to 5, characterized in that: The step of replacing the suspected word comprises: Masking the suspected words in the sentence to be processed; The original sentence to be processed and the masked sentence to be processed are connected and input into the BERT model to obtain the vocabulary probability distribution at the mask; The word with the highest probability is selected as the replacement word at the mask according to the vocabulary probability distribution.
7. The text backdoor defense method based on SHAP values according to any one of claims 1 to 5, characterized in that: The steps of deleting or replacing the suspected word include: According to the SHAP values and contextual semantics of the suspected words, the suspected words are selectively deleted and replaced to maintain the overall semantic consistency of the sentence to be processed.
8. A text backdoor defense system based on SHAP values, characterized in that: include: The backdoor trigger search module is used to obtain the SHAPvalues of each feature word in the sentence to be processed using the SHAP interpreter; A SHAP values analysis module, configured to select a first preset number of feature words with the largest SHAP values in the sentence to be processed as suspicious words; The suspicious word processing module is used to compare the SHAP values of the suspicious word with a preset threshold, and delete or replace the suspicious word according to the comparison result to obtain a new sentence.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the text backdoor defense method based on SHAP values as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the text backdoor defense method based on SHAP values as described in any one of claims 1 to 7 is implemented.