Word replacement-based text adversarial sample detection method and system

Through the text adversarial sample detection method based on word substitution, the word difference reaction score is calculated to screen important words and combined with auxiliary models for voting detection, the problem of insufficient accuracy of adversarial sample detection in the existing technology is solved, effective recognition and prediction correction of adversarial samples is achieved, and the robustness of the model is improved.

CN120353932APending Publication Date: 2025-07-22LIYANG RES INST OF SOUTHEAST UNIV +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510437862.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing text adversarial sample detection methods have insufficient accuracy, making it difficult to effectively identify and correct the adversarial attacks faced by deep learning models, especially the deception of adversarial samples on the model while maintaining semantic consistency.

Method used

The text adversarial sample detection method based on word substitution is used to filter important words by calculating word differential reaction scores, perform random replacement and combine auxiliary models for voting detection to correct the model's wrong prediction.

Benefits of technology

It significantly improves the detection accuracy and prediction correction capabilities of adversarial samples, can effectively identify and correct adversarial attacks to deep learning models, and improves the robustness and security of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353932A_ABST
    Figure CN120353932A_ABST
Patent Text Reader

Abstract

The invention discloses a text confrontation sample detection method and system based on word replacement. The method comprises the following steps: determining a text confrontation sample detection target; calculating word differential reaction scores and screening important words; carrying out random replacement and voting detection on the screened important words; detecting a sample, and if an adversarial sample is detected, giving a correct label of the adversarial sample; the text adversarial sample detection method based on word replacement comprises the following steps: calculating a word level differential reaction (WDR) based on logit, and capturing words having suspicious high influence on a classifier; then replacing words with synonyms, detecting the confrontation text by checking changes of previous and later tags and combining with a support model, and performing correct prediction; by detecting the adversarial text to identify whether there is an adversarial attack, the prediction results are corrected to protect the model from the adversarial attack.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of text adversarial sample detection, and particularly relates to a method and system for detecting text adversarial samples based on word replacement. Background Art

[0002] Currently, deep learning-based models have achieved high performance on many natural language processing (NLP) tasks. However, these models still show vulnerability to adversarial attacks. These attacks can deceive the model by perturbing only a small amount of the input text. More dangerously, the modified text still retains its original meaning, and humans cannot recognize the modifications in the text.

[0003] Although there are currently some defense techniques to solve this problem, DISP designed a perturbation recognizer and an embedding estimator to detect adversarial samples. The perturbation recognizer identifies the perturbation tokens in the text based on the context language model, and the embedding estimator restores the embedding vectors of the original words based on the context. Finally, the replacement tokens are selected through approximate kNN search, but the results are still limited in terms of performance. In the existing methods, the word frequencies in the text are used as the replacement objects, and many benign words may also be changed, resulting in misjudgment of clean samples. With the continuous development of text attacks, this naturally raises concerns about the safe and ethical deployment of NLP systems in real-world processes. Therefore, it is necessary to develop a new method and system for detecting text adversarial samples based on word replacement to solve the existing problems. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for detecting text adversarial samples based on word replacement to solve the problem of low accuracy in detecting text adversarial samples.

[0005] To achieve the above purpose, the present invention provides the following technical solutions: A method for detecting text adversarial samples based on word replacement, comprising the following steps:

[0006] Step 1, determining the text adversarial sample detection target;

[0007] Step 2, calculating the word differential response score to screen important words;

[0008] Step 3, randomly replacing the selected words and combining with an auxiliary model for voting detection;

[0009] Step 4, detecting the sample, and if an adversarial sample is detected, giving its correct label.

[0010] Preferably, the determining the text adversarial sample detection target in Step 1 is specifically as follows:

[0011] Given N texts and labels Accordingly, the deep learning model maps the input space to the label from any one of the N texts, the original text X org The generated valid adversarial text X adv must meet two criteria:

[0012] F(X adv )≠F(X org ) and Sim(X adv , X org )≥∈ (1)

[0013] where Sim(X adv , X org ) is the similarity between X adv and X org , ∈ is the minimum similarity between them. ε is the threshold that makes X adv and X org closer in meaning, for example, by semantic and syntactic criteria;

[0014] Two goals are set to process the input text X Input =[w1, w2,..., w n : (1) Determine whether the input text X Input is an adversarial text or an original text, (2) Correct the label distorted by the adversarial attack while maintaining the accuracy of the benign text.

[0015] Preferably, the calculating the word differential response score described in step 2 to screen important words is specifically:

[0016] Adversarial attacks based on semantic similarity change the prediction of the deep learning model by replacing as few words as possible. Therefore, when the attacker makes an adversarial sample, the words expected to be selected should have a greater impact on the prediction of the deep learning model. The present invention uses the word differential response (WDR) metric to measure the impact of words. Specifically, replacing the word w i has the following impact on the model prediction:

[0017]

[0018] In formula (2), Y * represents the final predicted category, that is, when a given input sample X is given, the attacked text model considers the most likely category; in formula (3), F(X\w i ) Y represents the logical value of the deep learning model F outputting its label Y when the word w i in the input sample X is replaced with an unknown word token; Denotes the logical value with the highest predicted probability in the error category after removing the replacement word w i ; The difference between the two represents the importance of the word w i . If this value is large, it indicates that after removing w i , the model still tends to predict the correct category; if this value is negative, it means that after removing w i , the model is more inclined to predict the wrong category. That is, if the input sample X is adversarial, the detector can expect to find a perturbation word with a negative WDR(w i ,F), because if there is no perturbation word, the input text sample should restore its original prediction.

[0019] The adversarial detector will first calculate the WDR scores WDR(w i ,F) of all replacement words w i in the input sample X for further detection later. It should be noted that a word with a negative WDR value is not necessarily always an adversarial word, but it usually indicates that the contribution of this word to the predicted category is "in the opposite direction", that is, its deletion will make the prediction more inclined to another category. If a word has a strong "misleading" effect on the classification result and deleting it will make the prediction result return to a more reasonable category, such words usually have negative WDR values. Therefore, a negative WDR value is only a hint indicating the contribution direction of the word to the model prediction. Whether it is an adversarial word needs to be judged by the significance change of word replacement and classification results in subsequent steps, and cannot simply rely on the positive or negative of the WDR value.

[0020] Preferably, the random replacement of the words screened in step 3 and the voting detection in combination with the auxiliary model are carried out as follows:

[0021] Screen out the words with negative WDR scores WDR(w Input ,F) from the input text X i to generate a reverse word set W neg , which plays an adversarial role in the prediction of the deep learning model and is more likely to be concentrated in potential perturbation positions; then use the auxiliary attack A to create a transformation set W neg for each word in the reverse word set W trans ;

[0022] For each word in the transformation set W trans , replace the corresponding original word in the input text X Input to generate a transformed text transformation set X trans . To ensure the rationality of the transformation, the following two aspects of checks need to be carried out on the text transformation set X trans : Similarity constraint, ensuring that the text transformation set X transSatisfy the similarity constraint of formula (1) to ensure that the deviations in semantics and text structure are controlled within the threshold range. Follow the specific restrictions of the auxiliary attack A. For example, PWWS prohibits modifying stop words to avoid affecting text fluency and semantic integrity. This step generates a text transformation set X through targeted similarity constraints trans closer to the distribution of actual attack samples.

[0023] Subsequently, the valid text transformation set X trans is input into the deep learning model F and the support model F sup to generate hard labels, that is, the class labels with the highest probability. These class labels are stored in the local list Y trans corresponding to the transformation set X trans . If the labels in Y trans are inconsistent with those in Y input , the inconsistent labels are stored in the global list Y correct for later corrective prediction. Y input represents the label value generated after the input text X input is input into the model F. The support model F sup is another selected deep learning model, different from the attacked model, used to assist in detection, and different support models can be selected.

[0024] Preferably, for the detection of samples in step 4, if an adversarial sample is detected, its correct label is given, and the specific process is as follows:

[0025] To quantify the adversarial nature of the input text, calculate the difference rate R, defined as the ratio of the number of labels conflicting in the local list Y trans and Y input ; define two thresholds, the adversarial detection value λ adv and the prediction correction value λ mis ; the adversarial detection value. λ adv is used for adversarial detection. If the difference rate R is greater than the adversarial detection value λ adv , the input sample X is considered an adversarial sample; the prediction correction value λ mis is used for prediction correction. If the difference rate R is greater than the prediction correction value λ mis , then by voting on Y correct , select the label with the highest proportion as the final prediction result to correct the original misclassification. The adversarial detection value λ adv and the prediction correction value λ mis are optimized through the validation set to achieve the best performance in adversarial detection and prediction correction.

[0026] The present invention further provides a text adversarial sample detection system based on word replacement, including:

[0027] A determination module, configured to determine a text adversarial sample detection target;

[0028] A calculation module, configured to calculate word differential response scores to screen important words;

[0029] A replacement detection module, configured to randomly replace the screened important words and perform voting detection;

[0030] A labeling module, configured to detect a sample, and if an adversarial sample is detected, give its correct label.

[0031] Technical effects and advantages of the present invention: The text adversarial sample detection method and system based on word replacement calculates the word-level differential response (WDR) based on logical values to capture words with suspiciously high impact on the classifier; then replaces the words with their synonyms, detects adversarial text by checking the changes in labels before and after and combining with a support model, and makes correct predictions; detects adversarial text to identify whether there is an adversarial attack, and corrects the prediction results to protect the model from adversarial attacks. Description of the Drawings

[0032] Figure 1 It is a schematic flowchart of the method of the present invention;

[0033] Figure 2 It is a schematic diagram of the overall architecture of the text adversarial sample detection of the present invention. Specific Embodiments

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] Embodiment 1

[0036] The present invention provides a text adversarial sample detection method based on word replacement as shown in Figure 1 , Figure 2 , including the following steps:

[0037] Step 1, determine a text adversarial sample detection target;

[0038] Step 2, calculate word differential response scores to screen important words;

[0039] Step 3, randomly replace the screened words and perform voting detection in combination with an auxiliary model;

[0040] Step 4: Detect the samples. If adversarial samples are detected, give their correct labels.

[0041] In step 1, the determination of the detection target of text adversarial samples is as follows:

[0042] Given N texts and labels Correspondingly, the deep learning model maps the input space to the label The text is a sentence set containing multiple sentences; the valid adversarial text X org generated from the original text X adv must meet two criteria:

[0043] F(X adv ) ≠ F(X org ) and Sim(X adv , X org ) ≥ ∈ (1)

[0044] where Sim(X adv , X org ) is the similarity between the valid adversarial text X adv and the original text X org , and ∈ is the minimum similarity between them. ∈ is the threshold that makes the valid adversarial text X adv and the original text X org closer in meaning, for example, through semantic and syntactic criteria.

[0045] Two goals are set to process the input text X Input = [w1, w2,..., w n : (1) Determine whether the input text X Input is an adversarial text or an original text, (2) Correct the label distorted by the adversarial attack while maintaining the accuracy of the benign text.

[0046] Furthermore, the calculation of the word differential response score in step 2 to screen important words is specifically as follows:

[0047] Adversarial attacks based on semantic similarity change the prediction of the deep learning model by replacing as few words as possible. Therefore, when making adversarial samples, the attacker hopes to select words that have a greater impact on the prediction of the deep learning model. The word differential response (WDR) metric is used to measure the impact of words. Specifically, the impact of replacing the word w i on the model prediction is as follows:

[0048]

[0049] Y in formula (2) * represents the final predicted category, that is, when given the input sample X, it is the category that the attacked text model considers most likely; in formula (3): WDR(w i ,F) represents the WDR score, and F(X\w i ) Y represents the logical value of the output label Y of the deep learning model F when the word w i in the input sample X is replaced by an unknown word token; the label Y is one of the labels in the label set.

[0050] represents the logit value with the largest predicted probability in the wrong category by the attacked text model after removing the replacement word w i . The difference between the two represents the importance of the word w i . If this value is large, it means that after removing w i , the model still tends to predict the correct category; if this value is negative, it means that after removing w i , the model is more inclined to predict the wrong category. That is, if X is adversarial, the detector can expect to find perturbed words with negative WDR(w i ,F), because without them, the input text sample should restore its original prediction.

[0051] Table 1 shows an example of a pair of original text and adversarial text and their corresponding WDR(w i ,F) scores. After removing the perturbed words in the adversarial sample, the original category is restored. This switch results in a negative WDR value. However, even if the most important word ("boring") in the original sentence is removed, the predicted category does not change, so the WDR value is greater than 0.

[0052] The adversarial detector will first calculate the WDR scores WDR(w i of all words w i, F), for further detection later. It should be noted that words with negative WDR values are not necessarily always adversarial words, but they generally indicate that the contribution of the word to the predicted category is in the "opposite direction", that is, deleting it will make the prediction more inclined to another category. If a word has a strong "misleading" effect on the classification result and deleting it makes the prediction result return to a more reasonable category, such words usually have negative WDR values. For example, in the example of Table 1, "thrilling" is an adversarial word, and deleting it makes the prediction shift from category 1 to category 0. For some non-adversarial words, they may only have a slight negative contribution to the classification, and deleting them will not significantly improve the classification accuracy, but the WDR values of such words may be negative. Therefore, a negative WDR value is only a hint indicating the contribution direction of the word to the model prediction. Whether a word is an adversarial word or not needs to be judged by word replacement and significant changes in the classification result in subsequent steps, and cannot simply rely on the positive or negative of the WDR value.

[0053] WDR scores calculated for the original samples in Table 1 and their corresponding adversarial samples

[0054]

[0055] The process of randomly replacing the selected words as described in Step 3 and combining with an auxiliary model for voting detection is as follows:

[0056] Select words with negative WDR scores from the input text X Input to generate a reverse word set W neg , these words usually play an adversarial role in the prediction of the deep learning model and are more likely to be concentrated in potential perturbation positions. Subsequently, use the auxiliary attack A to create a transformation set W neg for each word in the reverse word set W trans . The auxiliary attack A refers to any attack method;

[0057] For each word in the transformation set W trans , replace the corresponding original word in the input text X Input to generate a transformed text transformation set X trans . To ensure the rationality of the transformation, the following two aspects of checks need to be performed on the text transformation set X trans : Similarity constraint, ensure that the text transformation set X trans satisfies the similarity constraint of Equation (1) to ensure that the deviation in semantics and text structure is controlled within the threshold range. Follow the specific restrictions of the auxiliary attack A. For example, PWWS prohibits modifying stop words to avoid affecting text fluency and semantic integrity. Through targeted similarity constraints in this step, the generated X trans is closer to the distribution of actual attack samples.

[0058] Subsequently, the valid text conversion set X trans is input into the deep learning model F and the support model F sup to generate hard labels, i.e., the class labels with the highest probability. These labels are stored in the corresponding local list Y of the conversion set trans . If the labels in the local list Y trans are inconsistent with those in Y input , the inconsistent labels are stored in the global list Y correct for later corrective prediction; Y input represents the labels generated after the input text X input is input into the deep learning model F. The local list Y trans is a table containing multiple label values, and Y input is the label value of the original sample. The valid conversion set X trans is a set of words that satisfy all constraints.

[0059] For the samples detected in step 4, if adversarial samples are detected, their correct labels are given. The specific process is as follows:

[0060] To quantify the adversarial nature of the input text, the difference rate R is calculated, which is defined as the ratio of the number of labels conflicting in the local list Y trans to those in Y input . Two thresholds, the adversarial detection value λ adv and the prediction correction value λ mis , are defined. The prediction correction value λ adv is used for adversarial detection. If the difference rate R is greater than the prediction correction value λ adv , the input sample is considered an adversarial sample. The prediction correction value λ mis is used for prediction correction. If the difference rate R is greater than the prediction correction value λ mis , then by voting on Y correct , the label with the majority, i.e., the label with the highest proportion, is selected as the final prediction result to correct the original misclassification. The prediction correction value λ adv and the prediction correction value λ mis are optimized through the validation set to achieve the best performance in adversarial detection and prediction correction.

[0061] Experiment

[0062] The present invention conducts experiments using two widely recognized benchmark datasets: the IMDB dataset and the SST-2 dataset. Widely recognized natural language processing (NLP) models are used as victim models: RoBERTa, BERT, WordCNN, and WordLSTM.

[0063] The present invention utilizes the pre-trained RoBERTa (base) model provided by the Hugging Face Transformers library. After byte-pair encoding, for the IMDb and SST-2 datasets, the maximum input sequence lengths are set to 256 and 128 respectively. The RoBERTa model contains 125 million parameters. During the training process, the model uses batch sizes of 32 (SST-2) and 16 (IMDb) respectively, and the learning rate is 1·10 -5 , and it is trained for 10 epochs. After each epoch ends, the model performance is evaluated on the validation set, and the checkpoint with the best performance is selected for testing.

[0064] The convolutional neural network (CNN) architecture consists of 3 convolutional layers with kernel sizes of 2, 3, and 4 respectively, and each convolutional layer has 100 feature maps. The hidden state dimension of the long short-term memory network (LSTM) is 128. The LSTM and CNN are initialized using pre-trained GloVe word vectors. During the training process, both the LSTM and CNN adopt Dropout before the output layer with a ratio of 0.1. The Adam optimizer is used to train the two models for 20 epochs. After each epoch ends, the model performance is evaluated on the validation set, and the checkpoint with the best performance is selected for testing. The training batch size for the CNN and LSTM models is 100, and the learning rate is 1·10 -3 .

[0065] The adversarial text detection task includes Recall, False Positive Rate (FPR), and F1-score. Recall represents the proportion of adversarial texts that are correctly detected as adversarial samples. The False Positive Rate represents the proportion of normal texts that are misdetected as adversarial texts. The F1-score is the harmonic mean of TPR and precision, which is a comprehensive metric for measuring the detection performance.

[0066] The definition of the False Positive Rate (FPR) is:

[0067]

[0068] The definition of Recall is:

[0069]

[0070] The definition of F1-score is:

[0071]

[0072] The prediction correction metrics include Adv Accuracy, which is the classification accuracy of adversarial samples before correction, and AdvCorrection Accuracy, which is the classification accuracy of adversarial samples after correction.

[0073] Select FGWS as the comparison baseline. This method uses the word frequency feature of adversarial word replacement to detect text adversarial examples and has been proven to be a relatively simple and effective method conceptually.

[0074] First, on the SST-2 and IMDB datasets, use PWWS to generate adversarial examples for three common models (CNN, LSTM, BERT) to evaluate this method and the FGWS method. Use RoBERTa as the supporting model for the CNN, LSTM, and BERT deep learning models. The experimental results are shown in Table 2.

[0075] The experimental results show that this method is significantly better than the FGWS method in the tasks of adversarial text detection and correction. For the SST-2 and IMDB datasets, for the adversarial examples generated for the three models of CNN, LSTM, and BERT, it performs excellently in terms of F1-score and the accuracy after correction (Adv Correction Accuracy). In terms of F1-score, this method significantly exceeds FGWS in all combinations. For example, the F1-score of the BERT model on the SST-2 dataset increases from 84.4 to 88.2, and the BERT model on the IMDB dataset increases from 91.8 to 94.8, which indicates that the improved method can detect adversarial examples more accurately. In terms of the accuracy after correction, the improvement of this method is even more significant. For example, the accuracy after correction of CNN on the SST-2 dataset increases from 65.5 to 95.2, and the accuracy after correction of CNN on the IMDB dataset increases from 71.6 to 98.5. This significant improvement benefits from the fact that the improved method generates a transformation set that is closer to the distribution of real adversarial examples by screening words with negative WDR scores and concentrating on possible adversarial perturbation positions. In addition, using RoBERTa as the supporting model provides stronger semantic support information for the deep learning models and significantly enhances the accuracy of correction. Although the original accuracy of the adversarial examples (Adv Accuracy) is low, indicating that the adversarial examples generated by PWWS are highly aggressive to the deep learning models, this method can effectively detect and correct these examples, showing excellent generality for both short texts (SST-2) and long texts (IMDB).

[0076] Table 2 Detecting Adversarial Texts Generated by PWWS and Making Predictive Corrections

[0077]

[0078] In addition, experiments were also conducted on the adversarial samples generated by different adversarial attack methods for the CNN model on the SST-2 dataset, as shown in Table 3. The experimental results show that this application is significantly superior to the FGWS method under various adversarial attacks (TextFooler, BAE, PSO, PWWS) against the CNN model, demonstrating its powerful ability to detect and correct different adversarial samples. In terms of the balance between TPR (True Positive Rate) and FPR (False Positive Rate), the TPR was significantly improved while maintaining a low FPR. For example, under the TextFooler attack, the TPR increased from 43.2 to 87.6, while the FPR only slightly increased to 11.8. In terms of F1-score, this application was significantly improved compared to FGWS. For example, under the PWWS attack, the F1-score increased from 76.3 to 87.9. For Adv Correction Accuracy (accuracy after correction), the improved method performed excellently under all attack types. For example, under the PSO attack, it increased from 45.8 of FGWS to 93.6. These results indicate that this application can not only effectively detect the adversarial samples generated by various attacks, but also accurately correct the predictions of adversarial samples, especially showing strong robustness in complex attack scenarios. The improvement in performance benefits from the accuracy of the method in screening potential perturbation positions and the semantic enhancement ability of RoBERTa as a support model.

[0079] Table 3 Adversarial Text Detection and Prediction Correction for CNN Models on SST-2

[0080]

[0081]

[0082] Example 2

[0083] The present invention further provides a text adversarial sample detection system based on word replacement, including:

[0084] A determination module for determining the text adversarial sample detection target;

[0085] A calculation module for calculating the word difference response score to screen important words;

[0086] A replacement detection module for randomly replacing the screened important words and voting for detection; a labeling module for detecting samples and giving the correct label if adversarial samples are detected.

[0087] In summary, the present invention studies the robustness of natural language models in the field of text adversarial defense. First, it calculates the word-level differential response (WDR) based on logical values to capture words that have a suspiciously high impact on the classifier. Subsequently, it replaces the words with their synonyms, detects adversarial texts by checking the changes in labels before and after and combining with a support model, and makes correct predictions. The effectiveness of the present invention is demonstrated through comparative experiments.

[0088] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A text adversarial sample detection method based on word substitution, characterized in that: It includes: Determine the detection target of text adversarial samples; Calculate the word differential response score to screen important words; Randomly replace the screened important words and conduct a voting detection; Detect the sample, and if an adversarial sample is detected, give its correct label.

2. The text adversarial sample detection method based on word substitution according to claim 1, characterized in that: The determination of the detection target of text adversarial samples includes: Given N texts and labels Input the texts as samples into the deep learning model F and map to obtain the labels The expression is as follows: From the original text X org The generated valid adversarial text X adv Satisfy: F(X adv ) ≠ F(X org ) and Sim(X adv , X org ) ≥ ∈ (1) where: Sim(X adv , X org ) is the similarity between the adversarial text X adv and the original text X org , and ∈ is the minimum similarity between the adversarial text X adv and the original text X org ; Set two targets to process the input text X Input = [w1, w2,..., w n : where: W represents the words in the input text, w1 represents the first word, w n represents the nth word.

3. The method for detecting text adversarial samples based on word replacement according to claim 2, wherein: The two goals include: determining the input text X Input is adversarial text or the original text X org and correcting the labels distorted by adversarial attacks 4. A text adversarial sample detection method based on word replacement according to claim 1, characterized in that: The calculation of the word differential response score includes: Replace the word w i The impact on the prediction of the attacked text model is as follows: Y in formula (2) * represents the final predicted class, that is, when a given input sample X is provided, it is the class that the attacked text model believes is the most likely; In formula (3): WDR(w i , F) represents the WDR score, and F(X\w i ) Y represents the logical value of the label Y output by the deep learning model F when the word w i in the input sample X is replaced with an unknown word token; represents the logical value with the highest predicted probability in the error category of the attacked text model after removing the replacement word w i .

5. A text adversarial sample detection method based on word replacement according to claim 4, characterized in that: The random replacement and voting detection of the screened important words includes: Filter out the words in the input text X whose WDR scores in formula (3) are negative to generate the reverse word set W Input ; neg ; Create a transformation set W for each word in the reverse word set W by means of the auxiliary attack A neg ; trans ; For each word in the transformation set W trans replace the corresponding original word in the input text X Input to generate the transformed text transformation set X trans ; Constrain the conversion set X trans Perform the constraint; Convert the valid text conversion set X trans and input it into the deep learning model F and the support model F sup to generate hard labels, that is, the class labels with the highest probability; the class labels with the highest probability are stored in the conversion set X trans into the corresponding local list Y trans ; if the labels in the local list Y trans are inconsistent with Y input , the inconsistent labels are stored in the global list Y correct to correct the prediction; where Y input represents the label value generated after the input text X input is input into the deep learning model F 6. The method for detecting text adversarial samples based on word replacement according to claim 5, characterized in that: The constraint on the conversion set X trans includes:; Make the conversion set X trans Satisfy the similarity constraints of formula (1), ensuring that the deviations in semantics and text structure are controlled within the threshold range and follow the specific limitations of the auxiliary attack A.

7. A text adversarial sample detection method based on word replacement according to claim 1, characterized in that: The detection of the sample and giving the correct label if an adversarial sample is detected includes: Calculate the difference rate R, defined as the ratio of the number of labels in the local list Y trans that conflict with Y input in Y; Define the adversarial detection value λ adv and the prediction correction value λ mis ; If the difference rate R is greater than the adversarial detection value λ adv , then the input sample X input is an adversarial sample; If the difference rate R is greater than the predicted correction value λ mis , then by voting on the global list Y correct , select the label with the highest proportion as the final prediction result and correct the original misclassification; Adversarial detection value λ adv and the prediction correction value λ mis Optimized by the validation set.

8. A text adversarial sample detection system based on word substitution, characterized in that: It includes: A determination module for determining the detection target of text adversarial samples; A calculation module for calculating the word differential response score to screen important words; A replacement detection module for randomly replacing the screened important words and conducting a voting detection; A label module for detecting the sample and giving the correct label if an adversarial sample is detected.

Citation Information

Cited By

  • Text adversarial sample construction method and system based on key semantic node reconstruction

    CN122432662A