A method for detecting adversarial examples based on perturbation sensitivity differences

By extracting the characteristics of text adversarial samples based on the difference in perturbation sensitivity, the detector is trained, and the problems of high computational complexity and poor versatility in the prior art are solved, and efficient and general adversarial samples detection is achieved.

CN115408516BActive Publication Date: 2025-06-06GUANGZHOU UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210807194.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-11
Publication Date
2025-06-06
Estimated Expiration
2042-07-11

AI Technical Summary

Technical Problem

The existing adversarial detection technology has high computational complexity, poor versatility, and low defense efficiency in the application of text fields, making it difficult to quickly extract features and be suitable for different types of attack scenarios.

Method used

Adversarial sample detection method based on the difference in perturbation sensitivity is used to determine important words through gradient estimation, perturb these words and extract adversarial features, and a binary classification adversarial detector is trained to identify adversarial samples.

Benefits of technology

This method can quickly extract features under constant-order complexity, has high versatility and generalization, and has a high detection rate. It is suitable for different data sets, attack methods and target models, with a detection rate of 90%-100%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115408516B_ABST
    Figure CN115408516B_ABST
Patent Text Reader

Abstract

The present invention discloses an adversarial sample detection method based on perturbation sensitivity difference, comprising the following steps: Step 1: Generate adversarial samples using an attack algorithm; Step 2: Determine important words using gradient estimation; Step 3: Perturb important words to extract adversarial features; Step 4: Use adversarial features as training data to train a binary adversarial detector; Step 5: Input the text to be tested into the adversarial detector and output the result. The present invention uses perturbation sensitivity difference to extract adversarial features, which greatly improves the extraction efficiency compared to the complex characterization vector construction method in the prior art. The adversarial feature extraction method of the present invention is based on the universality of adversarial samples and has strong versatility. Compared with the prior art that can only detect adversarial samples generated for a certain or a certain type of attack method, it has universal adaptability and generalizability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security technology of deep learning model applications, and in particular to a method for detecting adversarial samples based on perturbation sensitivity differences. Background Art

[0002] The development of deep learning methods and the innovative evolution of computer systems have promoted the widespread application of deep learning models in the real world. At the same time, the security issues of model deployment have received widespread attention, especially in some security-sensitive applications such as face recognition and fraud detection. Adversarial attacks are a new type of attack method that has attracted much attention in recent years. They impose imperceptible perturbations on the original input to generate adversarial samples, and use adversarial samples to deceive the model, causing the model to make incorrect judgments. Such attacks will cause serious threats to the reliability of the model. Therefore, research on defense technologies against adversarial sample attacks has received widespread attention in academia and industry.

[0003] In the field of natural language processing, the defense methods against adversarial samples are mainly divided into adversarial training, input reconstruction, and adversarial detection. Although adversarial training has shown good results in the image field, due to the discrete nature of text, this method has little improvement in adversarial robustness in the text field; the input reconstruction method performs better than adversarial training in terms of improving robustness, but these methods are all aimed at adversarial attacks at the word level, with limited scope of application and difficult to promote; adversarial detection is different from the first two defense methods. It only needs to identify whether the output is an adversarial sample, without giving the true label of the adversarial sample. Adversarial detection can prevent maliciously constructed samples from entering the target model, which is of great significance in real applications. This type of method has been widely used in the image field, but there is little related research in the text field.

[0004] Among the existing adversarial detection technologies, the Chinese invention patent with publication number CN114169443A disclosed a word-level text adversarial detection method on March 11, 2022, which includes: using an adversarial attack algorithm to generate adversarial samples of corresponding clean samples, extracting feature vectors for representing the clean samples and adversarial samples respectively; secondly, using a deep learning model to build a detector. This method is only applicable to word-level synonym replacement attack methods, and its construction of the representation feature vector requires relying on the synonym table for permutations and combinations and requires a large number of access models. Its computational complexity is exponential, and the construction process is cumbersome, resulting in poor versatility and low defense efficiency of the method.

[0005] Therefore, how to design a universal adversarial detection method that can reduce calculations and achieve a faster feature extraction speed is an urgent problem that technicians in this field need to solve. Summary of the invention

[0006] In view of the above problems, the present invention proposes an adversarial sample detection method based on perturbation sensitivity difference to provide a detection method for text adversarial defense of deep neural network models to solve the above problems.

[0007] The technical solution provided by the present invention is:

[0008] A method for detecting adversarial examples based on perturbation sensitivity differences comprises the following steps:

[0009] Step 1: Generate adversarial samples using attack algorithms;

[0010] Step 2: Determine important words using gradient estimation;

[0011] Step 3: Perturb important words and extract adversarial features;

[0012] Step 4: Use the adversarial features as training data to train a binary adversarial detector;

[0013] Step 5: Input the text to be tested into the adversarial detector and output the result.

[0014] Preferably, in step 1, the process of generating adversarial samples includes constructing a new dataset T = {x j , x j *}, 0<j<t, where x j is a clean sample from D, x j * Some attack algorithm generates x j Corresponding adversarial examples; D = {x i ,y i},0<i<l,x i is the text in dataset D, y i is x i The corresponding label, that is, f(x i )=y i , x i It can be further expressed as x i =[w 1 , w 2 , ..., w i , ..., w n ], where w i Is the text x i In the word, n is the number of words, t is the number of samples in the new data set T, and l is the number of samples in the data set D.

[0015] Preferably, the method for constructing the dataset T includes randomly sampling a data (x, y) from the dataset D, and generating an adversarial sample x corresponding to x through a deep neural network model.* =[w 1 * , w 2 * , ..., w n * ], its towel is an adversarial sample x * The words in the, n is the number of words, if the attack is successful, that is, f(x * )≠y, then (x, x * ) is added to T. Repeat step 1 until there are t text pairs in T; where random sampling from D is a sampling strategy without replacement.

[0016] Preferably, in step 2, the most important k words are selected from all texts x in T, including adversarial samples and clean samples, and are sorted, which is recorded as C(x).

[0017] Preferably, in step 3, for each word in C(x), the predicted labels before and after the word is deleted are obtained by accessing the deep neural network model F, and the words with inconsistent predicted labels before and after the word is deleted are defined as sensitive words, and vice versa. The deletion operation means replacing the original word with a token. For the pre-trained models Bert and RoBERTa, the replacement token is [MASK], and for the traditional DNNs models LSTM and CNN, the replacement token is <unk>.

[0018] Preferably, the signal set of the sensitive words is used to measure the sensitivity of the text x to F, expressed as S(x, f(x)), which is formulated as:

[0019]

[0020] in Mathematical expression for word sensitivity:

[0021]

[0022] in is the text x without word w i , f(x) represents the output function of the deep neural network model F.

[0023] Better, through JSD divergence To calculate the similarity between probability distributions, the larger the JSD divergence value, the more similar they are and the closer the distributions are. Conversely, a smaller value indicates that the distributions have changed significantly:

[0024]

[0025] where f s (x) represents the probability distribution of the softmax layer, KL stands for Kullback-Leibler divergence and is expressed as:

[0026]

[0027] For each word in C(x), the JSD divergence is calculated and used as the distribution variance features of x, expressed as:

[0028]

[0029] Preferably, the input feature E(x, f(x)) of the adversarial detector consists of the sensitive signal JSD value, expressed as:

[0030] E(x, f(x))=S(x, f(x))*J(x, f(x))

[0031] Therefore, the input features of the adversarial detector are a set of continuous vectors of size k, and the labels are binary, 0 represents clean samples and 1 represents adversarial samples.

[0032] Preferably, in step 4, during the training phase of the adversarial detector, the data is divided into a training set and a test set in a ratio of 8:2.

[0033] Preferably, in step 5, for any text, steps 2 and 3 are repeated to extract features, and the features are input into the adversarial detector trained in step 4; if the adversarial detector outputs 0, the text is judged to be a normal sample; if the output is 1, the text is judged to be an adversarial sample.

[0034] Compared with the prior art, the adversarial sample detection method based on perturbation sensitivity difference provided by the present invention has the following advantages:

[0035] 1. The present invention is applicable to different types of attack scenarios, including word-level and character-level attacks. The present invention uses the sensitivity difference between adversarial samples and clean samples to design an adversarial feature extraction method based on important word perturbation. Since the sensitivity difference is universal between adversarial samples and clean samples, the proposed adversarial feature extraction method is universal and extensible, which improves the defect that the previous technology can only be used in specific attack scenarios.

[0036] 2. The detection technology proposed in the present invention is highly scalable. The detection technology proposed in the present invention does not rely on a specific target model and can be deployed outside the target model as an additional component to improve the adversarial robustness of the target model without making any modifications to the target model.

[0037] 3. The detection method proposed in the present invention has a high detection rate, reaching 90%-100% in different data sets, attack methods, and target models.

[0038] 4. The adversarial feature extraction method used in the detection method proposed in the present invention has a constant-level complexity, which effectively improves the detection efficiency compared to previous technologies that require exponential computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The present invention is further described using the accompanying drawings, but the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative work.

[0040] Figure 1 It is a schematic diagram of the overall flow of the adversarial detection method based on disturbance sensitivity differences of the present invention. DETAILED DESCRIPTION

[0041] The adversarial sample detection method based on perturbation sensitivity difference is further described in detail below in conjunction with the accompanying drawings and specific embodiments. These embodiments are only for the purpose of comparison and explanation, and the present invention is not limited to these embodiments.

[0042] Example 1

[0043] See also Figure 1 , the adversarial sample detection method based on perturbation sensitivity difference provided by the embodiment of the present invention comprises the following steps:

[0044] Step 1: Generate adversarial samples using attack algorithms;

[0045] Step 2: Determine important words using gradient estimation;

[0046] Calculate the importance of each word in x, where the importance of the word is calculated by gradient estimation to measure the importance of the word. The contribution of each word to the prediction is estimated by using the gradient. The direction of gradient descent is to help the model.

[0047] The optimization signal with the minimum loss is obtained during the training phase; therefore, the closer the embedding vector is to the gradient direction, the greater the contribution of the word to the prediction of F. Specifically, by using the dot product to represent the angle between the word embedding vector and the word gradient, the calculation formula is:

[0048]

[0049] Where v is the word dimension and J is the loss function of model F. In this way, only one access to the model is required to obtain the importance scores of all words. In many existing works, word importance is obtained by comparing the changes in output probability and predicted label before and after deleting the word. The formula is expressed as:

[0050]

[0051] in is the text x without word w i The representation of f(x, y j ) is the model F predicting x as label y j In this way, the model needs to be accessed n times to calculate the importance scores of all words, where n is the number of words. The method proposed in the present invention only requires one query model to calculate the importance scores of all words.

[0052] Sort the words and represent the top_k words as C(x). Subsequent feature extraction is performed around k words. Here, top_k is a threshold set manually. The experimental conclusion is that its reasonable value range is [0.1n, 0.2n].

[0053] No matter how the attacker designs the attack strategy, the ultimate goal is the same: minimize the modification rate of the adversarial sample and maximize the semantic similarity between the adversarial sample and its corresponding clean sample, which is also the basic condition for satisfying the adversarial instance. In order to achieve these goals, the attacker usually finds important words and interferes with them, rather than making meaningless modifications to those unimportant words. Therefore, important words are strong signals reflecting the difference between adversarial samples and normal samples, and are also the most critical feature source in adversarial detection.

[0054] Step 3: Perturb important words and extract adversarial features;

[0055] Adversarial samples and clean samples react differently to the perturbation of important words. Due to their boundary sensitivity, the output labels of adversarial samples are extremely susceptible to changes. Under the same perturbation, the predicted probability of each class of the clean sample will change, but the final predicted label remains relatively stable, which is similar to the principle that partial distortion of the image does not affect the model decision. Based on the inconsistency of the response of adversarial samples and clean samples to perturbations, a method to define the sensitivity of the input x is proposed: for each word in C(x), the predicted label before and after the word is deleted is obtained by accessing the model F. Words with inconsistent predicted labels before and after the word is deleted are defined as sensitive words, and vice versa. Insensitive words. More precisely, the deletion operation means replacing the original word with a specified token representation. For the pre-trained models Bert and RoBERTa, the replacement token is [MASK], and for the traditional DNNs models LSTM and CNN, the replacement token is [MASK]. <unk>Furthermore, the signal set based on sensitive words is used to measure the sensitivity of text x to F, expressed as S(x, f(x)), which is formulated as:

[0056]

[0057] in Mathematical expression for word sensitivity:

[0058]

[0059] in is the text x without word w i , f(x) represents the output function of the deep neural network model F.

[0060] Discrete signals can easily cause the detector to detect input samples with a high misjudgment rate. Since each word in a short text plays a more important role, the words in the clean sample also have a higher sensitivity, and this error is more obvious in short texts. To solve this problem, the change in the probability distribution of the softmax layer (that is, the confidence score of F for all category predictions) is used as another feature, which represents a more subtle difference between adversarial samples and clean samples. Perturbing words will cause the probability distribution to change, but this change is more obvious in adversarial samples. Even if the clean sample also causes inconsistent predicted labels due to perturbations, the change in probability distribution is not as large as that of adversarial samples. The similarity between probability distributions is calculated by using Jensen-Shannon divergence (JSD). The larger the JSD value, the more similar it is and the closer the distribution is. Conversely, it means that the distribution has changed a lot.

[0061] Step 4: Use the adversarial features as training data to train a binary adversarial detector.

[0062] In the training phase, the data is divided into a training set and a test set in a ratio of 8:2. In the testing phase, the input features are calculated by requiring the model k+1 times. The proposed detector does not rely on a specific machine learning model architecture implementation, so multiple machine learning models are trained and the effects of different models are compared. The method of the present invention does not rely on a specific model or a specific classification task, that is, the detector can be deployed as a plug-and-play plug-in of the target model to improve robustness. The adversarial feature extraction method designed by the present invention is based on the general characteristics of AE, so that it is not limited to a specific attack method.

[0063] Step 5: Input the text to be tested into the adversarial detector and output the result; for any text, repeat steps 2 and 3, extract features, and input the features into the adversarial detector trained in step 4; if the adversarial detector outputs 0, the text is judged to be a normal sample; if the output is 1, the text is judged to be an adversarial sample.

[0064] Preferably, in step 1, the process of generating adversarial samples includes constructing a new dataset T = {x j , x j * }, 0<j<t, where x j is a clean sample from D, x j * Some attack algorithm generates x j Corresponding adversarial examples; D = {x i ,y i },0<i<l,x i is the text in dataset D, y i is x i The corresponding label, that is, f(x i )=y i , x i It can be further expressed as x i =[w 1 , w 2 , ..., w i , ..., w n ], where w i Is the text x i The words in , n is the number of words, t is the number of samples in the new dataset T, and 1 is the number of samples in the dataset D.

[0065] Preferably, the method for constructing the dataset T includes randomly sampling a data (x, y) from the dataset D, and generating an adversarial sample x corresponding to x through the deep neural network model F. * =[w 1 * , w 2 * , ..., w n * ],in is an adversarial sample x * The words in the, n is the number of words, if the attack is successful, that is, f(x * )≠y, then (x, x * ) is added to T. Repeat step 1 until there are t text pairs in T; where random sampling from D is a sampling strategy without replacement.

[0066] Preferably, in step 2, the most important k words are selected from all texts x in T, including adversarial samples and clean samples, and are sorted, which is recorded as C(x).

[0067] Preferably, in step 3, for each word in C(x), the predicted labels before and after the word is deleted are obtained by accessing the deep neural network model F, and the words with inconsistent predicted labels before and after the word is deleted are defined as sensitive words, and vice versa. The deletion operation means replacing the original word with a token. For the pre-trained models Bert and RoBERTa, the replacement token is [MASK], and for the traditional DNNs models LSTM and CNN, the replacement token is <unk>.

[0068] Preferably, the signal set of the sensitive words is used to measure the sensitivity of the text x to F, expressed as S(x, f(x)), which is formulated as:

[0069]

[0070] in Mathematical expression for word sensitivity:

[0071]

[0072] in is the text x without word w i , f(x) represents the output function of the deep neural network model F.

[0073] Better, through JSD divergence To calculate the similarity between probability distributions, the larger the JSD divergence value, the more similar they are and the closer the distributions are. Conversely, a smaller value indicates that the distributions have changed significantly:

[0074]

[0075] where f s (x) represents the probability distribution of the softmax layer, KL stands for Kullback-Leibler divergence, which is expressed as:

[0076]

[0077] For each word in C(x), the JSD values ​​are calculated and used as the distribution variance features of x, expressed as:

[0078]

[0079] Preferably, the input feature E(x, f(x)) of the adversarial detector consists of the sensitive signal JSD value, expressed as:

[0080] E(x, f(x))=S(x, f(x))*J(x, f(x))

[0081] Therefore, the input features of the adversarial detector are a set of continuous vectors of size k, and the labels are binary, 0 represents clean samples and 1 represents adversarial samples.

[0082] Preferably, in step 4, during the training phase of the adversarial detector, the data is divided into a training set and a test set in a ratio of 8:2.

[0083] Preferably, in step 5, for any text, steps 2 and 3 are repeated to extract features, and the features are input into the adversarial detector trained in step 4; if the adversarial detector outputs 0, the text is judged to be a normal sample; if the output is 1, the text is judged to be an adversarial sample.

[0084] Embodiment 2:

[0085] In another embodiment, the present invention specifically conducts experiments based on the IMDB dataset and the BERT model. IMDB is a binary classification dataset of English movie reviews, with a total of 25,000 training data and 25,000 test data, and an average text length of 251 words. The BERT model is trained on the IMDB dataset; during the training process, the batch size is set to 32, the epochs is set to 10, and the dropout is set to 0.1.

[0086] Step 1: Generate a dataset T containing 1500 clean samples and 1500 adversarial samples. Randomly sample a sample (x, y) from the IMDB test set, and attack the trained BERT model using the existing attack method Textfooler. Some restrictions are imposed on the attack method: the maximum perturbation amplitude is 0.1, and the minimum similarity between the adversarial sample and its corresponding clean sample is 0.84. Under this restriction, generate an adversarial sample (x * ,y * ), if the attack is successful, that is, f(x * )≠y i , then the sample (x, x * ) are added to T. The sampling and attacking steps are repeated until there are 1500 sample pairs in T.

[0087] Step 2: For all text x in T (including adversarial samples and clean samples), select the k most important words in x. Input each sample into the model once to obtain the vectorized representation and gradient value of the sample. According to the formula Calculate the importance score of each word, where and The dimension of is 768, which is determined by the internal architecture of BERT. The words are sorted according to their importance scores to get the top 20 words, denoted as C(x). Subsequent feature extraction is performed around these 20 words.

[0088] Step 3: Extract features from 3000 samples in T. Assume x is any sample in T. For each word w in C(x), the text after replacing the word with [MASK] is expressed as Will Input into the trained BERT model, the output label is The probability distribution of the output is expressed as According to formula (4) and formula (5), the sensitivity tag of word w is calculated Similarity to probability distribution Then the feature value corresponding to word w is The feature representation of sample x is a vector composed of the feature values ​​of each word in C(x), denoted by E x =[e w1 , e w2 , ..., e w20 ]. If sample x is an adversarial sample, then label E y is 1, otherwise it is 0. Further, (E x , E y ) are the training samples of the detector.

[0089] Step 4: Train the detector by dividing the data into The present invention selects 5 different machine learning model architectures as detectors, and trains and tests the detectors. The results are shown in Table 1. The data in the table show that different architectures can effectively detect adversarial samples as detectors, among which XGBoost has a slightly higher detection ability than other models.

[0090] Table 1 Comparison of detection effects of different machine models

[0091] Machine learning Model Accuracv Recall F1-score Random Foorest 95.7 96.0 95.1 XGBoost 96.1 97.2 96.4 LightGBM 96.3 94.1 95.8 SVM 92.3 93.2 92.1 AdaBoost It should be noted that "Accuracv" in the original text seems to be a misspelling. It might be "Accuracy". 95.2 96.0 94.9

[0092] Step 5: Input the text to be tested into the adversarial detector and output the result; for any text, repeat steps 2 and 3, extract features, and input the features into the adversarial detector trained in step 4; if the adversarial detector outputs 0, the text is judged to be a normal sample; if the output is 1, the text is judged to be an adversarial sample.

[0093] The above-mentioned embodiments of the present invention solve the existing problems in adversarial detection technology to a certain extent, and improve the versatility and detection efficiency of adversarial detection while ensuring that the accuracy of adversarial sample recognition is not reduced.

[0094] The present invention extracts adversarial features based on the difference in perturbation sensitivity and trains a binary adversarial detector. The present invention only examines the changes in the output of the target model before and after the important words in the perturbation input sample, without processing all the words, wherein the important words are reflected through gradient information and only require one access to the model. This process ensures the pertinence of feature extraction and improves the efficiency of feature extraction. The difference in perturbation sensitivity is common between adversarial samples and clean samples. Therefore, the adversarial feature extraction technology proposed in the present invention is universal and can be extended to different attack scenarios, rather than being limited to specific attack methods.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention, rather than to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the essence and scope of the technical solution of the present invention.< / unk> < / unk> < / unk>

Claims

1. Adversarial sample detection method based on perturbation sensitivity difference, It is characterized in that The following steps are involved: Step 1: Generate adversarial samples using attack algorithms; Step 2: Determine important words using gradient estimation; Step 3: Perturb important words and extract adversarial features; Step 4: Use the adversarial features as training data to train a binary adversarial detector; Step 5: Input the text to be tested into the adversarial detector and output the result; In step 2, the most important k words are selected from all texts x in the data set T, including adversarial samples and clean samples, and sorted, which is recorded as C(x); In step 3, for each word in C(x), the predicted labels of the word before and after being deleted are obtained by accessing the deep neural network model, and the words with inconsistent predicted labels before and after the word is deleted are defined as sensitive words, otherwise they are non-sensitive words; The signal set of the sensitive words is used to measure the sensitivity of the text x to F, which is expressed as S(x,f(x)), and its formula is expressed as: in Mathematical expression for word sensitivity: in is the text x without word w i , f(x) represents the output function of the deep neural network model F; By JSD divergence To calculate the similarity between probability distributions, the larger the JSD divergence value, the more similar they are and the closer the distributions are. Conversely, a smaller value indicates that the distributions have changed significantly: where f s (x) represents the probability distribution of the softmax layer, KL stands for Kullback-Leibler divergence; For each word in C(x), the JSD divergence is calculated and used as the distribution variance features of x, expressed as: The input feature E(x,f(x)) of the adversarial detector consists of the sensitive signal JSD value, expressed as: E(x,f(x))=S(x,f(x))*J(x,f(x)) Therefore, the input features of the adversarial detector are a set of continuous vectors of size k, and the labels are binary, 0 represents clean samples and 1 represents adversarial samples.

2. The adversarial sample detection method based on perturbation sensitivity difference according to claim 1, It is characterized in that In the step 1, the process of generating adversarial samples includes constructing a new dataset T = {x j , x j *}, 0 < j < t, where x j is a clean sample from D, and x j * is the adversarial sample generated by a certain attack algorithm corresponding to x j ; D = {x i , y i}, 0 < i < l, x i is the text in dataset D, and y i is the label corresponding to x i , that is, f(x i ) = y i , and x i is further represented as x i = [w 1 , w 2 , …, w i , …, w n , where w i is the word in the text x i , n is the number of words, t is the number of samples in the new dataset T, and l is the number of samples in dataset D.

3. The adversarial sample detection method based on perturbation sensitivity difference according to claim 2, It is characterized in that The method for constructing the dataset T includes randomly sampling a data (x, y) from the dataset D, and generating an adversarial sample x corresponding to x through a deep neural network model F. * =[w 1 * ,w 2 * ,…,w n * ],in is an adversarial sample x * The words in, n is the number of words, if the attack is successful, that is, f(x * )≠y, then (x,x * ) is added to T; step 1 is repeated until there are t text pairs in T; wherein random sampling from D is a sampling strategy without replacement.

4. The adversarial sample detection method based on perturbation sensitivity difference according to claim 1, It is characterized in that In step 3, the deletion operation means replacing the original word with a token. For the pre-trained models Bert and RoBERTa, the replacement token is [MASK], and for the traditional DNNs models LSTM and CNN, the replacement token is <unk> 。< / unk> 5. The adversarial sample detection method based on perturbation sensitivity difference according to claim 1, It is characterized in that KL is expressed as:

6. The adversarial sample detection method based on perturbation sensitivity difference according to claim 1, It is characterized in that In step 4, during the training phase of the adversarial detector, the data is divided into a training set and a test set in a ratio of 8:

2.

7. The adversarial sample detection method based on perturbation sensitivity difference according to claim 1, It is characterized in that In step 5, for any text, steps 2 and 3 are repeated to extract features, and the features are input into the adversarial detector trained in step 4; if the adversarial detector outputs 0, the text is judged to be a normal sample; if the output is 1, the text is judged to be an adversarial sample.

Citation Information

Patent Citations

  • Word-level text confrontation sample detection method

    CN114169443A

  • Text adversarial sample generation method and system, computer equipment and storage medium

    CN114091448A

  • Quantifying Vulnerabilities of Deep Learning Computing Systems to Adversarial Perturbations

    US20200285952A1