A general adversarial attack detection method based on token loss information
Patent Information
- Application Number
- CN202311394745.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-25
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-10-25
AI Technical Summary
然而基于DARCY陷阱门的防御形式需要重新训练文本分类器,从而达到诱导的目的,这不符合现实场景中的对抗检测任务
[0028] (1) Effectively improves the detection of non-instance learnable adversarial samples, and can be detected by utilizing the differences in the TLV metric of the underlying token sequence in general attacks on both long and short adversarial texts.
Smart Images

Figure CN117407870B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the problem of adversarial samples in the field of natural language processing technology, and in particular to a detection method based on learnable general text adversarial attacks or a general adversarial attack detection method based on token loss information. Background Technology
[0002] Deep neural networks have achieved remarkable performance in various machine learning tasks, particularly in Natural Language Processing (NLP). However, the structural complexity and lack of interpretability of deep networks make them highly susceptible to subtle perturbations. Adversarial examples, created by adding minute, imperceptible perturbations to clean samples, can severely disrupt a model's recognition. Examples include adding undetectable masking information to images or making minor modifications to text. Therefore, defense against adversarial attacks has become a key research focus. Existing work on adversarial examples primarily focuses on computer vision tasks, but research on adversarial examples for NLP tasks has gradually emerged in recent years. For instance, in malicious attacks targeting spam, attackers can set specific perturbation information to mislead the spam filtering system's recognition results. Furthermore, in real-world NLP tasks such as SMS fraud, advertising, malicious comments, and public opinion detection, corresponding attacks can be created to mislead text classifiers. Current attacks against text classifiers can be categorized into instance-based and non-instance-based attacks. Non-instance attacks, which "learn" to generate adversarial sequences within existing deep neural networks, are a more efficient method. They add the same perturbation to batches of samples of a specific class, causing the text classifier to misclassify. Detection methods for non-instance text attacks are relatively scarce. Current research proposes using "honeypots" (DARCY) to set up multiple trapdoors in the network to induce the generation of attack tokens, thereby detecting non-instance attacks. However, defenses based on DARCY trapdoors require retraining the text classifier to achieve the induction, which is not suitable for real-world adversarial detection tasks. Summary of the Invention
[0003] Considering the shortcomings and deficiencies of existing technologies, this invention, from the perspective of token loss weighting, achieves detection of non-instance text attacks in the general text adversarial domain, effectively improving performance in the detection of non-instance learnable adversarial samples. It achieves optimal detection results in both long and short text detection and in the generalization of multiple text classifiers, providing a more effective reference mechanism for text adversarial detection and possessing high application value.
[0004] This invention addresses the UniTrigger, a learnable general non-instance text adversarial attack in natural language processing (NLP), by proposing an adversarial sample detection method based on token loss weight information. Specifically, it proposes a method for detecting learnable non-instance adversarial samples in NLP based on token loss weight information. This involves segmenting the target sample into independent token sequences, calculating the loss weight information for each token sequence and converting it into a TLV (Token-Loss Value) metric, and establishing a full sample sequence lookup table T. A difference detector threshold is then set to identify adversarial samples. This method effectively solves the problem of perturbation detection for learnable non-instance adversarial samples and achieves high performance on both long and short texts, reaching the optimal detection results for current tasks.
[0005] The specific technical solution adopted is as follows:
[0006] A general adversarial attack detection method based on token loss information, for attack detection tasks in natural language processing, characterized by the following steps;
[0007] Step S1: Receive the text sequence X input from the front end, and split the text sequence into X = [x1, x2, ... x...]. n ], forming an independent token sequence block, where x i This represents each segmented unit block in the text, and n represents the length of the text sequence.
[0008] Step S2: Convert the token sequence X into a 300-dimensional word vector V(x) = [V(x1), V(x2), ..., V(x3)] using the pre-trained model word2vec. n ]], each token sequence block x i Use V(x) i This indicates that a set of token sequences to be detected, V(x), is constructed. i )∈D, containing all token sequences to be detected;
[0009] Step S3: Assume that the k-th token sequence V(x) is detected. k If ), then the k-th V(x) is extracted from the sequence set D. k ) makes D dec =[V(x1),V(x2),····V(x k-1 ),V(x k+1 ),····,V(x n )], D dec Input to the text classifier F(X,Y) using Log(x) i The function represents the logit value of the objective parameters in the model, and obtains the gradient information.
[0010] Step S4, with L CE The cross-entropy loss function is expressed by the formula... Obtain the loss weight information TLV of the k-th token in the final input sample x; where index represents the tag bit in the full token sequence lookup table; calculate the TLV for each token sequence in the input sample and build the full sequence lookup table T;
[0011] Step S5: Input the full sequence lookup table T into the detector, and find the maximum and minimum index values based on the TLV metric, where diff * The diff represents the maximum difference value in the query list. * =argmax(T) - argmin(T);
[0012] Step S6: Diff * The input sample is used as the data representation in the adversarial detector, and a threshold α is set to limit the amplitude change. If it exceeds α, the input sample is judged to be an adversarial sample; otherwise, it is a clean sample.
[0013] Further, in step S1, the entire text sequence X to be detected is converted to lowercase for unified word representation, and the text is segmented into individual word tokens. Each token is assigned a unique numerical index ID number, which is then split into x using a token. i The form makes the text sequence = [x1, x2, ... x n Each of them has a corresponding serial number.
[0014] Further, in step S2, the text sequence with all assigned ID sequence numbers is transformed into a 300-dimensional word vector semantic space representation using the pre-trained language model Word2vec and equation (1); in equation (1), C represents the context-related length of the word vector setting, x k The target text sequence to be transformed is represented by V(x). For each text sequence to be tested, the word vector representation becomes V(x). i Furthermore, the set D contains the representations of all text sequences to be detected in the semantic space vector: V(x) = [V(x1), V(x2), ..., V(x...]. n )];
[0015] V(x k )=1 / (2C)*ΣV(x k-c )+∑V(x k+c (1).
[0016] Furthermore, in step S3, it is assumed that all the sequences to be detected are potential adversarial perturbations, therefore the k-th V(x) is obtained from the sequence set D. k), and obtain D in the semantic space D of the sample word vectors. dec =V(x1),V(x2),····V(x k-1 ),V(x k+1 ),····,V(x n )], D dec The input is fed into the target text classifier F(X,Y), using a BiLSTM model as the target text classifier, D dec The feature values of the input samples are obtained through the last layer of the BiLSTM using a bidirectional recurrent neural network; Log(x) i The function represents the logit value of the target parameter feature in the model, and equation (2) is used to calculate the gradient information. Then, the gradient information of the token to be detected is obtained;
[0017]
[0018] Further, in step S4, L CE The cross-entropy loss function L represents the gradient information of the token to be detected obtained in step S3, which is input into the cross-entropy loss function L. CE In the process, the loss weight information TLV of the k-th token of the sample to be detected X is obtained through equation (3); in the UniTrigger attack sequence, it is specifically represented by the corresponding metric TLV(k,L). CE (Log(x k ,f)));where index represents the position information of the text sequence token to be detected; the text sequence token x is measured by establishing a full sequence lookup table T. k The weighting of metrics in the entire sequence input sample;
[0019]
[0020] Furthermore, in step S5, the full sequence lookup table T specifically includes the text to be detected, X = [x1, x2, ... x...]. n Each token sequence in the text is transformed into a semantic space vector representation, which is then processed by a text classifier F(X,Y) to obtain logit information and converted into TLV information. The full sequence lookup table T displays the position of the corresponding token in the text sequence and its TLV value.
[0021] Further, in step S5, the full sequence lookup table T is input into the detector, and all TVL function values are retrieved from table T to find the maximum and minimum index values, which are then analyzed using diff. * The maximum difference value in the query list is represented as shown in equation (4); so that in the non-instance learnable text adversarial attack of UniTrigger, the artificially added adversarial perturbation sequence X'=[x'1,x'2,…,x'n After the processing steps S1-S5, the differences can be detected and identified.
[0022] diff * =argmax(Detecter(T))-argmin(Detecter(T)) (4).
[0023] Further, in step S6, an adversarial sample detector is set up, which receives each input text sequence sample to be detected, performs TLV metric, and retrieves the diff from the full sequence lookup table T. * The difference values are used to set a threshold α constraint for the adversarial sample detector; if the sample to be detected exceeds the threshold constraint, it is judged as an adversarial sample, otherwise it is a clean sample, as shown in equation (5):
[0024] x→{x∣x∈adv}or x→{x∣x∈ori} (5).
[0025] Furthermore, in step S6, to address the issue that during classifier training, the limited token sequence in short text sequences leads to some tokens having a high weight in the word vector space, resulting in samples whose semantic features have not been well learned by the classifier being misclassified by the detector, a constraint on the minimum TLV of short sequence tokens is added to the short data samples. It is added to the detector as auxiliary information for judging adversarial samples.
[0026] Considering that existing attacks are mainly divided into instance-based attacks and learning-based general non-instance attacks, especially general trigger attacks represented by UniTrigger (Universal Trigger), this invention and its preferred scheme generate a fixed attack sequence that reduces the prediction accuracy of the target model to near zero. Existing detection methods mainly target instance-based adversarial samples and cannot effectively prevent the impact of general trigger attacks on the model. Based on these problems, this invention achieves detection of non-instance text attacks in the general text adversarial domain, effectively improving performance in the task of detecting non-instance learnable adversarial samples. It utilizes the token sequence of adversarial text as input to the model and transforms it into loss weight information. It proposes using TLV to measure the token sequence, representing it numerically, and establishing a full sequence lookup table T to traverse all token information of the input test sample. The lookup table T is designed to have significant differences compared to the original token sequence, thus enabling adversarial sample detection by adjusting the detector's threshold.
[0027] Compared with the prior art, the present invention and its preferred embodiments also have the following beneficial effects:
[0028] (1) Effectively improves the detection of non-instance learnable adversarial samples, and can be detected by utilizing the differences in the TLV metric of the underlying token sequence in general attacks on both long and short adversarial texts.
[0029] (2) It effectively enhances the practicality of the detection method. The detection method does not require adversarial training of the model. The defense by directly detecting samples is more in line with the adversarial detection task in real-world scenarios.
[0030] (3) Effectively improves the generalization of the model. The detection method has high performance for multiple attacked models and can maintain a high detection rate for classifiers of different precision. Attached Figure Description
[0031] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0032] Figure 1 This is a flowchart of the method according to an embodiment of the present invention;
[0033] Figure 2 This is a schematic diagram illustrating the principle of the method in an embodiment of the present invention;
[0034] Figure 3 This is a schematic diagram of the detection results of the detector threshold of the embodiment of the present invention on four long and short text datasets, where TPR represents the correct detection rate and FPR represents the false detection rate.
[0035] Figure 4 This is a schematic diagram illustrating the detection results of the method of this invention on adversarial samples of four different text types of different lengths. Detailed Implementation
[0036] To make the features and advantages of this patent more apparent and understandable, specific embodiments are provided below, along with accompanying drawings, for detailed explanation:
[0037] like Figures 1-4 As shown, this embodiment of the invention provides an adversarial sample detection method based on token loss weight information for attack detection tasks in natural language processing, including the following steps;
[0038] Step S1: Receive the text sequence X input from the front end, and split the text sequence into X = [x1, x2, ... x...]. n ], forming an independent token sequence block, where x i Each segment represents a block of text, and n represents the length of the text sequence. To evaluate the response of model F(X,Y) to a given input, the token loss metric TLV is used to reflect its impact.
[0039] Step S2: Convert the token sequence X into 300-dimensional word vectors V(x) = V(x1), V(x2), ..., V(x3) using a pre-trained model (word2vec). n ]], each token sequence block x i Use V(x) i This indicates that a set of token sequences to be detected, V(x), is constructed. i )∈D, which contains all token sequences to be detected;
[0040] Step S3: Assume that the k-th token sequence V(x) is detected. k If ), then the k-th V(x) is extracted from the sequence set D. k ) makes D dec =[V(x1),V(x2),····V(x k-1 ),V(x k+1 ),····,V(x n )], D dec Input to the text classifier F(X,Y) using Log(x) i The function represents the logit value of the objective parameters in the model, and can be used to calculate gradient information.
[0041] Step S4, L CE The cross-entropy loss function is expressed as follows: The above variables are expressed using the formula... The loss weight information TLV of the k-th token in the final input sample X is obtained. Here, index represents the tag bit in the full token sequence lookup table. The TLV is calculated for each token sequence in the input sample to form a full sequence lookup table T;
[0042] Step S5: Input the full sequence lookup table T into the detector, and find the maximum and minimum index values based on the TLV function value, where diff * The diff represents the maximum difference value in the query list. * =argmax(T) - argmin(T);
[0043] Step S6: Considering that the token sequence of non-UniTrigger attacks fluctuates less in the TLV metric, the diff is... * The input sample is used as the data representation in the adversarial detector, and a threshold α is set to limit the magnitude change. If the value exceeds this threshold, the input sample is determined to be an adversarial sample; otherwise, it is considered a clean sample.
[0044] In this embodiment, preferably, in step S1, the entire text sequence X to be detected is converted to lowercase for unified word representation, and the text is segmented into individual word tags, with each tag assigned a unique numerical index ID number, which is then split into x using a token. i The form makes the text sequence = [x1, x2, ... x n Each of them has a corresponding serial number.
[0045] In step S2, the text sequence with all assigned ID sequence numbers is transformed into a 300-dimensional word vector semantic space representation using a pre-trained language model (Word2vec) based on equation (1). In equation (1), C represents the context-dependent length of the word vector setting, and x... k This represents the target text sequence to be transformed. The word vector representation for each text sequence becomes V(x). i Furthermore, the set D contains the representations of all text sequences to be detected in the semantic space vector: V(x) = [V(x1), V(x2), ..., V(x...]. n )).
[0046] V(x k )=1 / (2C)*∑V(x k-c )+∑V(x k+c (1)
[0047] In step S3, it is assumed that all the sequences to be detected are potential adversarial perturbations, therefore the k-th V(x) is obtained from the sequence set D. k ), and obtain D in the semantic space D of the sample word vectors. dec =[V(x1),V(x2),····V(x k-1 ),V(x k+1 ),····,V(x n )], D dec The input is fed into the target text classifier F(X,Y). In this example, a BiLSTM model is used as the target text classifier. dec The feature values of the input samples are obtained through the last layer of the BiLSTM using a bidirectional recurrent neural network. Log(x) i The function represents the logit value of the target parameter feature in the model, and we can obtain equation (2) to calculate the gradient information. Obtain the gradient information of the token to be detected.
[0048]
[0049] In step S4, L CE The cross-entropy loss function L represents the gradient information of the token to be detected obtained in step S3, which is input into the cross-entropy loss function L.CE In this process, the above variables are used to obtain the loss weight information TLV (Token-Loss Value) of the k-th token of the sample X to be detected through equation (3). In the UniTrigger attack sequence, this is specifically represented by the corresponding metric TLV(k,L). CE (Log(x k ,f))). Here, index represents the position information of the text sequence token to be detected. This method establishes a full sequence lookup table T to measure the text sequence token x. k The weighting of metrics in the full sequence input sample.
[0050]
[0051] In step S5, the full sequence lookup table T is used, specifically containing the text to be detected X = [x1, x2, ... x...]. n Each token sequence in the text is transformed into a semantic space vector representation. After passing through the text classifier F(X,Y), the logit information is obtained and then transformed into TLV information. The full sequence lookup table T shows the position of the corresponding token in the text sequence and the TLV value. Table 1 below gives an example of a full sequence lookup table T, where the bold part is the position with the highest loss percentage for that token in the example sample table.
[0052]
[0053] Table 1 shows the input samples converted to TLV metrics, and a full sequence lookup table T is established.
[0054] In step S5, the full sequence lookup table T is input into the detector. All TVL function values are retrieved from table T to find the maximum and minimum index values, and then diffed. * This represents the maximum difference value in the query list, specifically as shown in equation (4). Considering that the token sequence of a non-UniTrigger attack exhibits relatively small fluctuations in the TLV metric, while the adversarial sample shows greater differences in the full sequence lookup table of the TVL metric, this indicates that in a non-instance learnable text adversarial attack using UniTrigger, the artificially added adversarial perturbation sequence X'=[x'1,x'2,…,x' n After undergoing the processing flow of this invention, the differences exhibited can be detected and identified.
[0055] diff * =argmax(Detecter(T))-argmin(Detecter(T)) (4)
[0056] In step S6, an adversarial sample detector is set up. This detector receives each input text sequence sample to be detected, performs a TLV metric, and calculates the diff in the full sequence lookup table T. * To assess the difference in performance, this invention sets a threshold α constraint for the adversarial sample detector. If the sample to be detected exceeds the threshold constraint, it is judged as an adversarial sample; otherwise, it is a clean sample, as shown in equation (5).
[0057] x→{x∣x∈adv}or x→{x∣x∈ori} (5)
[0058] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0059] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0060] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0061] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0062] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
[0064] This patent is not limited to the above-described preferred embodiments. Anyone can derive other forms of general adversarial attack detection methods based on token loss information under the guidance of this patent. All equivalent changes and modifications made within the scope of this patent application shall fall within the scope of this patent.
Claims
1. A token loss information based universal adversarial attack detection method for attack detection tasks in natural language processing, characterized in that: Includes the following steps; Step S1, receiving a text sequence input by a front end , the text sequence is split into , forming independent token sequence blocks, wherein represents each split unit block in the text, represents the length of the text sequence; Step S2, for the token sequence Converted into 300-dimensional word vectors using the pre-trained word2vec model. Each token sequence block use This means constructing a set of token sequences to be detected. It contains all token sequences to be detected; Step S3, assuming detecting the first token sequence , then extracting the first token sequence in the sequence set D , inputting it into the text classifier , and obtaining the target parameter logit value in the model represented by the function ; specifically: Assuming all the sequences to be detected are potential adversarial perturbations, therefore in the sequence set Get the first indivual In the semantic space of sample word vectors From ,Will Input to target text classifier The BiLSTM model is used as the target text classifier. The feature values of the input samples are obtained through the last layer of the BiLSTM using a bidirectional recurrent neural network. The function represents the logit value of the target parameter feature in the model, and equation (2) is used to calculate the gradient information. Then, the gradient information of the token to be detected is obtained; (2) Step S4, with The cross-entropy loss function is expressed by the formula... Get the final input sample No. Individual token loss weight information TLV; Where index represents the tag bit in the full token sequence lookup table; calculate TLV for each token sequence in the input sample and build the full sequence lookup table T; In step S4, The cross-entropy loss function takes the gradient information of the token to be detected obtained in step S3 and inputs it into the cross-entropy loss function. In the middle, the sample to be tested is obtained through equation (3). The The loss weight information TLV for each token; specifically represented as a corresponding metric value in the UniTrigger attack sequence. The index represents the location information of the text sequence token to be detected; the text sequence token is measured by establishing a full sequence lookup table T. The weighting of metrics in the entire sequence input sample; (3) Step S5: Input the full sequence lookup table T into the detector, and find the maximum and minimum index values based on the TLV metric. Represents the maximum difference value in the query list. ; Step S6, The data representation of the input samples in the adversarial detector is used, and a threshold for amplitude variation is set. Limitations; if exceeded If the input sample is considered an adversarial sample, it is considered a clean sample otherwise.
2. The general adversarial attack detection method based on token loss information according to claim 1, characterized in that: In step S1, the text sequence to be detected is... All text is converted to lowercase for unified word representation, and the text is segmented into individual word tags, each assigned a unique numerical index ID number, which is then split into tokens. Form makes text sequences Each of them has a corresponding serial number.
3. The general adversarial attack detection method based on token loss information according to claim 2, characterized in that: In step S2, the text sequence with all assigned ID sequence numbers is transformed into a 300-dimensional word vector semantic space representation using the pre-trained language model Word2vec and equation (1); in equation (1) This represents the context-dependent length of the word vector settings. The target text sequence to be transformed is represented by the word vector representation of each text sequence to be tested. And in the set It contains the representations of all text sequences to be detected in the semantic space vector. ; (1)。 4. The general adversarial attack detection method based on token loss information according to claim 1, characterized in that: In step S5, the full sequence lookup table T specifically contains the text to be detected. Each token sequence is transformed into a semantic space vector representation, which is then processed by a text classifier. The logit information is obtained and converted into TLV information. The full sequence lookup table T shows the position of the corresponding token in the text sequence and its TLV value.
5. The general adversarial attack detection method based on token loss information according to claim 4, characterized in that: In step S5, the full sequence lookup table T is input into the detector. All TVL function values are retrieved from table T to find the maximum and minimum index values. The maximum difference value in the query list is represented as shown in equation (4); This allows for the artificially added adversarial perturbation sequences in non-instance learnable text adversarial attacks using UniTrigger. After the processing steps S1-S5, the differences can be detected and identified: (4)。 6. The general adversarial attack detection method based on token loss information according to claim 5, characterized in that: In step S6, the adversarial sample detector is set. Each input text sequence sample to be detected is measured by TLV and then entered into the full sequence lookup table T. The difference values are displayed, and an adversarial sample detector is set up. threshold Constraints; if the sample to be detected exceeds the threshold constraint, it is judged as an adversarial sample, otherwise it is a clean sample, as shown in equation (5): or (5)。 7. The general adversarial attack detection method based on token loss information according to claim 6, characterized in that: In step S6, to address the issue that during classifier training, the limited token sequence in short text sequences leads to some tokens having a high weight in the word vector space, causing samples whose semantic features haven't been well learned by the classifier to be misclassified by the detector, a constraint on the minimum TLV of short sequence tokens is added to the short data samples. It is added to the detector as auxiliary information for judging adversarial samples.